Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Anthropic’s May 21, 2024 study used dictionary learning to identify millions of recurring activation patterns—called features—in a middle layer of Claude 3 Sonnet. Researchers found patterns associated with subjects from San Francisco and lithium to secrecy and inner conflict, then experimentally changed selected features and observed changes in the model’s responses. The result is a rough conceptual map of some internal states, not a complete picture of the model or a transcript of its thoughts.

What does it mean to map a language model’s mind?

A language model processes information through internal numerical states. Those states include activations across many neurons, but an individual neuron does not necessarily correspond to one clear idea: a concept may involve many neurons, and a neuron may participate in representing more than one concept.

Anthropic’s researchers applied dictionary learning to activations in a middle layer of Claude 3 Sonnet. The method identifies activation patterns that recur across different contexts. They called the resulting patterns features, using them as candidate, more interpretable descriptions of parts of the model’s internal state.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The researchers offer an analogy: features combine neurons somewhat as words combine letters. That is a way to picture the method, not a literal description of the model’s architecture. A feature label is an interpretation supported by examples of when it activates; it does not prove that the model represents a concept in exactly the way a person understands it.

Anthropic describes the work as a rough conceptual map. It does not reveal everything the model learned, explain every response, or directly expose a human-like inner monologue. Anthropic’s May 21, 2024 research article describes the method and its examples.

What kinds of features did the researchers find?

Anthropic reports millions of features in the studied middle layer of Claude 3 Sonnet. Its article gives no precise count, so “millions” is the appropriate level of specificity. Examples range from particular entities to abstract patterns:

  • Places and people: San Francisco, the Golden Gate Bridge, and Rosalind Franklin.
  • Science and technology: lithium, immunology, and programming syntax.
  • Abstract or behavioral concepts: code bugs, gender bias, secrecy, and inner conflict.

Some features reportedly responded not just to names, but also to images and descriptions in multiple languages. That breadth is evidence of varied activation contexts in this study; it does not establish that the features capture every way those concepts can be represented.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How were relationships between features explored?

The researchers looked for nearby features using a distance measure based on overlap among the neurons in their activation patterns. In this representation, features near the Golden Gate Bridge feature included ones associated with Alcatraz, Ghirardelli Square, the Golden State Warriors, Gavin Newsom, the 1906 earthquake, and Vertigo.

Near an inner-conflict feature, they reported patterns involving relationship breakups, conflicting allegiances, logical inconsistencies, and “catch-22.” These neighborhoods show relationships according to the study’s chosen representation and distance measure. They are not proof of a complete semantic map or of concepts organized in a way identical to human understanding.

Did changing features change Claude’s responses?

Yes. Anthropic reports experiments in which researchers artificially amplified or suppressed selected features and observed changes in responses. This is distinct from identifying a feature that merely activates alongside a subject: the intervention results suggest that manipulating particular features can causally affect behavior in the reported experiments.

Golden Gate Bridge feature

When researchers amplified the Golden Gate Bridge feature, Claude began identifying as the bridge and bringing it up in unrelated answers. This illustrates how strongly changing a selected internal pattern could affect output under the experimental setup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scam-email feature

The researchers also describe activating a feature associated with scam emails strongly enough that Claude generated a scam email, despite ordinarily refusing that request. Anthropic says ordinary users cannot strip safeguards and manipulate models this way.

These interventions support a limited conclusion: selected features can influence responses under the conditions tested. They do not show that researchers can control every behavior, or that the model’s full internal process is understood.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What does this reveal about safety?

The study reported features associated with code backdoors, biological weapons, gender discrimination, racist claims about crime, power-seeking, manipulation, secrecy, and sycophantic praise. A feature associated with a behavior is not proof that the model will exhibit that behavior in ordinary use. Anthropic specifically cautions that finding a sycophantic-praise feature does not mean Claude will necessarily be sycophantic.

In principle, identifying internal patterns could help researchers monitor behavior, steer models, or evaluate safety. But the study does not establish that these features have already produced a safety improvement. Anthropic presents such applications as possibilities for future work, not validated outcomes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What are the study’s limits?

  • Limited scope: The report concerns a middle layer of Claude 3 Sonnet. It should not be generalized automatically to every layer, every Claude model, or all large language models.
  • Incomplete coverage: Anthropic says, “The features we found represent a small subset of all the concepts learned by the model during training.” Finding a fuller set with the current approach would be prohibitively expensive; the article says the required computation would vastly exceed the compute used to train the model.
  • Unresolved mechanisms: Researchers still need to understand the circuits in which features participate. Identifying a feature is not the same as explaining the full mechanism that produces a response.
  • Unproven safety benefit: Whether safety-relevant features can actually be used to improve safety remains an open question in the report.

For those reasons, the work is significant as both a descriptive study and a set of reported causal interventions, but it is not a complete map of a model’s internal representations or a demonstration that interpretability has solved model safety.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.