That is why the “microscope” metaphor matters. Anthropic is not claiming to uncover a hidden paragraph of private chain-of-thought. It is trying to build tools that let researchers inspect pieces of the computation underneath Claude’s written answers .
Anthropic’s earlier interpretability work focused on locating interpretable concepts inside a model, which it calls “features” . In practical terms, a feature is a handle on a pattern of internal activity that researchers can name, inspect, and test, rather than treating the model as a wall of opaque numbers .
This is the first layer of the map: instead of asking only what Claude said, researchers try to identify which internal concepts became active while Claude was generating that response .
The newer step is connecting those features into computational “circuits.” Anthropic describes this as extending feature-level interpretability to reveal parts of the pathway that transforms the words going into Claude into the words coming out .
That distinction is important. A single feature may show that a concept is present somewhere inside the model, but a circuit can help show how multiple internal components influence one another during a response . For reasoning-like behavior, the pathway matters as much as the individual concepts.
In March 2025, Anthropic said it was sharing two papers: one extending its feature work into circuit tracing, and another applying the toolset to Claude 3.5 Haiku . The Claude 3.5 Haiku study looked at simple tasks representative of ten crucial model behaviors, which Anthropic framed as part of studying “AI biology” .
The phrase “AI biology” signals the kind of understanding Anthropic is pursuing. Rather than only evaluating Claude from the outside—by looking at whether an answer is correct, fluent, or safe—the company is trying to identify internal mechanisms that help explain why a model behaves the way it does .
Claude’s written explanation is still generated text. Anthropic’s interpretability work targets the underlying computations that help produce that text in the first place .
That makes circuit tracing a different kind of evidence. It is not a prompt asking the model to describe its reasoning. It is an attempt to inspect parts of the computational pathway directly, using tools designed to translate neural activity into more legible structures .
The work can make some of Claude’s internals more legible: which features appear relevant, how those features are connected, and which pathways seem involved in producing a response . It can also give researchers a way to compare surface behavior with internal mechanisms, instead of relying only on final outputs .
But Anthropic’s own framing is cautious. The papers are described as progress toward a microscope and as revealing “parts” of the pathway from input words to output words . That means the current tools should not be treated as a complete decoder for every computation in Claude, or as a reliable transcript of everything the model “thinks” internally .
Anthropic is making Claude’s hidden reasoning more understandable by translating some internal activations into interpretable features, tracing how those features interact as circuits, and applying that map to concrete model behaviors . The result is a partial scientific map of Claude’s computation—not full mind-reading, and not a complete explanation of every answer .