Definition: Mechanistic interpretability is the research field that reverse-engineers the computation inside a neural network. Instead of asking only what a model outputs, it asks which internal features switch on, how they connect into circuits, and which of those circuits actually cause the answer. The goal is an explanation you can test, the way a biologist tests a proposed pathway in a cell.
TL;DR: Mechanistic interpretability reads the inside of an AI model. In 2024 Anthropic pulled up to 34 million features out of Claude 3 Sonnet, including safety-relevant ones such as sycophantic praise that change its output when turned up. In 2025 circuit tracing showed a model planning rhymes ahead. MIT Technology Review named the field a 2026 breakthrough technology. Build one free →
Think of a mechanic who can hear that an engine runs rough, and a second mechanic who opens the hood, traces the fuel line, and points at the clogged injector. Both noticed a problem. Only the second one can tell you why it happened and what to fix. Most AI evaluation is the first mechanic: it watches outputs. Mechanistic interpretability is the second mechanic, and for large language models the hood has only started to open.
Why Mechanistic Interpretability Matters in 2026
Mechanistic interpretability matters because the most important AI failures are the ones an output test cannot see. A model can give a correct answer for the wrong reason, or explain its reasoning in a way that does not match what it computed. In January 2026, MIT Technology Review put mechanistic interpretability on its list of 10 Breakthrough Technologies, and cited Anthropic's "microscope" that finds features for concepts such as Michael Jordan and the Golden Gate Bridge.
The results from 2025 moved the field from finding single features to tracing whole circuits inside a production model. In March 2025, Anthropic's Tracing the Thoughts of a Large Language Model reported that Claude uses a shared conceptual space across English, French, and Chinese, that it picks a rhyming word before it writes the line that ends on it, and that it sometimes writes a plausible chain of reasoning that works backward from a hint instead of from the problem. That last result is why a readable chain of thought is not the same as a faithful one.
Labs now tie safety plans to this work. In The Urgency of Interpretability (April 2025), Anthropic CEO Dario Amodei wrote that Anthropic aims to reach the point where "interpretability can reliably detect most model problems" by 2027, and described the long-run aim as a true "MRI for AI". Three months earlier, a January 2025 survey, Open Problems in Mechanistic Interpretability, led by Lee Sharkey with 28 co-authors, grouped what still stands in the way into three areas: methods that need conceptual and practical improvement, open questions about how to apply them to specific goals, and socio-technical issues around the field.
Key Terms in Mechanistic Interpretability
Mechanistic interpretability has its own vocabulary, and most of the words describe one idea: a model stores more concepts than it has neurons, so you need a tool to pull them apart.
| Term | Plain meaning | Where it comes from |
|---|---|---|
| Feature | A direction inside the model that stands for one concept, such as "Golden Gate Bridge" or "code error" | Core unit of the field |
| Superposition | The model packs more features than it has dimensions, so features share neurons | Toy Models of Superposition, Elhage et al., 2022 |
| Polysemantic vs monosemantic | A polysemantic neuron fires for many unrelated concepts. A monosemantic unit fires for one | Superposition explains polysemanticity |
| Sparse autoencoder (SAE) | A second network trained to split activations into many sparse, readable features | Scaling Monosemanticity, Templeton et al., 2024 |
| Induction head | An attention circuit that completes [A][B] … [A] with [B] |
Olsson et al., 2022 |
| Circuit | A group of features and connections that together carry out one computation | The unit that circuit tracing maps |
| Transcoder | A readable stand-in for a model's MLP neurons, built from sparse features, used to trace flow between layers | Circuit Tracing, Anthropic, 2025 |
| Attribution graph | A map of which features pushed a specific output, with low-influence nodes and edges pruned away | Circuit Tracing, Anthropic, 2025 |
| Replacement model | The readable copy of the model that a graph actually describes. Error nodes mark the part it cannot explain | Circuit Tracing, Anthropic, 2025 |
How Mechanistic Interpretability Works
Mechanistic interpretability works in a fixed order: record what the model does inside, decompose it into readable features, connect those features into a graph for one prompt, then intervene to prove the graph is causal.
- Record the activations. Researchers run text through a large language model and save the internal numbers at each layer of the transformer.
- Decompose them into features. Raw neurons are polysemantic, so a sparse autoencoder learns a larger set of features where only a few are active at once. Anthropic's Scaling Monosemanticity trained SAEs at roughly 1 million, 4 million, and 34 million features on the middle-layer residual stream of Claude 3 Sonnet.
- Label the features. Each feature gets a name from the text that turns it on. The same paper found safety-relevant features tied to security vulnerabilities, bias, lying, deception, and power-seeking, sycophancy, and dangerous content. Its authors warn that a feature for lies is not proof of lying: knowing about lies, being able to lie, and actually lying are different things.
- Build the graph. For circuit tracing on Claude 3.5 Haiku, Anthropic replaced the model's MLP neurons with a cross-layer transcoder of 30 million features, then drew the path from input to output for single prompts.
- Intervene to prove cause. Researchers clamp a feature up or down. When Anthropic amplified the Golden Gate Bridge feature, the model described itself as the bridge. When it turned up the sycophantic praise feature, the model answered an overconfident user with flattery instead of a truthful correction.
- Report the limits. Anthropic states that its method "only captures a fraction of the total computation" and that it takes "a few hours of human effort" to understand the circuits for a prompt of tens of words. In the companion biology paper, the team reports that its attribution graphs gave satisfying insight for about a quarter of the prompts it tried.
Mechanistic Interpretability Milestones
Mechanistic interpretability went from small attention-only models to production models in about three years.
The induction-heads paper (Olsson et al.) showed that these heads form during a sudden "phase change" early in training, at the same point that in-context learning improves sharply. On May 29, 2025, Anthropic open-sourced its circuit-tracing library with an interactive frontend on Neuronpedia, so anyone can generate attribution graphs for open-weight models such as Gemma-2-2b and Llama-3.2-1b.
Limits and Open Problems in Mechanistic Interpretability
Mechanistic interpretability is still a partial view. Every flagship result so far comes with a limit its authors state in the same paper, and those limits decide how much weight a finding can carry.
| Limit | What the source reports | Why it matters |
|---|---|---|
| Incomplete dictionaries | Claude 3 Sonnet can name streets in every London borough, but only about 60% of the boroughs had a matching feature in the 34M SAE (Scaling Monosemanticity) | A missing feature does not mean a missing concept |
| Graphs describe a copy | Attribution graphs describe the local replacement model, "which may differ from the underlying model" (Circuit Tracing) | A graph is a hypothesis about the real model, not a readout |
| Low hit rate | Satisfying insight for about a quarter of prompts tried (biology paper) | Most prompts still resist a clean explanation |
| Human cost | A few hours of expert work for a prompt of tens of words (Anthropic) | Checking every production answer this way is out of reach today |
| Feature is not behavior | Knowing about lies, being able to lie, and lying are different (Scaling Monosemanticity) | A "deception feature" alone is not evidence of deception |
What each method can tell you about one answer
Output test -> "The answer is wrong."
Written reasoning -> "Here is why I said it."
(a claim, can be unfaithful)
Attribution graph -> "These features pushed it."
(describes a replacement model)
Intervention -> "Turn feature X down and the answer changes."
(causal, for this one case)
Mechanistic Interpretability vs Chain-of-Thought Monitoring
Chain-of-thought monitoring reads the reasoning a model writes out. Mechanistic interpretability reads the computation underneath it. Both are safety tools. Models can write reasoning that does not match their internal path, so treat written reasoning as a claim and an interpretability result as a tested hypothesis about one case.
| Question | Chain-of-thought monitoring | Mechanistic interpretability |
|---|---|---|
| What you read | The reasoning text the model writes | Internal features and circuits |
| Cost per check | Low, a reviewer or second model reads text | High, hours of expert work per prompt today |
| Scale | Every run, in production | Selected prompts and research cases |
| Main blind spot | The written reasoning can be unfaithful | Captures only part of the computation |
| Real use | OpenAI used it to catch one of its reasoning models cheating on coding tests, per MIT Technology Review | Found safety-relevant features, such as sycophantic praise, that steer output when amplified |
Connection to Taskade
Mechanistic interpretability shows how hard it is to read knowledge once it lives inside model weights, even for the lab that trained the model. Taskade keeps what your agents know outside the weights, where you can read it. You train a Taskade AI Agent on knowledge you choose: files, links, projects, and media. Taskade EVE, the agent that builds Taskade Genesis apps, saves its notes as real Taskade projects in a projects/memory folder that you can open, edit, share, or delete. Agents run on frontier models from top AI labs, with Auto as the default. Taskade does not make those models transparent inside. It makes the context they act on inspectable, which is the part your team controls. This is Workspace DNA in practice: Memory holds facts you can audit, Intelligence reasons over them, and Execution writes results back into projects. When an answer looks wrong, you open the project it came from.

What You Would Build in Taskade
You already audit people this way. When a new analyst makes a strange call, you ask which document they read and which rule they followed, and then you fix the document.
In Taskade you would describe an agent knowledge audit board. Each row is one answer an agent gave. The row stores the question, the answer, and the projects and files the agent had as knowledge. A reviewer agent, briefed only with your checklist, marks whether each claim traces back to a source in the workspace. Rows it cannot trace get assigned to a teammate for a human in the loop review, and the fix goes into the knowledge project, not into a hidden setting. Over time the board shows which sources your agents lean on and which ones need updates.
Describe yours and build it free →
Related Concepts
- What Is Mechanistic Interpretability?: the full guide, with deeper history and examples
- Neural Network: the system that mechanistic interpretability takes apart
- Attention Mechanism: the part of a transformer where induction heads live
- AI Safety and Alignment: the wider goal that interpretability serves
- Connectome: the neuroscience parallel, where a full wiring map still does not explain function
- Superposition: how a layer holds more features than dimensions
- AI Sycophancy: a behavior that interpretability found as a feature inside a model
- Reward Hacking: the kind of hidden shortcut that output tests miss and internal checks aim to catch
- Chain-of-Thought Faithfulness: why written reasoning can differ from the real computation
Frequently Asked Questions About Mechanistic Interpretability
What is mechanistic interpretability in simple terms?
Mechanistic interpretability is the effort to read the inside of an AI model. Researchers find the internal features that stand for concepts, connect them into circuits, and test whether those circuits cause the model's answer. The aim is to explain how a model reaches an output, not only what the output is.
How is mechanistic interpretability different from explainable AI?
Many explainable AI methods score which parts of the input mattered for an output. Mechanistic interpretability goes further and tries to identify the actual algorithm inside the network, step by step. It then checks each step by intervening on the model, so the explanation is a tested mechanism and not only a correlation.
What is a sparse autoencoder in AI interpretability?
A sparse autoencoder is a small network trained to rewrite a model's internal activations as a large set of features, where only a few are active at a time. That sparsity makes each feature easier to name. Anthropic's 2024 work trained sparse autoencoders with up to 34 million features on Claude 3 Sonnet.
What are polysemantic neurons?
A polysemantic neuron fires for several unrelated concepts. The 2022 paper Toy Models of Superposition explained why: a model stores more features than it has neurons, so features share neurons. Sparse autoencoders and transcoders try to recover monosemantic features that each stand for a single concept.
What did Anthropic find with circuit tracing?
In March 2025 Anthropic traced circuits in Claude 3.5 Haiku. It found a shared concept space across English, French, and Chinese, planning of rhyming words before a line is written, parallel paths for mental math, and cases where the model worked backward from a hint so its written reasoning did not match the internal computation. The team reports satisfying insight for about a quarter of the prompts it tried.
What are the limits of mechanistic interpretability today?
Anthropic states that its circuit-tracing method captures only a fraction of the model's computation and needs a few hours of human effort per prompt of tens of words. Its graphs describe a replacement model that can differ from the real one, so a graph is a hypothesis to test. Even a 34-million-feature dictionary missed about 40% of the London boroughs the model knows.
How does Taskade keep AI agent knowledge inspectable?
Taskade keeps agent knowledge in your workspace instead of in model weights. You give an agent files, links, and projects as knowledge, and Taskade EVE saves its memory as readable projects that you can open and edit. When an answer looks wrong, you can trace it to a source and fix that source directly.