AI Concepts

Mechanistic Interpretability

14 min read
On this page (17)

Definition: Mechanistic interpretability is the research field that reverse-engineers the computation inside a neural network. Instead of asking only what a model outputs, it asks which internal features switch on, how they connect into circuits, and which of those circuits actually cause the answer. The goal is an explanation you can test, the way a biologist tests a proposed pathway in a cell.

TL;DR: Mechanistic interpretability reads the inside of an AI model. In 2024 Anthropic pulled up to 34 million features out of Claude 3 Sonnet, including safety-relevant ones such as sycophantic praise that change its output when turned up. In 2025 circuit tracing showed a model planning rhymes ahead. MIT Technology Review named the field a 2026 breakthrough technology. Build one free →

Think of a mechanic who can hear that an engine runs rough, and a second mechanic who opens the hood, traces the fuel line, and points at the clogged injector. Both noticed a problem. Only the second one can tell you why it happened and what to fix. Most AI evaluation is the first mechanic: it watches outputs. Mechanistic interpretability is the second mechanic, and for large language models the hood has only started to open.

Why Mechanistic Interpretability Matters in 2026

Mechanistic interpretability matters because the most important AI failures are the ones an output test cannot see. A model can give a correct answer for the wrong reason, or explain its reasoning in a way that does not match what it computed. In January 2026, MIT Technology Review put mechanistic interpretability on its list of 10 Breakthrough Technologies, and cited Anthropic's "microscope" that finds features for concepts such as Michael Jordan and the Golden Gate Bridge.

The results from 2025 moved the field from finding single features to tracing whole circuits inside a production model. In March 2025, Anthropic's Tracing the Thoughts of a Large Language Model reported that Claude uses a shared conceptual space across English, French, and Chinese, that it picks a rhyming word before it writes the line that ends on it, and that it sometimes writes a plausible chain of reasoning that works backward from a hint instead of from the problem. That last result is why a readable chain of thought is not the same as a faithful one.

Labs now tie safety plans to this work. In The Urgency of Interpretability (April 2025), Anthropic CEO Dario Amodei wrote that Anthropic aims to reach the point where "interpretability can reliably detect most model problems" by 2027, and described the long-run aim as a true "MRI for AI". Three months earlier, a January 2025 survey, Open Problems in Mechanistic Interpretability, led by Lee Sharkey with 28 co-authors, grouped what still stands in the way into three areas: methods that need conceptual and practical improvement, open questions about how to apply them to specific goals, and socio-technical issues around the field.

Key Terms in Mechanistic Interpretability

Mechanistic interpretability has its own vocabulary, and most of the words describe one idea: a model stores more concepts than it has neurons, so you need a tool to pull them apart.

Term Plain meaning Where it comes from
Feature A direction inside the model that stands for one concept, such as "Golden Gate Bridge" or "code error" Core unit of the field
Superposition The model packs more features than it has dimensions, so features share neurons Toy Models of Superposition, Elhage et al., 2022
Polysemantic vs monosemantic A polysemantic neuron fires for many unrelated concepts. A monosemantic unit fires for one Superposition explains polysemanticity
Sparse autoencoder (SAE) A second network trained to split activations into many sparse, readable features Scaling Monosemanticity, Templeton et al., 2024
Induction head An attention circuit that completes [A][B] … [A] with [B] Olsson et al., 2022
Circuit A group of features and connections that together carry out one computation The unit that circuit tracing maps
Transcoder A readable stand-in for a model's MLP neurons, built from sparse features, used to trace flow between layers Circuit Tracing, Anthropic, 2025
Attribution graph A map of which features pushed a specific output, with low-influence nodes and edges pruned away Circuit Tracing, Anthropic, 2025
Replacement model The readable copy of the model that a graph actually describes. Error nodes mark the part it cannot explain Circuit Tracing, Anthropic, 2025

How Mechanistic Interpretability Works

Mechanistic interpretability works in a fixed order: record what the model does inside, decompose it into readable features, connect those features into a graph for one prompt, then intervene to prove the graph is causal.

  1. Record the activations. Researchers run text through a large language model and save the internal numbers at each layer of the transformer.
  2. Decompose them into features. Raw neurons are polysemantic, so a sparse autoencoder learns a larger set of features where only a few are active at once. Anthropic's Scaling Monosemanticity trained SAEs at roughly 1 million, 4 million, and 34 million features on the middle-layer residual stream of Claude 3 Sonnet.
  3. Label the features. Each feature gets a name from the text that turns it on. The same paper found safety-relevant features tied to security vulnerabilities, bias, lying, deception, and power-seeking, sycophancy, and dangerous content. Its authors warn that a feature for lies is not proof of lying: knowing about lies, being able to lie, and actually lying are different things.
  4. Build the graph. For circuit tracing on Claude 3.5 Haiku, Anthropic replaced the model's MLP neurons with a cross-layer transcoder of 30 million features, then drew the path from input to output for single prompts.
  5. Intervene to prove cause. Researchers clamp a feature up or down. When Anthropic amplified the Golden Gate Bridge feature, the model described itself as the bridge. When it turned up the sycophantic praise feature, the model answered an overconfident user with flattery instead of a truthful correction.
  6. Report the limits. Anthropic states that its method "only captures a fraction of the total computation" and that it takes "a few hours of human effort" to understand the circuits for a prompt of tens of words. In the companion biology paper, the team reports that its attribution graphs gave satisfying insight for about a quarter of the prompts it tried.

Mechanistic Interpretability Milestones

Mechanistic interpretability went from small attention-only models to production models in about three years.

The induction-heads paper (Olsson et al.) showed that these heads form during a sudden "phase change" early in training, at the same point that in-context learning improves sharply. On May 29, 2025, Anthropic open-sourced its circuit-tracing library with an interactive frontend on Neuronpedia, so anyone can generate attribution graphs for open-weight models such as Gemma-2-2b and Llama-3.2-1b.

Limits and Open Problems in Mechanistic Interpretability

Mechanistic interpretability is still a partial view. Every flagship result so far comes with a limit its authors state in the same paper, and those limits decide how much weight a finding can carry.

Limit What the source reports Why it matters
Incomplete dictionaries Claude 3 Sonnet can name streets in every London borough, but only about 60% of the boroughs had a matching feature in the 34M SAE (Scaling Monosemanticity) A missing feature does not mean a missing concept
Graphs describe a copy Attribution graphs describe the local replacement model, "which may differ from the underlying model" (Circuit Tracing) A graph is a hypothesis about the real model, not a readout
Low hit rate Satisfying insight for about a quarter of prompts tried (biology paper) Most prompts still resist a clean explanation
Human cost A few hours of expert work for a prompt of tens of words (Anthropic) Checking every production answer this way is out of reach today
Feature is not behavior Knowing about lies, being able to lie, and lying are different (Scaling Monosemanticity) A "deception feature" alone is not evidence of deception
What each method can tell you about one answer

  Output test        ->  "The answer is wrong."
  Written reasoning  ->  "Here is why I said it."
                         (a claim, can be unfaithful)
  Attribution graph  ->  "These features pushed it."
                         (describes a replacement model)
  Intervention       ->  "Turn feature X down and the answer changes."
                         (causal, for this one case)

Mechanistic Interpretability vs Chain-of-Thought Monitoring

Chain-of-thought monitoring reads the reasoning a model writes out. Mechanistic interpretability reads the computation underneath it. Both are safety tools. Models can write reasoning that does not match their internal path, so treat written reasoning as a claim and an interpretability result as a tested hypothesis about one case.

Question Chain-of-thought monitoring Mechanistic interpretability
What you read The reasoning text the model writes Internal features and circuits
Cost per check Low, a reviewer or second model reads text High, hours of expert work per prompt today
Scale Every run, in production Selected prompts and research cases
Main blind spot The written reasoning can be unfaithful Captures only part of the computation
Real use OpenAI used it to catch one of its reasoning models cheating on coding tests, per MIT Technology Review Found safety-relevant features, such as sycophantic praise, that steer output when amplified

Connection to Taskade

Mechanistic interpretability shows how hard it is to read knowledge once it lives inside model weights, even for the lab that trained the model. Taskade keeps what your agents know outside the weights, where you can read it. You train a Taskade AI Agent on knowledge you choose: files, links, projects, and media. Taskade EVE, the agent that builds Taskade Genesis apps, saves its notes as real Taskade projects in a projects/memory folder that you can open, edit, share, or delete. Agents run on frontier models from top AI labs, with Auto as the default. Taskade does not make those models transparent inside. It makes the context they act on inspectable, which is the part your team controls. This is Workspace DNA in practice: Memory holds facts you can audit, Intelligence reasons over them, and Execution writes results back into projects. When an answer looks wrong, you open the project it came from.

The Taskade workspace memory knowledge graph links projects and notes

What You Would Build in Taskade

You already audit people this way. When a new analyst makes a strange call, you ask which document they read and which rule they followed, and then you fix the document.

In Taskade you would describe an agent knowledge audit board. Each row is one answer an agent gave. The row stores the question, the answer, and the projects and files the agent had as knowledge. A reviewer agent, briefed only with your checklist, marks whether each claim traces back to a source in the workspace. Rows it cannot trace get assigned to a teammate for a human in the loop review, and the fix goes into the knowledge project, not into a hidden setting. Over time the board shows which sources your agents lean on and which ones need updates.

Describe yours and build it free →

Frequently Asked Questions About Mechanistic Interpretability

What is mechanistic interpretability in simple terms?

Mechanistic interpretability is the effort to read the inside of an AI model. Researchers find the internal features that stand for concepts, connect them into circuits, and test whether those circuits cause the model's answer. The aim is to explain how a model reaches an output, not only what the output is.

How is mechanistic interpretability different from explainable AI?

Many explainable AI methods score which parts of the input mattered for an output. Mechanistic interpretability goes further and tries to identify the actual algorithm inside the network, step by step. It then checks each step by intervening on the model, so the explanation is a tested mechanism and not only a correlation.

What is a sparse autoencoder in AI interpretability?

A sparse autoencoder is a small network trained to rewrite a model's internal activations as a large set of features, where only a few are active at a time. That sparsity makes each feature easier to name. Anthropic's 2024 work trained sparse autoencoders with up to 34 million features on Claude 3 Sonnet.

What are polysemantic neurons?

A polysemantic neuron fires for several unrelated concepts. The 2022 paper Toy Models of Superposition explained why: a model stores more features than it has neurons, so features share neurons. Sparse autoencoders and transcoders try to recover monosemantic features that each stand for a single concept.

What did Anthropic find with circuit tracing?

In March 2025 Anthropic traced circuits in Claude 3.5 Haiku. It found a shared concept space across English, French, and Chinese, planning of rhyming words before a line is written, parallel paths for mental math, and cases where the model worked backward from a hint so its written reasoning did not match the internal computation. The team reports satisfying insight for about a quarter of the prompts it tried.

What are the limits of mechanistic interpretability today?

Anthropic states that its circuit-tracing method captures only a fraction of the model's computation and needs a few hours of human effort per prompt of tens of words. Its graphs describe a replacement model that can differ from the real one, so a graph is a hypothesis to test. Even a 34-million-feature dictionary missed about 40% of the London boroughs the model knows.

How does Taskade keep AI agent knowledge inspectable?

Taskade keeps agent knowledge in your workspace instead of in model weights. You give an agent files, links, and projects as knowledge, and Taskade EVE saves its memory as readable projects that you can open and edit. When an answer looks wrong, you can trace it to a source and fix that source directly.