Definition: Predictive coding is a theory of how the brain processes information: higher brain areas continuously predict what the lower areas are about to receive, and the lower areas send back only the part the prediction got wrong, the prediction error. The free energy principle is the wider framework built around it, which says a brain, or any system that keeps itself intact, works to minimize surprise about its own inputs.
TL;DR: In predictive coding, predictions flow down and errors flow up. A language model is trained on a narrower version of the same idea: predict the next token. A 2023 fMRI study of 304 people found brain responses to stories were best explained by forecasts up to about eight words ahead, organized in a hierarchy. The theory is influential and still contested. Build one free →
Think of reading your own handwriting. You barely look at most words, because you already know roughly what you meant to write. Your eyes slow down only at the word that does not fit. That pause is the prediction error. Most of the page never needed your attention, because your expectation already covered it. Predictive coding proposes that the whole cortex works this way, all the time.
Why Predictive Coding Matters in 2026
Predictive coding matters because it is the neuroscience theory closest to how modern AI is trained, and the gap between the two is now measurable. In Evidence of a predictive coding hierarchy in the human brain listening to speech (Nature Human Behaviour, 2023), Charlotte Caucheteux, Alexandre Gramfort, and Jean-Rémi King analyzed fMRI from 304 people who listened to short stories from the Narratives dataset. Activations from GPT-2 mapped linearly onto the brain responses. Adding forecasts of words further in the future improved that mapping by 23% on average, with the best window about eight words ahead, roughly 3.15 seconds of speech. Frontoparietal areas forecast further ahead and more abstractly than temporal areas.
That is the sharpest published contrast between brains and language models. A model trained on next-token prediction learns to guess the adjacent word. The brain data looked like a stack of predictions running at several timescales at once, with semantic forecasts reaching further than syntactic ones.
The evidence is also moving in the other direction. A February 2025 review in Trends in Cognitive Sciences, Predictive coding: a more cognitive process than we thought? by Kaitlyn Gabhart, Yihan Xiong, and André Bastos, argues that recordings from single neurons show genuine prediction errors emerging in prefrontal cortex rather than in early sensory areas. In March 2026, Desegregation of neuronal predictive processing by Bin Wang, Nicholas Audette, David Schneider, and Johnatan Aljadeff (Nature Communications) found that recurrent networks doing robust prediction mix stimulus signals and prediction-error signals across the same neurons, contrary to earlier proposals of specialized "error cells". The team then confirmed the prediction in animals by violating their learned expectations. The core idea survives. The textbook circuit is being redrawn.
The language evidence is also under pressure. In April 2026, an eLife paper by Schönmann, Szewczyk, de Lange, and Heilbron, Stimulus dependencies, rather than next-word prediction, can explain pre-onset brain encoding, ran the standard "the brain encodes the next word before it arrives" analysis on systems that cannot predict at all: word embeddings and raw speech acoustics. Both control systems produced the same hallmarks. Natural language is full of correlations between neighboring words, so a signal that looks like a forecast can be an echo of the words already heard. In an accompanying eLife commentary, Richard Antonello summarized the result: apparent encoding of future words may come from the structure of language itself.
A Short History of Predictive Coding
The idea is older than neuroscience's tools for testing it. Each row below links to a source that was checked for this article.
| Year | Milestone | What it added |
|---|---|---|
| 1860s | Hermann von Helmholtz, unconscious inference | Perception as a best guess about hidden causes, not a passive recording |
| 1999 | Rao and Ballard, Nature Neuroscience | The first influential circuit model: predictions down, residual errors up |
| 2006 | Friston, Kilner, and Harrison, "A free energy principle for the brain" | One quantity that both perception and action can reduce |
| 2010 | Friston, Nature Reviews Neuroscience | The claim that several brain theories optimize the same thing |
| 2016 | PredNet, Lotter, Kreiman, and Cox | A deep network that forwards only its prediction errors, trained on video |
| 2017 | Whittington and Bogacz, Neural Computation | Predictive coding with local Hebbian updates can approximate backpropagation |
| 2023 | Caucheteux, Gramfort, and King, Nature Human Behaviour | A multi-timescale forecast hierarchy in 304 human brains |
| 2024 | Antonello and Huth, Neurobiology of Language | The model-brain fit may come from features, not prediction |
| 2025 | Gabhart, Xiong, and Bastos, Trends in Cognitive Sciences | Genuine prediction errors in spiking appear mainly in prefrontal cortex |
| 2026 | Wang et al. and Schönmann et al. | Error and stimulus signals share neurons, and a confound in the "pre-onset" language evidence |
Key Terms
| Term | Plain meaning | Where it shows up |
|---|---|---|
| Generative model | An internal model of how hidden causes produce the input | Every level of the predictive hierarchy |
| Prediction | The model's guess of the activity one level down | Feedback connections |
| Prediction error | Input minus prediction, the part the guess missed | Feedforward connections |
| Precision | How reliable a signal is, used to weight its error | A noisy input gets less say than a clean one |
| Variational free energy | Surprise plus the gap between belief and best explanation | The free energy principle |
| Active inference | Reducing error by acting on the world, not only by revising beliefs | Movement and behavior under the free energy principle |
How Predictive Coding Works
The classic formal model is Rao and Ballard's 1999 paper in Nature Neuroscience, "Predictive coding in the visual cortex". In their network, feedback connections from a higher visual area carry predictions of the activity in the area below, and feedforward connections carry the residual errors between those predictions and the actual activity.
- A higher area makes a guess. It holds a model of the current situation and sends down a prediction of what the lower area will see next.
- The input arrives. The lower area receives the real signal from the senses or from the area below it.
- Only the difference moves up. Whatever the prediction already explained is dropped. The residual, the prediction error, travels up the hierarchy. This is efficient, because a well-predicted world costs very little signal.
- The higher area updates. It adjusts its belief to reduce the error, and sends a better prediction down.
- The loop repeats at every level. Each level predicts the one below, so the hierarchy explains the input at several scales, from edges to objects, or from sounds to meaning.
When Rao and Ballard trained this network on natural images, its units developed receptive fields like simple cells in visual cortex, and the error units reproduced "endstopping", where a neuron responds less to a line that extends past its receptive field. Their reading: the surround is predictable from context, so there is less error to report.
The free energy principle puts this inside a single objective. Karl Friston, James Kilner, and Lee Harrison introduced it in 2006 in A free energy principle for the brain, which recast Helmholtz's ideas about perception in modern terms. Friston's 2010 review The free-energy principle: a unified brain theory? in Nature Reviews Neuroscience argued that many brain theories optimize the same quantity: value (expected reward) or its complement, surprise (prediction error). Surprise cannot be computed directly, so the brain minimizes variational free energy, which is surprise plus a KL divergence between its approximate belief and the true posterior. Because a KL divergence is never negative, free energy is an upper bound on surprise. Negative free energy is the same quantity machine learning calls the evidence lower bound (ELBO), the objective used to train variational autoencoders.
Variational free energy, in one line
F = −log p(input) + KL[ q(causes) ‖ p(causes | input) ]
surprise gap between the brain's belief q
(the thing we want) and the best explanation (never < 0)
so F ≥ surprise and −F = ELBO (the VAE training objective)
Two ways to push F down:
perception → change q, the belief, so it explains the input better
action → change the input, so it matches what q expected
Surprise here is the information-theory kind, measured in bits (log base 2) or nats (natural log), a direct relative of entropy and cross-entropy. The principle adds action: you can reduce surprise by changing your beliefs, or by acting on the world so the input matches the prediction. Friston calls that active inference.
Predictive Coding vs Next-Token Prediction vs the Free Energy Principle
These three ideas share a vocabulary and get used interchangeably, which blurs real differences in what each one claims.
| Question | Predictive coding | Next-token prediction | Free energy principle |
|---|---|---|---|
| What it is | A theory of brain circuits | A training objective for language models | A general principle for self-organizing systems |
| Key paper | Rao and Ballard, 1999 | Standard for GPT-style models | Friston, 2006 and 2010 |
| What gets predicted | Activity in the level below, at every level | The next token only | Sensory input, through a generative model |
| What travels upward | Prediction error | Nothing at inference time, gradients during training | Prediction error, weighted by confidence |
| Prediction horizon | Several timescales (up to about 8 words in the 2023 fMRI study) | One step ahead | Not fixed, it depends on the model |
| Role of action | Not part of the core theory | None, it only predicts text | Central: act to make input match prediction |
| Quantity minimized | Prediction error | Cross-entropy loss | Variational free energy, an upper bound on surprise |
| Can it be falsified? | Yes, circuit claims are testable and some are contested | Yes, it is an engineering choice | Friston says that, as a principle, it cannot be falsified |
The horizon row is the one the 2023 study measured. The chart shows the forecast distances that best explained brain responses, next to the one-step horizon of the training objective.
The brain values come from Caucheteux et al. (2023): syntactic forecasts peaked at about five words, in superior temporal and left frontal areas, and semantic forecasts at about eight words, in frontal and parietal areas. A trained model can still carry information about later words inside its hidden states. The difference is what the objective asks for.
The honest caveat. A good fit between a model and brain data does not prove that the brain predicts. In Predictive Coding or Just Feature Discovery? (Neurobiology of Language, 2024), Richard Antonello and Alexander Huth showed that, inside a language model, the representations that are best at predicting future words are strictly worse models of the brain than other representations in the same model. They argue that language models fit brain data because they capture a wide variety of linguistic phenomena, not because they predict. The 2026 eLife control study by Schönmann and colleagues, described above, adds a second confound: word-to-word correlations in natural language can mimic a forecast. At the circuit level, the Simons Foundation summarized the older problem: a reduced response to a repeated stimulus, often cited as evidence, can also come from plain neural adaptation.
Predictive Coding in Machine Learning
Predictive coding is not only a theory about brains. It is also a family of learning algorithms, and that line of work asks a pointed question: can a network learn well if every connection updates using only the signals at its two ends?
Backpropagation, the method that trains almost every modern neural network, cannot do that. It computes one error at the output and carries it backward through every layer, so a weight change depends on neurons far from the synapse. In 2017, James Whittington and Rafal Bogacz showed in Neural Computation that a predictive coding network can do supervised learning with only simple local Hebbian plasticity, and that for certain parameters its weight changes converge to those of backpropagation.
Two other results show the range of the idea:
- PredNet (2016). William Lotter, Gabriel Kreiman, and David Cox built a deep network for video in which each layer makes local predictions and forwards only the deviations. Trained on unlabeled video, including car-mounted camera footage, it learned representations that helped with tasks such as estimating steering angle.
- The 2022 survey. Beren Millidge, Tommaso Salvatori, Yuhang Song, Rafal Bogacz, and Thomas Lukasiewicz reviewed the field in Predictive Coding: Towards a Future of Deep Learning beyond Backpropagation?. They describe predictive coding networks that use local learning, run on arbitrary graph topologies, and act as classifiers, generators, and associative memories.
Associative memory is the thread that links predictive coding to the Hopfield network, and local learning is the thread that links it to Hebb. Frontier language models still train with backpropagation on a next-token loss. Predictive coding networks remain a research direction, not the method behind today's production models.
Connection to Taskade
The lesson that carries over to AI work is simple: a good prior saves effort, and the surprise is the signal worth recording. When an agent starts with a written plan, most of the work is already predicted, and the useful information is wherever the result differs from the plan.
Taskade is built to hand agents that prior. A Taskade AI Agent can use a project as its knowledge, so it reads your plan, spec, or checklist before it acts rather than starting from a blank prompt. On /connect/dna, the middle "association" ring is where projects and shared context live, and it is what agents read before they act. The brain terms on that page are a metaphor for how the parts connect. Taskade does not model a brain. Workspace DNA closes the loop in plain terms: Memory (projects and databases), Intelligence (agents running on frontier models from top AI labs), and Execution (automations across 100+ bidirectional integrations) that write results back into Memory.

| Predictive coding idea | Plain lesson | Where it lives in Taskade |
|---|---|---|
| A prior from the level above | Start from what you already expect | A project the agent reads as knowledge |
| Prediction error | The mismatch is the useful signal | A "surprises" table that logs where results differed from the plan |
| Update the belief | Revise the model where it was wrong | You, or an automation you wrote, edit the plan |
| Precision weighting | Trust reliable signals more | Agent instructions that say which sources come first |
The honest limit: nothing in Taskade updates its own beliefs the way a predictive-coding network does. Agents do not learn on their own. What improves is the content. When you log where reality differed from the plan and then revise the plan, the next run reads a better prior. You make that update, or an automation you wrote makes it.
What You Would Build in Taskade
You already work this way with people. A good manager writes the plan, lets the team run, and asks to hear only about what went differently. A status report that repeats the plan tells you nothing. The exceptions are the report.
In Taskade you would describe a plan-versus-actual tracker. One project holds the plan: each deliverable, its expected date, and what "done" looks like. An agent gets that project as knowledge. An automation triggers on each new update, asks the agent to compare it with the plan, and uses a branch step: when they differ, it writes a row into a "surprises" table with the gap, the likely cause, and who needs to decide. Updates that match the plan stay quiet. At the weekly review you read the surprises, change the plan where it was wrong, and every agent reads the revised plan on its next run. Taskade Genesis can build the first version from one prompt, with Taskade EVE planning the projects, agent, and automation it needs.
Describe yours and build it free →
Related Concepts
- Next-Token Prediction: the one-step training objective that predictive coding is most often compared with
- World Model: the internal model that makes predictions possible, in brains and in AI
- KL Divergence: the gap term inside variational free energy
- Entropy: the average surprise of a source, in bits
- Cross-Entropy: the loss a language model minimizes, a measure of average surprise under a model
- Hebbian Learning: the older theory of how connections between neurons strengthen
- Connectome: the wiring map that predictions and errors travel along
- Hopfield Network: an energy-minimizing associative memory, a close cousin of predictive coding networks
- Memory Consolidation: how the brain turns recent experience into long-term memory
- In-Context Learning: how a language model adapts to a prompt without changing its weights
- What Is Intelligence?: from neurons to AI agents, the longer story
Frequently Asked Questions About Predictive Coding
What is predictive coding in simple terms?
The brain guesses what it is about to sense and pays attention mainly to what it got wrong. Higher areas send predictions down, and lower areas send up only the prediction error. The theory was formalized for vision by Rao and Ballard in 1999 and is now applied to hearing, language, movement, and more.
What is the difference between predictive coding and the free energy principle?
Predictive coding is a specific theory about brain circuits, with predictions going down and errors going up. The free energy principle, from Karl Friston, is a broader claim that any self-organizing system minimizes surprise about its inputs, by updating beliefs or by acting. Predictive coding is often presented as one way a brain could carry that out.
Is the brain just a next-token predictor like ChatGPT?
No. The 2023 study by Caucheteux, Gramfort, and King found brain responses to stories were best explained by forecasts up to about eight words ahead, arranged in a hierarchy, while a language model is trained to predict one token at a time. Other researchers, including Antonello and Huth, argue that the model-brain fit comes from shared features, not shared prediction.
How does the free energy principle relate to machine learning?
Variational free energy is the negative of the evidence lower bound (ELBO), the objective used to train variational autoencoders. It equals surprise plus a KL divergence between an approximate belief and the true one. Minimizing it pushes the belief toward the best explanation of the data.
Is predictive coding proven?
Not fully. Neuroimaging shows many prediction-error-like responses, but a 2025 review by Gabhart, Xiong, and Bastos notes that single-neuron recordings place genuine prediction errors mainly in prefrontal cortex, and a 2026 Nature Communications study found stimulus and prediction-error signals mixed in the same neurons rather than held in dedicated error cells. Some effects credited to prediction can also come from ordinary neural adaptation.
Does the brain predict upcoming words before it hears them?
Some evidence says yes, and the debate is active. Brain signals that appear before a word is heard have been read as prediction, but a preprint by Azizpour and colleagues replicated those signals and found no stable overlap with the brain's response to the word itself. A 2026 eLife study by Schönmann and colleagues found the same pre-onset hallmarks in word embeddings and speech acoustics, which cannot predict anything, so correlations between neighboring words can produce them.
Can predictive coding replace backpropagation?
Not in production today. Whittington and Bogacz showed in 2017 that a predictive coding network can learn with only local Hebbian updates and, for certain parameters, match the weight changes of backpropagation. A 2022 survey by Millidge and colleagues argues these networks are more flexible. Frontier language models still train with backpropagation.
What can AI builders take from predictive coding?
Two habits. Give your agent a strong prior, such as a written plan, spec, or checklist it reads before it acts. And log the prediction error, the places where the result differed from the plan, because that is where the information is. A report that only confirms the plan tells you little.
Do Taskade agents learn the way the brain does?
No. A Taskade AI Agent reads your projects, files, and notes and keeps context across chats with agent memory, but it does not update its own model or learn on its own. What gets sharper over time is the content it reads, such as the plans and records your team keeps up to date.