AI Concepts

Continual Learning (and Catastrophic Forgetting)

17 min read
On this page (22)

Definition: Continual learning is the ability to keep learning from a stream of new tasks and data without erasing what was learned before. Its central obstacle is catastrophic forgetting, also called catastrophic interference: when a neural network is trained on something new, the weight updates overwrite the knowledge it already held. People learn continually every day. Today's large language models mostly do not.

TL;DR: Continual learning means learning new things without forgetting old ones. Researchers documented the problem in 1989, when they showed that training in sequence overwrites earlier knowledge. The brain handles it with two memory systems, a fast one and a slow one. In October 2025 Andrej Karpathy named the gap plainly: today's agents "don't have continual learning." Build one free →

Think of a chalkboard that only has room for one lesson. Every time the teacher writes a new one, the last lesson gets wiped. Your own memory does not work that way. You learned to ride a bike, then to drive, and driving did not erase the bike. A standard neural network is closer to the chalkboard: the same shared weights hold every skill, so writing a new skill in smudges the old ones out.

Why Continual Learning Matters in 2026

Continual learning is the gap between a model that was trained once and a colleague who gets better on the job. Frontier models are trained, frozen, and shipped. Anything they pick up during a conversation lives in the context window and is gone when the chat ends. In his October 2025 interview with Dwarkesh Patel, Andrej Karpathy explained why agents cannot yet do the work of an intern: "They don't have continual learning. You can't just tell them something and they'll remember it." He compared the missing piece to sleep, when a person's day is distilled back into the weights of the brain, and said large language models have no equivalent phase.

The research record explains why the problem is so stubborn. In 1989 Michael McCloskey and Neal Cohen published Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem, which showed that teaching a network new items in sequence could badly disrupt what it had learned before. In 2024 Richard Sutton's group reported a second failure in Nature: trained on task after task, standard deep-learning methods gradually lose plasticity, the capacity to learn at all, "until they learn no better than a shallow network." The Bitter Lesson explainer covers that result in depth.

The field is moving again. In November 2025 Google Research introduced Nested Learning, published at NeurIPS 2025 and framed explicitly as a paradigm for continual learning. The authors compare the limits of current LLMs to anterograde amnesia, where a person is limited to immediate context and cannot form new long-term memories. In September 2025, researchers at MIT's Improbable AI Lab showed in RL's Razor that fine-tuning with on-policy reinforcement learning preserves prior knowledge significantly better than supervised fine-tuning, and that the amount of forgetting tracks how far the new model drifts from the base model.

How Continual Learning Works

The brain's answer came first. In 1995, James McClelland, Bruce McNaughton and Randall O'Reilly proposed Complementary Learning Systems (CLS) theory in Psychological Review. A later review summarizes it this way: the hippocampus is a sparse, pattern-separated system that learns episodes fast, and the neocortex is a distributed, overlapping system that learns slowly and extracts structure across many episodes. Consolidation connects the two.

  1. Shared weights cause the forgetting. A network stores every skill across the same overlapping weights. Training on task B moves those weights toward B and away from the settings that solved task A. That is catastrophic interference.
  2. A fast store captures the new thing first. In CLS theory the hippocampus records a new episode quickly, in a sparse code that does not overwrite older episodes.
  3. Replay interleaves old and new. The stored episodes are reinstated in the neocortex alongside older memories. The 1995 paper argues that the cortex discovers structure when learning of each item is gradual and interleaved with learning about other items. DeepMind's 2016 update of CLS theory points out that the experience replay buffer in the DQN Atari agent plays a hippocampus-like role, interleaving stored experiences with new play.
  4. The slow store changes a little each time. Each reinstatement nudges cortical synapses slightly, so the new skill is folded in without wiping the old ones.
  5. Important connections get protected. Elastic Weight Consolidation (Kirkpatrick et al., DeepMind, PNAS 2017) copies the idea of synaptic consolidation. It remembers old tasks "by selectively slowing down learning on the weights important for those tasks."
  6. Memory can run at several speeds. Google's Continuum Memory System, part of Nested Learning, treats memory as a spectrum of modules that each update at their own frequency, instead of one fast context and one frozen set of weights.

Here is the failure in its simplest form:

STEP 1  Train on Task A   ->  model is good at A
STEP 2  Train on Task B   ->  model is good at B
STEP 3  Test Task A again ->  much of A is gone
        ^ nothing told the network to protect A.
          The same shared weights were rewritten for B.

The Stability-Plasticity Dilemma

Every continual learner faces the stability-plasticity dilemma. A learning system needs plasticity to take in new knowledge and stability to keep what it already knows, and the two pull against each other. In a 2013 article in Frontiers in Psychology, Martial Mermillod, Aurélia Bugaiska and Patrick Bonin argued that the two failures are opposite ends of one continuum, not two separate problems.

   TOO PLASTIC                                         TOO STABLE
   <------------------------------------------------------------>
   learns B fast,            the useful zone:          keeps A safe,
   forgets A                 learns B, keeps A         cannot learn B
   (catastrophic                                       (entrenchment;
    forgetting)                                         loss of plasticity
                                                        is a related failure)
Too much of What goes wrong Where it shows up
Plasticity New learning overwrites old knowledge Catastrophic forgetting, first described in 1989
Stability Old knowledge blocks new learning Entrenchment, the effect Mermillod and colleagues link to age-limited learning
Long training on a task stream The network loses the capacity to learn at all Loss of plasticity, reported in Nature in 2024

Three Scenarios: Task, Domain, and Class Incremental

In 2019 Gido van de Ven and Andreas Tolias described three scenarios for continual learning, sorted by what the model is told at test time. The split matters because a method that looks strong in one scenario can fail in another.

Scenario Is the task named at test time? What the model must do Difficulty
Task-incremental Yes Solve the named task Easiest
Domain-incremental No, and the model does not need to infer it Solve the same kind of problem under new conditions Middle
Class-incremental No, and the model must infer it Choose among every class it has ever seen Hardest

On split and permuted MNIST benchmarks, the authors found that regularization-based methods such as EWC fail in the class-incremental scenario, and that replaying representations of earlier experiences seemed to be required. That result is one reason replay, the brain's own strategy, keeps coming back.

The Main Families of Continual Learning Methods

The research groups continual learning methods into three families: replay, regularization, and parameter isolation. That is the taxonomy of the widely cited survey A Continual Learning Survey: Defying Forgetting in Classification Tasks (De Lange et al., IEEE TPAMI). Each family answers the same question differently: how do you keep task A safe while the network learns task B?

  • Replay stores a sample of old data and mixes it into new training. Deep Generative Replay (Shin et al., 2017) skips the storage. A generator, "inspired by the generative nature of hippocampus," produces pseudo-examples of old tasks that are interleaved with the new task.
  • Regularization adds a penalty that keeps important weights close to their old values. EWC is the best-known example.
  • Parameter isolation gives each task its own part of the network. Progressive Neural Networks (Rusu et al., DeepMind, 2016) freeze each trained column and add a new one with lateral connections to the old features, so earlier tasks are immune to forgetting. The price is a network that grows with every task.

Continual Learning for Large Language Models

Large language models meet catastrophic forgetting every time they are fine-tuned. The practical question is how much general skill a model gives up when it learns a new domain.

  • Low-rank adapters trade learning for retention. The 2024 study LoRA Learns Less and Forgets Less (Biderman et al., TMLR) compared LoRA with full fine-tuning on programming and mathematics. LoRA learned the target domain less well, but it better maintained the base model's performance outside that domain.
  • On-policy reinforcement learning forgets less. RL's Razor reports that, among the many ways to solve a new task, on-policy RL prefers solutions closest to the original model, which keeps prior skills intact.
  • New architectures aim at the root cause. Nested Learning and its proof-of-concept model, Hope, try to build memory at several update speeds into the architecture itself.
  • External memory skips weight updates altogether. Retrieval-augmented generation and persistent memory keep new knowledge in documents the model reads at answer time. Nothing is overwritten, because nothing in the weights changes.

Continual Learning in the Brain vs in AI

Most machine-learning fixes for forgetting map onto an idea the brain already uses, some closely and some only loosely. The table pairs them up. The last row is the approach you can use today without retraining anything.

Brain mechanism Machine-learning technique What it protects Trade-off
Hippocampal replay of recent episodes Experience replay, rehearsal buffers, and generative replay Old tasks, by mixing their examples into new training Old data must be stored, or generated again
Synaptic consolidation (less plastic synapses) Elastic Weight Consolidation The specific weights an old task depends on Protected weights leave less room for new skills
Fast and slow learning systems (CLS) Nested Learning and the Continuum Memory System Knowledge held at different update speeds New research architecture, not in today's mainstream models
Loose analogy: keeping some capacity fresh Continual backpropagation (reinitializing a small share of less-used units) The ability to keep learning at all Adds a random, non-gradient step to training
Small, gradual cortical change On-policy RL fine-tuning (RL's Razor) Prior skills, by keeping the new model close to the base model Needs a reward signal and samples drawn from the model itself
Notes, diaries, and shared records External memory: retrieval and persistent memory Everything, because the weights never change The model knows only what it can find and read

The neighbor term is fine-tuning. Fine-tuning is one training pass on new data, and it is where catastrophic forgetting shows up in practice. Continual learning is the goal of running that pass again and again, forever, without the losses adding up. The other neighbor is in-context learning, where a model picks up a pattern from the prompt alone and keeps nothing once the chat ends.

Connection to Taskade

Frontier models do not learn continually in their weights, and Taskade does not pretend they do. Every Taskade AI Agent runs on one of the frontier models from top AI labs, and those weights stay fixed. Taskade keeps knowledge outside the weights instead, in the Memory pillar of Workspace DNA: Memory + Intelligence + Execution. Your projects and databases hold the facts your team works from. You train each agent on files, links, projects, and media as agent knowledge, and persistent memory keeps context across chats.

The workspace memory knowledge graph links the projects and notes that Taskade agents read

Taskade EVE, the agent that builds Taskade Genesis apps, saves notes as real, readable projects in a projects/memory folder and keeps a running TASKS.md list of what is done and what remains (how Taskade EVE memory works). Automations write their results back into Memory, so what your agents know grows with every run. This is the external-memory row of the table above: parametric knowledge stays frozen, contextual knowledge keeps growing, and you can open, edit, or delete any of it.

Nothing in Taskade learns on its own. As the connectome page puts it, the workspace gets sharper the more you use it, because it holds more of your work.

What You Would Build in Taskade

You already do this by hand. A new teammate joins, and someone forwards them a pile of old threads and says "read these first." The knowledge survives, but only because a person carried it across.

In Taskade you would describe a team playbook that keeps itself current. Each closed project, support case, or post-mortem lands as a row with what happened and what worked. An automation adds each new row when the work closes, and an agent trained on the playbook project answers questions from it: how did we handle the last refund dispute, which vendor missed a deadline, what is the checklist for a launch. Last quarter's answers stay in the project because nothing overwrites them, and this week's lessons sit right beside them. The agent does not retrain. It reads a record that grows. The agent knowledge guide shows how to connect a project as knowledge.

Adding files as knowledge to train a Taskade AI agent

Describe yours and build it free →

Frequently Asked Questions About Continual Learning

What is continual learning in AI?

Continual learning is a model's ability to keep learning from new tasks and data over time without losing what it learned before. It is also called lifelong or incremental learning. The main obstacle is catastrophic forgetting, where training on new material overwrites the weights that stored older skills.

What is catastrophic forgetting?

Catastrophic forgetting, also called catastrophic interference, happens when a neural network learns a new task and its performance on an old task collapses. McCloskey and Cohen described it in 1989. It happens because every skill is stored across the same shared weights, so training for one moves the settings another depended on.

What is the stability-plasticity dilemma?

The stability-plasticity dilemma is the trade-off every learning system faces. Plasticity lets it take in new knowledge, and stability lets it keep old knowledge. Too much plasticity causes catastrophic forgetting. Too much stability blocks new learning. Continual learning methods try to hold a useful balance between the two.

How does the human brain avoid catastrophic forgetting?

Complementary Learning Systems theory says the brain uses two systems. The hippocampus learns new episodes quickly in a sparse code. The neocortex learns slowly and gradually. Replay reinstates recent memories in the neocortex, interleaved with older ones, so new knowledge is folded in without wiping out the old.

What are the main approaches to continual learning?

A widely cited survey groups them into three families. Replay methods mix old or generated examples back into training. Regularization methods, such as Elastic Weight Consolidation, protect the weights old tasks depend on. Parameter isolation methods, such as progressive neural networks, give each task its own part of the network.

What is Elastic Weight Consolidation?

Elastic Weight Consolidation (EWC) is a 2017 DeepMind method for learning tasks in sequence. It estimates which weights matter most for an old task and slows learning on those weights when a new task is trained. The authors modeled it on synaptic consolidation, and tested it on MNIST classification tasks and a sequence of Atari games.

Does LoRA prevent catastrophic forgetting?

LoRA reduces forgetting but does not remove it. A 2024 study titled "LoRA Learns Less and Forgets Less" found that LoRA better maintained a base model's performance outside the target domain than full fine-tuning did, but it also learned the new domain less well.

Do ChatGPT and Claude learn continually?

No. Frontier models are trained, then frozen before release. What they seem to learn inside a conversation lives in the context window and disappears when the chat ends. Memory features store notes outside the model and load them back into the prompt. Andrej Karpathy named the lack of continual learning as a core gap in today's agents.

What is loss of plasticity?

Loss of plasticity is the gradual decline in a network's ability to learn anything new after long training on a stream of tasks. A 2024 Nature paper from Richard Sutton's group showed standard methods decline until they learn no better than a shallow network. Their fix, continual backpropagation, reinitializes a small share of less-used units. See the Bitter Lesson explainer.

What is Nested Learning?

Nested Learning is a Google Research approach published at NeurIPS 2025. It treats a model as a set of smaller, nested optimization problems, each with its own internal flow of information. Its Continuum Memory System spreads memory across modules with different update frequencies, and its proof-of-concept architecture, Hope, is a self-modifying model built as a variant of the Titans architecture.

Do Taskade AI Agents learn continually?

Not in their weights, and Taskade does not claim they do. The models stay fixed. What grows is the knowledge the agents read: your projects, the files and links you train them on, persistent memory, and the notes Taskade EVE saves as readable projects. Agents know more because your workspace holds more, and you stay in control of every record.