Definition: In-context learning (ICL) is the ability of a large language model to pick up a new task from the text in its prompt, an instruction, a few worked examples, or both, and then carry that task out on a fresh input, with no change to the model's weights. The term was popularized by the 2020 GPT-3 paper, which applied the model to every task "without any gradient updates or fine-tuning," with the task specified purely through text.
TL;DR: In-context learning is a model learning a task from its prompt instead of from retraining. Show it three examples of the format you want and it returns the fourth in the same shape. Nothing about the model changes, so the skill lasts exactly as long as the examples stay in the context window. That is why what you put in the context matters as much as which model you pick. Build one free →
You already do this with people. You hand a new colleague two finished invoices and say "do the rest like these." You do not send them on a course or rewire how they think. They read the examples, infer the pattern, and apply it to the pile on their desk. Take the examples away tomorrow and they fall back on their general habits. A language model works the same way, except that "tomorrow" is the next conversation.
Why In-Context Learning Matters in 2026
In-context learning is the reason you can steer a general model toward your task by writing, not by training. When OpenAI introduced GPT-3 in Language Models are Few-Shot Learners (Brown et al., 2020), the 175-billion-parameter model was never fine-tuned for the benchmarks it was tested on. The authors used "in-context learning" for the inner loop of the process, the adaptation that happens inside a single forward pass, and found that larger models made increasingly efficient use of the information in their context.
Longer context windows then turned a handful of examples into hundreds. In Many-Shot In-Context Learning (NeurIPS 2024 spotlight), a Google DeepMind team reported significant gains going from few-shot to many-shot prompts with hundreds or thousands of examples, across a wide variety of generative and discriminative tasks. With enough examples, the model could override biases it picked up in pretraining, which few-shot prompts could not do, and it performed comparably to fine-tuning. The same paper notes the price: inference cost grows linearly with the number of examples.
The mechanism is still an active research question, and the 2025 findings changed the picture. At ICML 2025, Kayo Yin and Jacob Steinhardt compared two kinds of attention head across 12 language models in Which Attention Heads Matter for In-Context Learning?. Few-shot performance depended primarily on function vector heads, especially in larger models, rather than on the induction heads that earlier work pointed to (more on both below). A Google Research team argued in Learning without training (2025) that, inside a transformer block, a forward pass with context is mathematically equivalent to a forward pass without it, but with a minimal low-rank update to the MLP weights. In plain terms, the prompt acts like a temporary, invisible edit that disappears when the prompt does.
| Year | Milestone | What it added |
|---|---|---|
| 2020 | GPT-3 paper (Brown et al., OpenAI) | Used the term "in-context learning" for the inner loop and defined zero-, one-, and few-shot |
| 2021 | Implicit Bayesian inference (Xie et al.) | A theory: the model infers a hidden "concept" shared by the examples |
| 2022 | Rethinking the Role of Demonstrations (Min et al.) | Random labels in the examples barely hurt, so format carries much of the signal |
| 2022 | In-context Learning and Induction Heads (Olsson et al., Anthropic) | Tied a specific attention circuit to the moment ICL appears in training |
| 2022 | Transformers learn in-context by gradient descent (von Oswald et al.) | Showed attention layers can carry out gradient descent steps inside one forward pass |
| 2023 | Larger models do ICL differently (Wei et al.) | Only large models follow flipped labels over what they already believe |
| 2024 | Many-shot ICL (Agarwal et al., Google DeepMind) | Hundreds or thousands of examples in long windows |
| 2024 | Many-shot jailbreaking (Anthropic) | The same mechanism, turned against a model's safety training |
| 2025 | Function vector heads (Yin and Steinhardt) | Few-shot ICL depends primarily on function vector heads |
| 2025 | Learning without training (Dherin et al., Google Research) | Context acts like a temporary low-rank weight update |
How In-Context Learning Works
In-context learning happens entirely at inference time. The model's weights are frozen, and everything it adapts to arrives as tokens in the prompt.
- Pretraining builds the raw material. During next-token prediction over huge text collections, the model develops broad skills and pattern recognition. GPT-3's authors framed this as the outer loop of meta-learning.
- The prompt specifies the task. An instruction, a few examples, or both go into the context window. No label says "this is training data." It is all just text.
- Attention connects the new input to the examples. Through the attention mechanism, tokens in the new input look back at similar tokens earlier in the prompt and at what followed them.
- Induction heads copy the pattern forward. Induction heads are attention heads that complete a sequence like
[A][B] ... [A] → [B]. Anthropic researchers first described them in their 2021 transformer circuits work, and in In-context Learning and Induction Heads (Olsson et al., 2022) they tied them to ICL. The heads form at the same point as a sudden jump in in-context learning ability, visible as a bump in the training loss. For small attention-only models the authors gave strong causal evidence. For larger models with MLP layers the evidence was correlational, and the authors said so. - Richer heads carry the task itself. The 2025 follow-up by Yin and Steinhardt found that few-shot performance depends primarily on function vector heads, whose activations compute a hidden encoding of the task being demonstrated, especially in larger models. Many of those heads start out as induction heads during training before they switch to the function vector mechanism.
- The adaptation ends with the prompt. Nothing is saved. Start a new conversation without the examples and the model is back to its pretrained defaults.
The three classic settings from the GPT-3 paper differ only in how many demonstrations the prompt carries. The few-shot learning page covers how to write them.
| Setting | What the prompt holds | Where to read more |
|---|---|---|
| Zero-shot | An instruction only | Zero-shot learning |
| One-shot | An instruction plus one example | Few-shot learning |
| Few-shot | As many examples as fit in context (typically 10 to 100 for GPT-3's 2,048-token window) | Few-shot learning |
| Many-shot | Hundreds or thousands of examples in a long context window | Many-Shot ICL paper |
One more finding keeps expectations honest. In Rethinking the Role of Demonstrations (Min et al., EMNLP 2022), randomly replacing the labels in the examples barely hurt performance on a range of classification and multiple-choice tasks, consistently across 12 models, GPT-3 included. The examples often teach the format, the label space, and the kind of input to expect more than they teach the right answer.
Scale changes that picture. In Larger language models do in-context learning differently (Wei et al., 2023), researchers flipped the labels on purpose. Small models ignored the flipped examples and kept answering from what they learned in pretraining. Large models followed the flipped examples, and the authors called overriding those semantic priors an emergent ability of model scale. So the same prompt can teach a small model the format and a large model the rule.
Theories of Why It Works
No single theory explains in-context learning yet. Five views have strong papers behind them, and they are not mutually exclusive: several describe the same behavior at different levels.
| View | Core idea | Key paper |
|---|---|---|
| Pattern copying | Induction heads find an earlier match and copy what came next | Olsson et al., 2022 |
| Task vectors | Function vector heads compress the demonstrated task into a hidden encoding | Yin and Steinhardt, 2025 |
| Bayesian inference | Pretraining on coherent documents teaches the model to infer the hidden concept shared by the examples | Xie et al., ICLR 2022 |
| Gradient descent in the forward pass | Trained self-attention layers can act like steps of gradient descent on the examples | von Oswald et al., 2022 |
| Implicit weight update | Context works like a temporary low-rank edit to the MLP weights | Dherin et al., 2025 |
The Bayesian view also predicts behavior you can see in practice: Xie and colleagues reproduced better performance with scale, sensitivity to the order of the examples, and cases where zero-shot beats few-shot on their synthetic dataset. The gradient-descent view came from small self-attention-only models trained on regression tasks, so treat it as a mechanism that can exist, not a proof of what every large model does. The mechanistic interpretability page covers the tools researchers use to test claims like these.
The top half of that chart happens once, during training. The bottom half happens on every request and is gone when the request ends.
Limits and Risks
In-context learning is powerful because the model follows what it reads. The same property is its main weakness.
| Limit | What happens | What helps |
|---|---|---|
| Token cost | Every example is paid for on every call. Many-shot inference cost grows linearly with the example count | Keep a small, well-chosen set, and reuse a stable prefix with prompt caching |
| Context rot | Long contexts bury the parts that matter | Fewer, better tokens. See context rot |
| Order sensitivity | The same examples in a different order can change the answer | Keep one consistent template and test more than one ordering |
| Hostile examples | Text in the context steers the model, whoever wrote it | Treat retrieved and pasted text as untrusted. See prompt injection |
| No lasting change | The skill ends with the prompt | Store good examples somewhere the next request can read them |
The hostile-examples row has a name. In April 2024 Anthropic published many-shot jailbreaking: a single prompt filled with faux dialogues in which an assistant answers harmful questions, followed by the real target question. The researchers tested up to 256 such dialogues and found the rate of harmful responses rose as the count grew past a certain number. It is many-shot in-context learning pointed the wrong way. Anthropic reported that one method, which classifies and modifies the prompt before the model sees it, cut the attack success rate in one case from 61% to 2%.
In-Context Learning vs Fine-Tuning, RAG and Continual Learning
In-context learning is one of four ways to make a model better at your task, and they differ in what changes and how long the change lasts.
| Question | In-context learning | Fine-tuning | RAG | Continual learning |
|---|---|---|---|---|
| What changes | The prompt | The model's weights | The prompt, filled by a search step | The weights, updated over time |
| How long it lasts | One request | Until the next training run | One request, re-retrieved each time | Ongoing |
| Setup cost | Write the examples | Labeled data and a training run | A searchable index of your sources | Repeated training plus safeguards against forgetting |
| Cost per request | More tokens per call | Normal | Retrieval plus extra tokens | Normal |
| Easy to update | Edit the prompt | Retrain | Edit the source documents | Retrain on new data |
| Best for | Format, tone, fast iteration | Stable, high-volume, narrow tasks | Facts that change or are private | Research systems that must adapt over months |
In practice these stack. Retrieval-augmented generation is in-context learning with a search step in front: the retriever picks the passages, and the model learns from them in context. Context engineering is the discipline of deciding what goes into that context, and context rot is what happens when too much of the wrong material goes in.
Connection to Taskade
Every Taskade AI Agent runs on in-context learning on every request. The model stays the same. What changes is what the agent reads before it answers. You train an agent by adding knowledge: files, links, projects, and media. Its persistent memory keeps context across chats. Taskade EVE, the agent that builds Taskade Genesis apps, saves its notes as readable Taskade projects in a projects/memory folder and keeps a TASKS.md list of what is done and what remains, so the next session starts from the record instead of from zero. You can open, edit, or delete any of it.

This is Workspace DNA seen through ICL. Memory, your projects and databases, decides what goes into the context. Intelligence, your agents running on frontier models from top AI labs, learns the task from that context. Execution, your automations across 100+ bidirectional integrations, writes results back into Memory. It also explains why a well-kept project with clear examples gives an agent better answers without any retraining. The agent did not learn on its own. It read better material. /connect/dna draws that loop as a wiring map.
Each ingredient of a good in-context prompt has a home in the workspace:
| ICL ingredient | Where it lives in Taskade | Learn more |
|---|---|---|
| Instruction | The agent's role, tone, and instructions | Custom agents |
| Worked examples | A project of model outputs you add as knowledge | Agent knowledge |
| Retrieved facts | Files, links, projects, and media the agent reads | Agent knowledge |
| Running state | Persistent memory, plus Taskade EVE's notes and TASKS.md | Taskade EVE |
| Fresh input | The chat message, or a trigger from one of your automations | Automations |
What You Would Build in Taskade
You already teach by example when you forward a teammate "the good version" of a report and ask for the next one in the same style. The trouble is that the good examples live in someone's inbox, so every new request starts from a blank prompt.
In Taskade you would describe an example library for a support desk. One project holds a short list of model replies: a refund, a bug report, a feature request, and an angry customer, each with the reply your team was proud of. A support agent is given that project as knowledge and a brief that says "match these." An automation sends each new ticket to the agent and posts the draft back to the ticket row, where a person reviews it before it goes out. When a reply is especially good, you add it to the library, and every draft after that has a better example to learn from in context.
Describe yours and build it free →
Related Concepts
- Few-Shot Learning: how to write the examples that in-context learning reads
- Zero-Shot Learning: the instruction-only end of the scale
- Context Engineering: choosing what goes into the context, the practical side of ICL
- Context Window: the hard limit on how much a model can learn from at once
- Fine-Tuning: the weight-changing alternative
- Retrieval-Augmented Generation: in-context learning with a search step in front
- Continual Learning: models that change their weights as they go
- Emergent Behavior: abilities, ICL among them, that grow sharply with scale
- Mechanistic Interpretability: the research field that found induction heads
- Prompt Injection: what happens when untrusted text in the context steers the model
Frequently Asked Questions About In-Context Learning
What is in-context learning in simple terms?
In-context learning is a language model learning a task from the text in its prompt. You give an instruction and a few examples, and the model continues the pattern on a new input. The model is not retrained. It adapts only for that request.
Who introduced the term in-context learning?
The term was popularized by the 2020 GPT-3 paper, Language Models are Few-Shot Learners by Brown et al. at OpenAI. The authors used it for the inner loop of meta-learning, the adaptation that happens within one forward pass, and split it into zero-shot, one-shot, and few-shot settings.
Is in-context learning the same as few-shot learning?
Few-shot prompting is the most common form of in-context learning, but not the only one. Zero-shot (instruction only), one-shot, and many-shot prompts are in-context learning too. The few-shot learning page covers how to write good examples.
Does in-context learning change the model's weights?
No. The weights stay frozen, and the adaptation disappears when the prompt ends. A 2025 Google Research paper argued that the effect is mathematically equivalent to a small, temporary low-rank update to part of the model, but nothing is stored after the request.
What are induction heads?
Induction heads are attention heads that complete patterns of the form [A][B] ... [A] → [B]: after seeing a pair once, they predict the second token when the first appears again. Olsson et al. (2022) presented preliminary, indirect evidence for the hypothesis that they drive most in-context learning, with strong causal evidence in small attention-only models and correlational evidence in larger ones. A 2025 study found that in larger models, function vector heads matter more for few-shot performance.
When should I use in-context learning instead of fine-tuning?
Use in-context learning when you need to iterate fast, change behavior often, or cannot collect a labeled dataset. Consider fine-tuning when the task is narrow, high-volume, and stable enough that paying for a training run beats paying for long prompts on every call.
Do more examples always help?
Often, up to a point. Google DeepMind's many-shot study found strong gains from hundreds or thousands of examples. But every example costs tokens, and a crowded context can bury the parts that matter, a problem known as context rot. Near-duplicate examples add cost without adding much new signal, so varied examples that cover the real cases are a better use of the space.
Can in-context learning be used against a model?
Yes. Because the model follows the examples it reads, an attacker can fill a prompt with fake examples of the behavior they want. Anthropic's 2024 many-shot jailbreaking research showed this with up to 256 faux dialogues. The same risk applies to any untrusted text an agent reads, which is why prompt injection defenses matter.
How do Taskade agents use in-context learning?
Every Taskade AI Agent learns each task from its context: the knowledge you add, its persistent memory, and the projects it reads. Better-organized projects give it better examples, so its answers improve without retraining. The agent does not learn on its own. It works from what your workspace holds.