AI Concepts

Jagged Intelligence

11 min read
On this page (15)

Definition: Jagged intelligence is the uneven skill profile of modern language models: a model can solve competition-level math and still stumble on a question a child answers, because its abilities do not rise together the way a person's do. The edge between what it does well and what it does badly is ragged, and it is not always obvious where the edge sits.

TL;DR: Jagged intelligence means an AI can be brilliant at one task and unreliable at the one right next to it. Andrej Karpathy named it in 2024, and a Harvard and BCG study of 758 consultants showed the cost: AI users were 19 percentage points less likely to be correct on a task outside the frontier. Trust the model per task, not per reputation. Build one free →

Think of a friend who can recite every capital city, then cannot find their own car keys. Or a chess grandmaster who is lost in a grocery store. With people, skill in one area tells you a lot about skill in nearby areas, so you extend trust without thinking about it. A model breaks that habit. Its strengths come in patches, and the patches do not follow the map you would draw for a person.

Why Jagged Intelligence Matters in 2026

Jagged intelligence is the reason a strong benchmark score does not tell you whether a model will handle your task. In July 2024, Andrej Karpathy described the pattern on X: state-of-the-art models do very impressive things, such as solving complex math problems, while failing at some very simple ones. His example was the question of whether 9.11 or 9.9 is bigger, which a leading model got wrong. He contrasted this with people, whose knowledge and problem-solving ability are correlated and grow together from birth to adulthood.

The cost shows up when people extend trust in the wrong direction. In the 2023 field experiment Navigating the Jagged Technological Frontier, researchers from Harvard Business School and others worked with 758 consultants at Boston Consulting Group. On 18 realistic tasks inside the frontier of AI ability, consultants using GPT-4 completed 12.2% more tasks, finished 25.1% faster, and produced results rated more than 40% higher in quality. On one task deliberately placed outside the frontier, the AI users were 19 percentage points less likely to produce a correct solution. The answers still sounded persuasive, which is what makes the failure hard to see.

Princeton cognitive scientist Tom Griffiths, author of The Laws of Thought (Henry Holt, 2026), names jaggedness as a top concern. In a 2026 interview with Brian Keating, he put it this way: if you had a friend who could solve International Math Olympiad problems at a gold medal level, you would trust them with all sorts of other things, but you should not extend that trust to an AI system, because it does not generalize across problems the way people do. Wrong intuitions about these systems, he argued, are a major bottleneck to using them well.

How Jagged Intelligence Works

Jaggedness is not a bug in one model. It follows from how a model gets its skills: each ability grows where its training data, feedback, and practice are dense, and stays thin where they are not.

  1. Skills grow where the signal is dense. A model learns from next-token prediction over huge text collections, then from feedback in fine-tuning. Wherever examples and clear right answers are plentiful, ability climbs fast.
  2. Skills stay thin where it is not. Rare formats, unusual phrasings, and problems that look like a familiar pattern but are not can fall between the peaks.
  3. The peaks do not predict the dips. In a person, hard skills sit on top of easy ones. In a model, a peak can sit next to a dip, so success on a hard task says little about the easy one beside it.
  4. The surface hides the edge. A wrong answer arrives in the same fluent voice as a right one, so nothing on the screen tells you which side of the edge you are on.
  5. Scale and training move the edge, not remove it. Scaling and new training methods smooth some dips, and new emergent abilities add new peaks. The shape stays uneven, and it changes with each model.
  6. You find the edge by testing. Run the task on your own examples, watch where it fails, and keep the test as an eval so you notice when a model update moves the edge.

The practical rule that falls out of this is simple: trust per task, verify per task. Do not transfer trust from one task to another because they seem related.

Jagged Intelligence vs the One-Dimensional AGI Scale

Griffiths argues that the popular picture of intelligence as a single line, with humans at one point and AI climbing toward and past them, is not a productive way to think about AI. Human minds and AI systems both solve computational problems, but they were optimized under different constraints. Humans evolved under limits of a short lifetime, a couple of pounds of neurons, and narrow bandwidth for sharing what we know. AI systems are shaped by large amounts of data, large scale, and the ability to copy weights. Different constraints produce different skill shapes.

Question One-dimensional scale Jagged intelligence
Picture of intelligence One line, from low to high A profile with peaks and dips
What a hard win means The model is "smarter" overall The model is strong at that task
Where trust comes from General reputation and benchmark rank Testing on the task you care about
What an easy failure means A fluke A sign of where the edge sits
How humans compare Ahead, behind, or level Optimized for different problems under different limits
How you plan Wait for the model that does everything Match each task to what the model does well, and check the rest

The two pictures lead to different habits. The single scale invites you to hand over everything once a model is "good enough". The jagged picture asks you to map the strengths, put a check where the edge is, and keep a person or a second reviewer on anything that matters.

Connection to Taskade

Taskade's answer to jaggedness is task-level structure, not blanket trust. Every Taskade AI Agent works against your own projects, records, and files, so a task starts from what is written down in your workspace and not from a model's general reputation. The built-in tools behind each agent let it look a fact up, and 100+ bidirectional integrations bring in the live source of truth from the apps your team already uses. Taskade agents read the web through search and fetch, and you write each agent's brief, so you can scope an agent to the jobs where it performs well. Agents run on frontier models from top AI labs, with Auto as the default. You can set up a second agent as reviewer in a team of agents, and an automation can include steps that check work or hold a decision for a person. This is Workspace DNA in practice: Memory holds your ground truth, Intelligence does the task, and Execution runs the checks.

What You Would Build in Taskade

You already do this by hand when you spot-check a new hire on one skill before you hand them a bigger one. A model needs the same treatment for each task, and most teams skip it because the model looked so capable on something else.

In Taskade you would describe a task trust board. Each row is a task you want an agent to handle: draft a client summary, categorize an invoice, compare two contract clauses, total a column of figures. For each row, an agent runs the task on a small set of examples you wrote, and a second reviewer agent, briefed only on the checklist and the source documents, marks each result pass or fail. The board shows a running record of where each task stands: trusted, trusted with review, or not yet. Tasks marked trusted run in an automation with a spot-check step. Anything the reviewer flags goes to a person, using the same human in the loop line you would draw for a new colleague. When a model changes, you rerun the board and see whether the edge moved.

Describe yours and build it free →

  • AI Sycophancy: another way a fluent answer hides how reliable it is
  • AI Hallucinations: confident wrong answers, often found right at the edge of a model's skill
  • Evals: the tests that map where a model's peaks and dips fall
  • Emergent Behavior: new abilities that appear with scale and add new peaks
  • Scaling Laws: why bigger models raise the average but keep the uneven shape
  • Inductive Bias: the built-in assumptions that shape what a learner generalizes to
  • AI Agents in Taskade: agents scoped to jobs you have tested

Frequently Asked Questions About Jagged Intelligence

What is jagged intelligence in AI?

Jagged intelligence is the uneven skill profile of current AI models. A model can be excellent at a hard task, such as competition math, and unreliable at a simple one, such as comparing two decimals. Andrej Karpathy popularized the term in 2024 to describe how abilities that rise together in people can come apart in a model.

Who coined the term jagged intelligence?

Andrej Karpathy used the term in a July 2024 post on X and gave the 9.11 versus 9.9 comparison as his example. A related idea, the "jagged technological frontier," comes from the 2023 Harvard Business School and BCG field experiment by Dell'Acqua and colleagues. That study measured what the uneven edge costs in real work.

Why does AI fail easy tasks but pass hard ones?

A model learns skills where its training data and feedback are dense, and stays thin elsewhere. A hard task with plenty of practice examples can sit on a peak, while an easy task in an unusual form can sit in a dip. Difficulty for a person is not the same as difficulty for a model.

What did the Harvard and BCG study find about the jagged frontier?

Across 758 BCG consultants, AI users did better on tasks inside the frontier: 12.2% more tasks completed, 25.1% faster, and quality rated more than 40% higher. On one task outside the frontier, they were 19 percentage points less likely to be correct. The lesson is to check which side of the edge a task is on.

Will bigger models remove jagged intelligence?

Bigger models and better training smooth many dips, but the shape stays uneven, and it moves with each new model. Some old failures disappear while new ones appear. Plan for the profile to keep changing, and retest your important tasks when the model changes.

How should I work with a jagged AI?

Trust per task, verify per task. Test each task on your own examples, do not carry trust from one task to a related one, and add a check where the cost of an error is high. For important work, use a second reviewer that never saw the first answer, such as a separate agent with its own brief.

Is jagged intelligence the same as a hallucination?

No. A hallucination is one confident wrong answer, such as a source that does not exist. Jagged intelligence describes the overall shape of what a model can and cannot do. Hallucinations often show up at the dips, so mapping the shape helps you predict where to look.

Are Taskade AI Agents jagged too?

Taskade agents run on the same frontier models as other AI tools, so the uneven shape is present. What Taskade adds is structure around it. An agent works against your own projects, follows a brief you wrote, and can be paired with a reviewer agent or a human check, so each task earns trust on its own record.