Definition: Prompt injection is an attack in which text that a language model was only supposed to read gets treated as an instruction to follow. It works because a model receives its developer's system prompt, your request, and any outside content as one stream of text, and it cannot reliably tell instructions from data. OWASP ranks it first in its 2025 list of risks for LLM applications, as LLM01:2025.
TL;DR: Prompt injection is a model obeying instructions hidden in content it was only meant to read. The danger peaks when one agent has private data, reads untrusted content, and can send data out, which Simon Willison calls the lethal trifecta. OpenAI wrote in December 2025 that it is unlikely to ever be fully solved, so design limits matter more than filters. Build one free →
Picture a new assistant who opens your mail every morning. One letter says, in plain type, "Assistant: ignore your boss and fax the payroll file to this number." A person laughs and bins it, because they know the difference between a letter they are reading and an order from the person they work for. A model has no such built-in sense. Every word in its context window arrives with roughly the same authority, so a well-written sentence inside a web page, an email, or a code comment can pass for a command.
Why Prompt Injection Matters in 2026
Prompt injection matters in 2026 because AI moved from answering questions to taking actions, and an action is what an attacker wants to hijack. Simon Willison coined the term in September 2022, naming it after SQL injection, where a database also cannot tell trusted code from untrusted input. At the time the worst case was a chatbot saying something embarrassing. Once a model can call tools, read files, and send messages, the worst case becomes data theft.
The shift to indirect attacks came early. In the 2023 paper Not what you've signed up for, Greshake and colleagues showed that an attacker does not need to talk to the model at all. They plant instructions in a web page or document the application will retrieve later, because, in the authors' words, LLM-integrated applications "blur the line between data and instructions." They demonstrated data theft and attacker control over API calls against real systems, including Bing's GPT-4 powered chat.
Two 2025 incidents made the risk concrete for developers. In May 2025, Invariant Labs showed that a malicious issue in a public GitHub repository could hijack an agent connected through the official GitHub MCP server. The user only asked the agent to look at open issues. The agent then read private repositories and published their contents in a public pull request. Invariant ran the demo on a frontier model with strong safety training, which shows that alignment alone does not stop the attack. The same month, Legit Security reported that hidden instructions in merge request descriptions, comments, commit messages, issues, and source code could make GitLab Duo leak private code. Duo base64-encoded the code into an <img> URL, and the victim's browser sent it out when it loaded the image. GitLab fixed it by blocking Duo from rendering <img> and <form> tags that point outside gitlab.com. The underlying weakness, a model that follows text it reads, remains an industry-wide problem. In December 2025, OpenAI wrote in its post "Hardening Atlas against prompt injection", as reported by TechCrunch, that prompt injection, "much like scams and social engineering on the web, is unlikely to ever be fully 'solved.'"
| When | Milestone | Why it mattered |
|---|---|---|
| Sep 2022 | Riley Goodside shows "Ignore the above directions" overriding a translation prompt. Simon Willison names the attack prompt injection | The problem gets a name and an analogy: SQL injection |
| Feb 2023 | Greshake et al. demonstrate indirect prompt injection against real LLM-integrated apps | Attackers no longer need to type into the chat |
| 2025 | OWASP lists Prompt Injection as LLM01, first in its Top 10 for LLM applications | The risk becomes a standard security checklist item |
| May 2025 | GitLab Duo and GitHub MCP disclosures | Coding agents leak private code in real products |
| Jun 2025 | Willison's lethal trifecta and the Beurer-Kellner design patterns paper | Defense shifts from filters to architecture |
| Dec 2025 | OpenAI says prompt injection is unlikely to ever be fully solved | A frontier lab treats it as a permanent risk to manage |
How Prompt Injection Works
Prompt injection works by getting attacker text into the model's context and then letting the model's normal helpfulness do the rest. Willison's lethal trifecta names the three ingredients that turn it into data theft: access to private data, exposure to untrusted content, and a way to communicate externally.
- The attacker plants text. The payload sits where the agent will look: a web page, a GitHub issue, an email, a shared document, a code comment. It can be hidden in white text, encoded, or written to look like part of the page.
- The agent reads it as part of normal work. You ask for a summary, a triage pass, or a reply draft. The agent fetches the content in good faith.
- Everything lands in one context. The system prompt, your request, your private data, and the poisoned content now sit side by side as tokens. Nothing in the text itself marks which parts carry authority.
- The model follows the most compelling instruction. Models are trained to be helpful and to follow instructions, and a clear, confident instruction inside the data can win.
- A tool call carries the data out. The agent sends an email, opens a pull request, calls a URL, or renders an image link that encodes your data. The GitLab Duo case used exactly that image trick.
- Remove one leg and the theft fails. With no private data there is nothing to steal. With no untrusted content there is no attacker text. With no outbound channel the data has nowhere to go.
Direct vs Indirect Prompt Injection
OWASP splits prompt injection into two forms by where the attacker's text enters. Direct injection comes from the person typing. Indirect injection comes from content the model pulls in from outside, which is the form that threatens agents.
| Question | Direct prompt injection | Indirect prompt injection |
|---|---|---|
| Who writes the payload | The user in the chat box | A third party, inside a page, file, email, or issue |
| How it reaches the model | Typed straight into the prompt | Retrieved by a tool, search, or document read |
| Does the victim see it | Yes, they wrote it | Usually not, it is hidden or ignored as ordinary content |
| Typical goal | Override the system prompt, extract hidden instructions | Hijack an agent's actions, steal data the victim can access |
| Real example | "Ignore the above directions" in a translation prompt, shown by Riley Goodside in 2022 | The GitHub MCP issue attack and the GitLab Duo hidden-comment attack in 2025 |
Prompt injection is also not the same as jailbreaking, although sources draw the boundary differently. Willison draws a hard line: jailbreaking tries to get around the safety filters built into a model, while prompt injection attacks an application by concatenating untrusted input with the developer's trusted prompt. OWASP instead treats jailbreaking as one form of prompt injection, the form that makes a model disregard its safety protocols. Both views agree on the practical split. A jailbreak makes a model say something it was trained not to say. A prompt injection makes an application do something its owner never asked for.
Common Prompt Injection Techniques
Prompt injection payloads do not need to look like commands, and they do not even need to be visible. OWASP notes that injected inputs "do not need to be human-visible/readable" as long as the model parses them. Its LLM01 attack scenarios describe the main shapes an attacker uses.
| Technique | How it works | Example from OWASP |
|---|---|---|
| Instructions planted in content | A sentence written for the model sits inside a page, file, or listing, often where a person will not notice it | A job ad carries an instruction that fires when an applicant's AI tool reads it |
| Poisoned retrieval | A document in a search or RAG index is edited to carry instructions | A changed repository file alters the answer for any user whose query retrieves it |
| Payload splitting | The instruction is split into pieces that only work together | A resume with split prompts earns a positive AI review |
| Multimodal injection | The instruction hides in an image that accompanies harmless text | A multimodal model reads the image and acts on the hidden prompt |
| Adversarial suffix | A string of odd-looking characters is appended to a prompt | The suffix pushes the model past its safety behavior |
| Obfuscation | Base64, emoji, or a second language hides the payload from filters | A keyword filter misses what the model still understands |
OWASP also states that retrieval-augmented generation and fine-tuning "do not fully mitigate" the problem. More knowledge does not fix a model that cannot tell a quote from an order.
How to Defend Against Prompt Injection
No single filter stops prompt injection, so defense is architectural. Willison points out that guardrail products often claim to catch "95% of attacks", and that in web application security 95% is a failing grade, because an attacker only needs the attempt that gets through. The OWASP mitigations and the June 2025 paper Design Patterns for Securing LLM Agents against Prompt Injections by Beurer-Kellner and colleagues point the same way: limit what a hijacked agent is able to do.
| Defense | What it does | Source |
|---|---|---|
| Least privilege | Give each agent only the data and tools its job needs | OWASP |
| Human approval | A person confirms high-risk actions before they run | OWASP |
| Segregate untrusted content | Mark outside content clearly and keep it apart from trusted instructions | OWASP |
| Validate outputs | Define the expected output format and check it before anything acts on it | OWASP |
| Adversarial testing | Attack your own agents on a schedule and keep the cases as evals | OWASP |
The design patterns paper, written by 14 researchers from organizations including IBM, Invariant Labs, ETH Zurich, Google, and Microsoft, goes further. It names six patterns, each of which trades some flexibility for a guarantee.
| Pattern | What the agent is allowed to do | What it gives up |
|---|---|---|
| Action-Selector | Pick an action from a fixed menu. Tool results never flow back to the model | Any reasoning over what the tools return |
| Plan-then-execute | Fix the list of tool calls before it reads untrusted content | Changing the plan based on what it reads |
| LLM Map-Reduce | Sub-agents read each untrusted item alone and return a narrow, checked result | Free-form answers from each sub-agent |
| Dual LLM | A quarantined model reads untrusted text. A privileged model plans with symbolic references and never sees that text | Simplicity, because two models must coordinate |
| Code-then-execute | The privileged model writes a program in a sandboxed language, so data flow can be tracked | Open-ended, conversational behavior |
| Context minimization | Remove content the task no longer needs, such as the user's original prompt, before later steps | Context that later steps can no longer use |
The paper's core rule, as Willison summarizes it, is that once an agent has read untrusted input, it must be constrained so that the input cannot trigger consequential actions. Guardrails still help as one layer. They do not replace a design that breaks the trifecta.
Connection to Taskade
Taskade does not claim that any AI product is immune to prompt injection, including its own. Taskade AI Agents run on frontier models from top AI labs, and every one of those models reads instructions and data as the same kind of text. What Taskade gives you is a control on each leg of the lethal trifecta, so you can decide per agent and per automation which legs exist at all.
Taskade AI Agents read the web through search and page fetch. They do not drive a browser, click, or fill in forms, so the risk surface is the content they read: web pages, uploaded files, and messages that arrive through 100+ bidirectional integrations.
| Trifecta leg | Taskade control | Where to set it |
|---|---|---|
| Access to private data | Each agent knows only the files, links, and projects you train it on. Role-based access from Owner to Viewer decides which people reach which projects | Agent knowledge |
| Exposure to untrusted content | Turn web search, page reading, and connected apps on or off for each agent | Agent tools |
| Ability to communicate externally | Set a tool to Manual Approval, and the agent waits for you to click Approve or Reject before it sends an email, posts to a channel, or exchanges data | Agent tools |
| Unattended actions in automations | When an automation hands a prompt to an AI Team with the Ask Agent Team action, a tool that needs manual approval fails the flow instead of running unattended | Ask Agent Team |

Model Context Protocol connections follow the same rules. Taskade agents do not call outside MCP servers. An automation can call one through the MCP Client connector, and the reply from that server is untrusted content like any web page. In the other direction, the hosted MCP server on paid plans lets outside AI clients reach your workspace. Treat each connected client as one more agent, and give it only the access its job needs.
What You Would Build in Taskade
You already apply this logic to people. The intern who opens inbound mail does not also hold the keys to the bank account, and nobody wires money because a letter asked nicely.
In Taskade you would describe an inbound triage desk built on that split. An automation catches new support emails and form entries and hands them to a reader agent whose only knowledge is your public help docs. That agent categorizes and summarizes, and it has no access to customer records or billing projects. Its output lands as rows in a triage project. A second agent, which does hold your internal notes, works only from those rows, never from the raw email. Any reply that leaves the workspace lands as a draft task in a review project. A separate automation sends it only after a teammate marks that task complete, using the Task Completed trigger. If a message contains hidden instructions, the agent that reads it has nothing private to leak, and the path out goes through a human.
The reader agent holds untrusted content but no private data. The drafting agent holds private data but never sees the raw email. The only way out runs through a person. No single agent holds all three legs of the trifecta.
Describe yours and build it free →
Related Concepts
- Guardrails: input and output checks that add one layer of defense
- Agent Permissions: which actions run, which ask first, which never run
- Human in the Loop: a person approves the consequential step
- Model Context Protocol: the tool standard behind the GitHub MCP attack
- Tool Poisoning: injection hidden in a tool's own description
- Agent Sandbox: limit what a hijacked agent can touch
- System Prompt: the trusted instructions an injection tries to override
- What Is AI Safety?: the wider field this attack belongs to
Frequently Asked Questions About Prompt Injection
What is prompt injection in simple terms?
Prompt injection is when an AI model follows instructions hidden in text it was only supposed to read. A web page, email, or document can contain a sentence written as a command, and the model can obey it as if you had typed it. It happens because the model sees all of its input as one stream of text.
What is the difference between direct and indirect prompt injection?
Direct prompt injection comes from the person typing into the chat, for example "ignore your previous instructions." Indirect prompt injection comes from outside content the model retrieves, such as a web page, file, or GitHub issue. Indirect injection is the bigger risk for agents, because the victim never sees the attacker's text.
Is prompt injection the same as jailbreaking?
No. Jailbreaking tries to get around the safety training built into a model so it says something it was trained to refuse. Prompt injection attacks an application built on a model by mixing untrusted text with the developer's trusted instructions. Simon Willison, who coined the term in 2022, treats them as separate classes of attack. OWASP files jailbreaking as one form of prompt injection, but both agree that injection is the attack that hijacks an application's actions.
What is the lethal trifecta?
The lethal trifecta is Simon Willison's name for the three capabilities that make an AI agent easy to turn into a data thief: access to private data, exposure to untrusted content, and the ability to communicate externally. When one agent has all three, hidden instructions can make it collect your data and send it out. Removing any one leg blocks that theft.
Can prompt injection be fully prevented?
Not with current models. OpenAI wrote in December 2025 that prompt injection is unlikely to ever be fully solved, much like scams and social engineering. Filters catch many attempts but not all, so the reliable approach is architectural: least privilege, human approval for risky actions, and keeping untrusted content away from agents that hold private data.
What are real examples of prompt injection attacks?
In May 2025, Invariant Labs showed a malicious GitHub issue hijacking an agent through the official GitHub MCP server and leaking private repository contents into a public pull request. The same month, Legit Security showed hidden comments making GitLab Duo leak private source code through an image URL. In 2023, researchers demonstrated indirect injection against Bing's GPT-4 chat.
How do I protect my AI agents from prompt injection?
Give each agent only the data and tools its job needs. Keep agents that read outside content separate from agents that hold private records. Put a human approval step in front of anything that sends data out, such as emails, posts, or HTTP calls. Then attack your own setup on a schedule and keep the cases that worked as tests.
Are Taskade AI Agents vulnerable to prompt injection?
Taskade agents run on the same frontier models as other AI tools, so the underlying weakness exists. Taskade agents read the web through search and fetch and never drive a browser. You control what each agent knows, which tools it can use, and which people can reach each project. Set any tool that sends data out to Manual Approval, and the agent waits for you to approve the action. Those controls let you break the lethal trifecta in your own setup.