Prompt Injection
What Is Prompt Injection?
Prompt injection is an AI security vulnerability in which a language model follows untrusted or unauthorized instructions instead of the rules it was meant to follow. A prompt injection attack can change an AI's answer, manipulate a recommendation, attempt to expose information, or cause an AI agent to request an unintended action.
This is not simply an unusual or poorly written prompt. It is a security problem that arises when an application gives a generative AI system access to user messages, documents, websites, emails, images, private data, or software tools. The model may struggle to treat some of that material as data only, rather than as instructions. OWASP identifies prompt injection as a leading risk for LLM applications.
How Prompt Injection Works in Generative AI
Generative AI applications combine several kinds of input before producing an answer or taking action. An attacker tries to make untrusted content compete with, or override, the application's intended instructions.
- The application provides trusted system or developer instructions, such as its role, safety rules, and permitted tasks.
- A user asks the AI to complete a task, such as summarize an email, research a topic, or update a record.
- The application may retrieve untrusted content from a web page, uploaded file, knowledge base, email, or image.
- The model interprets all of this material and generates a response. It may mistakenly treat hostile content as a command.
- If the system can use tools, it may attempt a search, database query, message draft, file operation, or other tool call based on that interpretation.
- The application returns an answer or performs an action, unless permission checks, validation, or human approval stop it.
For example, a research assistant might read a web page containing a hidden instruction that says it should favor a particular result. The core problem is not that the page contains text. It is that the AI may give that text influence it should not have.
Key Components and Trust Boundaries
A useful way to assess prompt injection is to map trust boundaries. Trusted instructions define what the application is supposed to do. Untrusted content includes anything supplied by users or obtained from outside sources, even if it looks legitimate.
Other important components are sensitive data, connected tools, permissions, and human approval. Sensitive data can include private documents, customer records, credentials, or internal policies. Tools can include email, calendars, payment systems, databases, browsers, and code execution environments. The most important security question is often not whether an AI can read malicious content, but what it can access or do after reading it.
Clear prompts still matter for reliable AI use. Guidance on writing prompts that work can help teams express intended tasks clearly, but wording alone is not a security boundary. Access controls and application design must enforce the boundary.
Types of Prompt Injection Attacks
Prompt injection can arrive through more than a chat message. The source and the AI's available permissions shape the likely impact.
| Type | Where the instruction appears | Safe high-level example | Likely impact |
|---|---|---|---|
| Direct prompt injection | A user message sent directly to the AI | A user asks a chatbot to disregard its stated task and follow a conflicting request. | Misleading output, policy bypass attempts, or system prompt disclosure attempts. |
| Indirect prompt injection | External content the AI reads, such as an email, web page, or document | An email contains text intended to influence an AI assistant summarizing the inbox. | Manipulated summaries, recommendations, or attempted data exposure. |
| Multimodal or image injection | An image, audio file, PDF, or other non-text input | Small text inside an image tries to influence an AI that can interpret images. | Hidden instructions may affect analysis or downstream decisions. |
| Multi-turn or persistent injection | Several messages or stored content over time | An attacker gradually introduces conflicting instructions across a long conversation. | Context drift, altered behavior, or delayed unsafe requests. |
| Agentic prompt injection | Content read by an AI with access to tools or workflows | A web-research agent encounters content that attempts to redirect its next task. | Unauthorized workflow requests or harmful tool use if safeguards fail. |
Prompt Injection Examples and Real-World Scenarios
The same weakness looks different depending on what the AI application can read and what it is allowed to do.
- A customer support chatbot receives a message designed to make it give inaccurate return-policy guidance. The immediate harm is poor service and loss of trust.
- An email assistant summarizes a message containing embedded instructions that attempt to alter the summary or influence a reply. The user may receive misleading advice.
- A document search or retrieval-augmented generation system reads an uploaded report that tries to influence its answer. Retrieved material should be evidence, not authority over application rules.
- A web-browsing research assistant encounters a page crafted to promote a particular vendor regardless of the user's stated criteria. The result may be a biased recommendation.
- An AI agent connected to business tools reads hostile content that attempts to trigger a payment, data export, or external message. Strong authorization can prevent the request from becoming an actual action.
These examples show why organizations should treat untrusted AI inputs much like untrusted attachments or web content. Teams handling private information should also consider how sensitive data in LLM prompts and traces can be exposed through logs, context, or connected systems.
Prompt Injection vs. Jailbreaking
Prompt injection and jailbreaking overlap, but they describe different goals and attack paths. Both may involve attempts to influence a model beyond its intended behavior.
| Factor | Prompt injection | Jailbreaking |
|---|---|---|
| Primary goal | Make an AI treat untrusted content as instructions. | Make an AI bypass safety rules or restrictions. |
| Instruction source | May come directly from a user or indirectly from a document, website, email, or image. | Usually comes from the user interacting with the model. |
| Affected party | Can affect a user, an organization, or an automated workflow that reads external content. | Usually affects the conversation and the user's requested output. |
| Typical risk | Data exposure attempts, manipulated recommendations, or unauthorized tool requests. | Generation of restricted content or behavior outside the intended policy. |
A user trying to override a chatbot's safeguards can be both a jailbreak attempt and a direct prompt injection attempt. Indirect prompt injection is the clearer distinction because the hostile instruction may be placed by someone other than the person using the AI.
Why Understanding Prompt Injection Matters
Recognizing this risk helps people choose safer AI designs and use AI systems more carefully.
- It encourages safer assistants that limit access to the data and tools needed for a task.
- It helps protect private data by separating information retrieval from permission to disclose it.
- It improves recommendation quality by treating web pages and documents as evidence to evaluate, not commands to obey.
- It supports controlled automation by adding approval steps before high-impact actions.
- It creates clearer accountability because teams can identify who authorized an action and what information influenced it.
Practical Limits and Common Pitfalls
There is no single prompt that permanently solves prompt injection. Language models are designed to interpret language flexibly, so a stronger system prompt alone cannot reliably turn all hostile language into harmless data.
- Putting passwords, access tokens, or other secrets in a system prompt and assuming they cannot be revealed.
- Giving an AI broad permissions, such as unrestricted email access or the ability to make irreversible changes.
- Trusting retrieved documents, search results, or web pages simply because they came from an approved source.
- Relying only on keyword filters, which can miss indirect, rewritten, encoded, or image-based instructions.
- Assuming a final refusal proves nothing happened. A tool call may have been attempted before the final response was shown.
- Testing only one-message chat attacks while ignoring files, images, long conversations, and multi-step workflows.
How to Prevent Prompt Injection Attacks
Effective prevention is layered. The goal is to reduce what an injected instruction can influence, then verify sensitive actions outside the model.
- Minimize access. Give the AI only the data, tools, and permissions required for its current task.
- Separate instructions from untrusted content wherever possible, and clearly label retrieved material as reference data.
- Require explicit human confirmation before consequential actions, such as sending messages, transferring money, changing records, or sharing sensitive information.
- Use allowlisted tools, scoped credentials, and short-lived permissions rather than broad standing access.
- Validate tool inputs and outputs with ordinary application rules. For example, enforce recipient, amount, and data-access policies outside the model.
- Isolate sensitive data so a model that reads external content cannot freely retrieve unrelated private records.
- Log tool requests, authorization decisions, and outcomes. Monitoring helps teams investigate suspicious behavior and improve controls.
- Test realistic adversarial content across documents, websites, images, and multi-turn conversations. The OWASP prompt injection guidance emphasizes defense in depth rather than reliance on instructions alone.
- Give users clear controls to review sources, limit connected accounts, and approve actions before they occur.
A Safer Design Checklist for AI Agents
Before launching an AI agent, evaluate the full application path, not only the model's final answer. Application-level controls are as important as model-level instructions.
- Map every source of data the agent can read, including user input, files, inboxes, websites, and knowledge bases.
- Classify each source by trust level and assume external content can contain hostile instructions.
- Inventory each tool the agent can use and document the precise data or action each permission enables.
- Define approval gates for actions that are irreversible, expensive, externally visible, or privacy-sensitive.
- Create harmless test markers to detect whether protected context is being exposed, without placing real secrets in test prompts.
- Run attack simulations using realistic documents, webpages, images, and conversation sequences.
- Review logs for unexpected tool requests, unusual data flows, and differences between the user's request and the action attempted.
- Reduce access or disable automation when the trust level, authorization, or intended outcome is uncertain.
Frequently Asked Questions
Your Questions, Answered
Don't change this element unless you know what you are doing
How serious is a prompt injection?
Prompt injection can be serious when an AI can access private information or take actions through connected tools. A basic chatbot may only produce an incorrect answer, while an AI agent with email, database, or payment access could create larger privacy, financial, or operational risks. The severity depends on permissions, data access, and whether human approval is required.
What is the difference between jailbreak and prompt injection?
Jailbreaking usually means trying to bypass a model's safety restrictions through a direct interaction. Prompt injection means causing an AI to treat unauthorized text as instructions, including text hidden in external content such as emails, documents, or web pages. A direct request to override safeguards may fit both categories.
How does prompt injection work in generative AI?
It works when a model receives trusted instructions alongside untrusted content and gives the untrusted content too much influence. An attacker may place conflicting instructions in a user message, file, website, email, or image. If the model follows those instructions, it can produce manipulated output or request an unintended tool action.
How do you prevent prompt injection attacks?
Use layered controls. Limit data and tool access, separate untrusted content from application instructions, validate every sensitive tool action, require human approval for consequential steps, use scoped credentials, and test with adversarial documents and multi-step workflows. Do not rely on a system prompt or keyword filter as the only defense.
Are prompt injections illegal?
Prompt injection is a technique, not automatically a crime. Its legality depends on intent, authorization, jurisdiction, and harm. Testing an AI system with permission as part of security research may be lawful, while using prompt injection to access private data, disrupt services, or cause unauthorized actions can violate computer misuse, privacy, fraud, or contract laws.
What is the purpose of a prompt injection attack?
An attacker may seek to manipulate an AI's output, bypass intended restrictions, influence recommendations, obtain sensitive information, or cause an AI-connected system to request an unauthorized action. In indirect attacks, the attacker may also aim to affect users who never directly interact with the malicious content.
on Emergent today


