Indirect Prompt Injection: Risks in AI Browsing, RAG and Document Workflows
How indirect prompt injection works when AI systems read web pages, retrieved documents, emails and tool outputs, why it is hard to prevent, and the trust boundaries, content isolation, least privilege, action confirmation and monitoring that limit damage.
Quick answer
Indirect prompt injection hides instructions in content an AI system reads while working for someone else: web pages, retrieved documents, emails, reviews and tool outputs. Because models cannot reliably separate data from instructions, design so that a successful injection has limited effect: treat all retrieved and tool content as untrusted, keep sensitive data and powerful tools away from contexts that process untrusted content, enforce permissions outside the model, require user confirmation for consequential or outbound actions, block exfiltration channels and test with planted content.
Where This Fits
The general prompt injection guide, including direct attacks and defence in depth, is prompt injection prevention. This article goes deeper on content that enters through retrieval, browsing, communications and tools. Tool controls are in AI tool security, MCP-specific risks in MCP security and testing in AI red teaming.
How the Attack Works
A user asks an assistant to summarize a web page, triage their inbox or answer a question from the company wiki. Somewhere in that content, an attacker has placed text written as instructions: perhaps visible, perhaps hidden in white text, metadata or alt text. The model reads the content to do its job and may treat the planted text as instructions, for example to reveal information, change its answer or call a tool.
The core problem is that language models process instructions and data in the same channel. Training and filtering reduce susceptibility, but no current approach makes models reliably ignore well-crafted instructions in content.
Entry Points
| Entry point | Example | Who can write it |
|---|---|---|
| Web pages and search results | Browsing agent summarizes a page | Anyone on the internet |
| Emails and messages | Assistant triages an inbox | Anyone who can send email |
| Shared documents and wikis | RAG over company drive | Employees, guests, compromised accounts |
| User-generated content | Product reviews, support tickets, forum posts | Customers or the public |
| Tool and API outputs | Third-party API or MCP server response | The tool provider or anyone who can influence its data |
| Code repositories and issues | Coding agent reads issues and READMEs | Contributors, external reporters |
Why Exfiltration Is the Main Danger
The most damaging pattern combines three ingredients: access to private data, exposure to untrusted content and a channel to send data out. If an assistant can read your inbox, process an attacker's email and render images or links, planted instructions might encode private data into a URL that the attacker's server receives when loaded. Removing any one ingredient breaks the chain. OpenAI's guidance on designing agents to resist prompt injection describes analysing where data flows to (sinks) and asking users to confirm or blocking steps that would send conversation data to third parties.
Building an assistant that reads email, documents or the web?
ZSpace Labs designs AI architectures with trust boundaries and confirmation flows that contain injection risk. See AI development services.
Trust Boundaries and Content Isolation
Draw trust boundaries explicitly. Developer-authored system instructions are trusted. User input, retrieved documents, web content, emails and tool outputs are not. Microsoft's agent safety guidance makes the same point: only developer-controlled content belongs in system messages, and tool and retrieved content must be treated as untrusted.
Isolation patterns reduce exposure. A quarantined model can process untrusted content and return only constrained, structured results, such as a classification or extracted fields, to a privileged component that holds tools. Research on design patterns for securing LLM agents describes several such patterns, all based on preventing untrusted input from triggering consequential actions.
Least Privilege and Action Confirmation
Give each assistant only the tools and data its task needs, scoped to the current user. A summarizer does not need a send-email tool; an inbox triage assistant does not need access to file shares. Enforce permissions in the systems the tools call, not in the prompt.
Require explicit user confirmation for consequential actions, especially outbound ones: sending messages, sharing files, making payments, changing settings, visiting URLs constructed from private data. Show the user exactly what will happen, including recipients and content. See AI agent access control and human-in-the-loop AI.
Additional Controls
- Block or proxy automatic loading of external images and links in AI output
- Restrict network egress for agents and tools to allow-listed domains
- Mark retrieved content with source and trust level in the context
- Limit who can write to indexed sources, and monitor changes
- Scan incoming content with injection classifiers as one signal, not the only defence
- Log tool calls with the content that preceded them for investigation
Advantages and Limitations of Current Defences
Architectural controls such as least privilege, isolation, confirmation and egress restrictions are robust because they do not depend on the model resisting manipulation. They also reduce capability and add friction. Detection classifiers and model hardening lower success rates but can be bypassed. Combine both, and decide consciously which capabilities are worth the residual risk.
How to Reduce Indirect Injection Risk Step by Step
- 1. Map every untrusted content source the system reads
- 2. Map data and tools reachable in the same context
- 3. Remove unnecessary tools and data from those contexts
- 4. Add confirmation for outbound and consequential actions
- 5. Close exfiltration channels such as auto-loaded links
- 6. Plant test injections in each source type and test
- 7. Monitor tool calls following untrusted content
Indirect Injection in Coding Agents and Browsing Agents
Coding agents read issues, pull request comments, documentation and dependency files, any of which outside contributors may write. Restrict which events can trigger agents, run them with tokens limited to branches and pull requests, and require human review before merge; see AI coding agents. Browsing agents read arbitrary web pages, so keep them separate from private data and sensitive tools, restrict where they can submit forms or send data, and require confirmation before purchases, logins or sharing information.
Monitoring and Response
Assume some injections will get through and design detection. Log tool calls together with the content that preceded them, alert on unusual patterns such as outbound messages to new recipients or tool calls right after reading external content, and review samples of agent trajectories. Keep the ability to disable specific tools or sources quickly, and when an injection is found in indexed content, remove it, search for similar content and add the case to the regression suite. See LLM observability.
Designing Confirmation That Works
Confirmation is one of the strongest controls against injected instructions, but only if users actually review what they confirm. Show the concrete action: recipient, content, amount, destination URL. Highlight anything unusual, such as an external recipient, a link to an unfamiliar domain or data the user did not mention. Avoid confirmation prompts for trivial actions, which train users to click through. For high-risk actions, require users to edit or type a value rather than clicking a default button.
Never let the model write the confirmation text alone; generate the summary from the structured tool call so a manipulated model cannot describe a harmful action innocently. Interface patterns are covered in AI copilot UX.
Worked Example
An illustrative scenario, not a client case: a support team's assistant answers from tickets, including text customers submit. Testing shows a planted instruction in a ticket can make the assistant append a link to its answer. The team stops rendering links from model output except to allow-listed domains, marks customer text as untrusted in the prompt, removes an unnecessary tool that could update ticket priority and adds the planted ticket to the regression suite.
Common Mistakes
- Assuming only users can inject instructions
- Giving reading assistants sending or sharing tools
- Rendering model-generated links and images automatically
- Relying on 'ignore instructions in documents' prompts
- Indexing content anyone can edit without monitoring
Want your assistant tested against injected content?
Talk to ZSpace Labs about an AI security review focused on retrieval, browsing, email and tool risks.
Conclusion
Any content an AI system reads can carry instructions. Treat it as untrusted, keep powerful tools and private data out of reach of untrusted contexts, confirm consequential actions with users, close exfiltration channels and keep testing with planted content.
Common questions
An attack where malicious instructions are placed in content an AI system will read later, such as a web page, document, email, review or tool response, rather than typed by the user, so the model may follow them while performing an unrelated task.