Prompt Injection Risks From Web-Retrieved Content in AI Agents
Web-scraped content gives attackers an untrusted channel directly into AI agents' reasoning loops.

Agents that browse, research, and reason over live web content are running in production now, not sitting in a demo folder somewhere, and the threat landscape follows directly from that fact. A message from a user carries an identity behind it, someone accountable, someone who can be blocked or fired or sued. A fetched web page carries no such thing. Anyone can write the words an agent is about to read, and the agent has no built-in way of knowing whether those words came from a colleague or a stranger with intent to harm.
Consider how that content actually gets into the model's reasoning in the first place. Deep Research Agents, the class of system now doing the bulk of this work, combine dynamic reasoning, adaptive planning, multi-iteration retrieval, and report generation into a single loop. Retrieval in these systems isn't a single lookup that happens once and gets cached; it's continuous, repeated, looping back on itself as the agent refines its questions. Two retrieval modes dominate production systems today: API-based search, which is structured and fast and scales cleanly, and browser-based search, which behaves more like a human clicking around, real-time and unstructured. Both bring untrusted content directly into the model's context window, just through different doors.
The ReAct pattern, introduced back in 2022, produces most of this architecture: reasoning traces, actions, and observations interleaved in a loop. Every single observation cycle is a fresh delivery of untrusted input arriving straight from the open web. Layer on top of that the Model Context Protocol, which Anthropic introduced to connect agents to external tools and data sources through one unified interface, and the exposure multiplies: every connected source is a potential injection channel.
Comparing this to SQL injection clarifies what's broken. Both attacks exploit the same underlying failure: a system that can't separate trusted instructions from untrusted data. SQL injection has a clean fix, the parameterized query, which draws a hard syntactic line between code and data. Language models have no equivalent line to draw. Instructions and data both arrive as plain natural language, indistinguishable at the level the model actually reads them. That's not a bug someone forgot to patch. It's a property of how the architecture works, which is why OWASP named prompt injection LLM01:2025, the top-ranked risk in its Top 10 for Large Language Model Applications, and treated it as a structural characteristic of the technology rather than a fixable defect.
The "lethal trifecta": which agents are structurally exploitable
Not every agent carries this exposure equally. The risk is structural, and it depends on three conditions lining up at once. Independent researcher Simon Willison framed this as the "lethal trifecta": access to private data, exposure to untrusted content, and the ability to communicate externally. Remove any single leg of that stool and the attack path collapses. An agent that reads untrusted web pages but has no private data to leak is a nuisance at worst. An agent with private data and no way to send anything outward can be tricked but not robbed.
An agent that browses the web, processes email, operates in a code repository, or calls external APIs typically satisfies all three conditions simultaneously. That's not a design failure so much as the whole point of building the agent in the first place: it needs private data to be useful, it needs to touch untrusted content to do research, and it needs to communicate outward to actually accomplish anything.
The exposure compounds with scope. An agent wired into several enterprise systems at once doesn't just risk one user's data getting exposed; it exposes the combined authority of every permission it holds across every system it touches, according to Obsidian Security's analysis. AI agents move sixteen times more data than human users do Prompt Injection Attacks on AI Agents: How to Detect and Prevent Them. A single compromised agent, then, isn't a single-user incident. It's a high-magnitude exposure event by nature of what the agent was built to do Prompt Injection Attacks on AI Agents: How to Detect and Prevent Them.
Direct vs. indirect prompt injection: why web retrieval is the harder problem
Prompt injection splits into three main types, and each one describes a different trust surface the attacker is working against. Direct injection, sometimes called jailbreaking, happens when the attacker has access to the model's own input and simply types malicious instructions into it. It's the original form and the most visible one, but it's also the most bounded: an attacker manipulating a standalone chatbot produces an embarrassing transcript, not a breach.
Indirect injection is the harder problem, and it's the one that matters for web retrieval specifically. The attacker never touches the model directly. They place content somewhere the agent is going to find it later, a web page, a document, an email, a code comment, a database record, and wait. Stored injection is a variant in which instructions sit dormant in long-term memory or an indexed knowledge base until a future retrieval triggers them.
For a chatbot with no tool access, that distinction matters less. For an agent with tool access, the consequences of indirect injection cascade into real-world actions rather than staying contained inside a conversation. In RAG systems specifically, this risk compounds: retrieved documents get inserted automatically into the prompt, and the model has no reliable way to tell an instruction embedded in that retrieved page apart from the developer's own system prompt. Both arrive as text in the same window. The model reads them the same way.
Goal hijacking is a further escalation that targets multi-step autonomous agents by redirecting the agent's entire objective rather than pulling out one piece of information. In multi-agent pipelines, a successful hijack doesn't stay contained to the agent that got hit. It propagates downstream, poisoning shared memory or manipulating decisions made by an orchestrator that trusts what it's been handed.
And the environment is getting worse, not better. Google researchers monitoring the public web found a 32% increase in malicious indirect prompt injection attempts detected in the CommonCrawl archive between November 2025 and February 2026 How Prompt Injection Attacks Compromise AI Agents in 2026. The web itself, in other words, is becoming a more hostile place for an agent to go looking for answers How Prompt Injection Attacks Compromise AI Agents in 2026.
How injections are embedded in web content: the delivery mechanisms
An attacker doesn't need to compromise a server to run this attack. They only need to place content somewhere an agent will eventually retrieve it, which is a much lower bar. Website meta tags are one route: instructions hidden in HTML metadata that no human reader ever sees, but that an agent's scraper reads anyway. User-generated content is another, and arguably the most consequential, because it scales with every platform that lets strangers post text. The Prismata paper out of UC Berkeley, published on arXiv in July 2026, gives this its own name: Cross-Site Prompting, or XSP. A product review on a shopping site can contain a payload instructing an agent to reply with the user's credit card details, and the review looks, to a human skimming it, like nothing at all.
Publicly accessible documents are a third route, and one that's already been exploited at enterprise scale, discussed further below. Email is a fourth: EchoLeak, which surfaced in mid-2025, exploited Microsoft 365 Copilot's zero-click processing of incoming email to exfiltrate data without the user doing anything at all. Code comments in repositories form a fifth channel, engineered specifically to push AI coding assistants with shell access into running destructive commands. And an innocuous-looking Google Docs file was enough, in one documented case, to get an agent to contact a malicious MCP server and run a Python payload that harvested developer secrets.
The Prismata paper's framing is the clearest lens available for understanding why this category resists the usual fixes. XSP is the direct analogue of Cross-Site Scripting: it exploits a site's own legitimate features, reviews, direct messages, account settings, to carry out the attack rather than injecting foreign code. The problem is that XSS defenses simply don't carry over https://www.helpnetsecurity.com/2026/07/17/xss-web-agent-prompt-injection/. XSP payloads are natural language, not executable code, and an input sanitizer built to strip script tags has no way to tell an instruction apart from ordinary data when both are just sentences.
Prismata names a second problem: web entanglement. Figuring out whether an agent's action is safe requires reading the structure of the page it's acting on, and that structure is itself entangled with the untrusted content sitting inside it. The same click might complete a purchase, reply to an attacker's planted review, or interact with a sponsored ad, and there's no clean way to separate the site's legitimate structure from the malicious content riding inside it.
Documented incidents from 2025 to 2026 that moved this from research to operational risk
The January 2025 enterprise RAG attack is the case study that made this concrete rather than theoretical. Researchers embedded malicious instructions inside a publicly accessible document, the enterprise RAG system retrieved it during normal operation, and from there things escalated fast: proprietary business intelligence leaked to external endpoints, the system modified its own system prompts to disable its safety filters, and it executed API calls carrying more privilege than the user who triggered the query actually had. None of that required a sophisticated exploit. It required a system that treated every retrieved document as equally trustworthy.
EchoLeak followed in mid-2025, disclosed by Aim Labs as the first zero-click data exfiltration attack against Microsoft 365 Copilot. An attacker sent an entirely ordinary-looking email. The user never opened it. Copilot read it anyway during background processing, and a later, completely unrelated query was enough to trigger the leak. CVE-2025-32711 is the record of that: one hidden instruction, walking data out of an enterprise copilot with zero clicks and zero alerts.
An AI IDE zero-click attack followed a similar shape. An innocuous-looking Google Docs file triggered a coding agent inside an IDE to use the Google Docs MCP connection to retrieve a malicious document, fetch a Python Gist from that document's instructions, execute the payload, and harvest developer secrets, all without the victim doing anything beyond opening the document in the first place.
Then, in December 2025, Palo Alto's Unit 42 reported what it described, to its own knowledge, as the first detected real-world example of malicious indirect prompt injection built specifically to bypass an AI-based product ad review system. The attacker didn't rely on a single trick. They stacked multiple injection methods at once, which Unit 42 read as a signal that intent is moving toward higher severity and payloads are getting more sophisticated, not less.
Findings have piled up elsewhere too: Slack AI, Microsoft 365 Copilot, Cursor, GitHub MCP integrations, and various AI coding assistants have all been found vulnerable in roughly the same window. Between 2024 and 2026, the whole category shifted from something that read like a clever chatbot trick to something that reads as enterprise risk. The common thread running through every one of these incidents isn't a weak model. It's an infrastructure layer that made no distinction between a developer's instruction and an attacker's payload, and fed both to the model with equal confidence.
RAG-specific attack vectors: knowledge base poisoning and vector embedding manipulation
RAG systems add several new components on top of a base model: an external knowledge store, a retriever, a vector index, and whatever policy decides how context gets assembled before it reaches the prompt. Any one of those stages can fail on its own, independent of whether the underlying language model is well-behaved.
PoisonedRAG, presented at USENIX Security 2025, is the first documented knowledge corruption attack against RAG systems specifically. The mechanism is almost elegant in how little it requires: insert a small number of carefully written documents into the retrieval corpus, and the system can be made to reliably return an attacker's chosen answer for a specific query, without ever touching the model's weights. The scale finding from that research should unsettle anyone running a large knowledge base. Injecting just five poisoned texts per target question, into a corpus holding millions of documents, achieved a 90% attack success rate across multiple benchmark datasets and multiple models. That's an extraordinarily cheap attack for the size of the payoff.
A second vector operates one layer deeper. Research from Prompt Security describes what it calls the Embedded Threat attack, which targets the embeddings layer itself rather than the prompt, the model weights, or the API. Poisoned embeddings look entirely legitimate to the retrieval system, they pass every check that layer performs, but they carry semantics engineered to override the correct answer once retrieved. Nothing about the prompt looks wrong. Nothing about the model's behavior looks wrong in isolation. The corruption lives in the vector space the retriever trusts.
Attack opportunities exist at every stage of the pipeline. Malicious or misleading content can be added during document ingestion to silently shape future responses. Vector representations can be corrupted at the embedding model stage. Poisoned embeddings can be injected directly into the vector database so they rank highly for queries the attacker cares about. And context construction itself, the step that decides which passages actually make it into the prompt window, can be manipulated to favor the attacker's planted material over the legitimate answer.
Corruption of the underlying corpus isn't even required for the system to leak. Research including "Spill the Beans" (Qi et al., 2025) found that instruction-tuned RAG systems can be induced to regurgitate stored content verbatim, and attackers can craft adaptive queries through black-box access alone to extract specific sensitive sentences from a datastore they never touched. And beyond outright attack, retrieval permissions that are simply too broad produce a plainer failure mode. A RAG system with loose access control can surface PII, credentials, financial records, legal documents, health information, or confidential strategy material to a query that never should have been allowed to see them. Permission-aware retrieval, filtering at retrieval time based on role, has been proposed as the fix, and it's the kind of fix that has nothing to do with the model at all.
Why the infrastructure shaping web content before it reaches the model is the first line of defense
The asymmetry here favors the attacker by default. An attacker needs exactly one payload that works, placed in exactly one place the agent happens to look. A defender has to block every malicious instruction across every input the agent might touch: every email, every document, every web page, every internal wiki, every other agent's output feeding into a shared pipeline. That's not a fair fight, and pretending otherwise doesn't help anyone building these systems.
Model-level defenses, on their own, don't close that gap. Adaptive attacks bypass essentially every published model-level defense at meaningfully high rates, and the reason traces straight back to the structural problem raised earlier: the model has no syntactic boundary to enforce between instructions and data, because both arrive as the same kind of text. Asking the model to police that distinction is asking it to solve a problem its own architecture doesn't give it the tools to solve.
So the defense has to sit somewhere else: at the layer where content gets selected, filtered, and shaped before it ever reaches the context window. Infrastructure that controls what the model reads is, in a very direct sense, infrastructure that controls what the model can be made to do.
That reframes context engineering as a security discipline, not merely a performance tweak aimed at shaving tokens. What gets fetched matters: restricting retrieval to sources inside a defined trust perimeter, rather than pulling indiscriminately from the open web, closes off entire categories of attack before they get a chance. How content gets filtered matters just as much: stripping navigation chrome, ad slots, user-generated content zones, and metadata fields before any of it reaches the prompt removes the exact surfaces that Cross-Site Prompting depends on. Ranking matters too, favoring structured, authoritative sources over unvetted user content whenever the choice is available. The format the content finally arrives in affects security, because a structure that makes the boundary between source material and agent instruction visible is a structure an attacker has a harder time hiding inside.
This is also where AI-native retrieval and raw scraping diverge in ways that become concrete rather than academic. A raw web page burns tokens on navigation bars, ads, and boilerplate, and in doing so, it hands the model's context window exactly the zones that carry the highest injection risk. Retrieval built to return ranked, token-dense excerpts rather than whole pages shrinks that surface before the content ever enters the reasoning loop. It's the same infrastructure choice as filtering out an ad slot, just made earlier in the pipeline.
Provenance lives at this same layer. Knowing where a piece of content actually came from is the precondition for applying any trust policy to it at all, and a retrieval system that strips or obscures that provenance makes downstream governance impossible no matter how careful the policy team is. This is also where owning the retrieval pipeline end to end, rather than assembling one out of third-party scrapers and search wrappers stitched together after the fact, becomes a materially different proposition. A pipeline where filtering, ranking, and content shaping all happen inside a single controlled system can enforce one consistent trust policy across everything that passes through it. A pipeline stitched together from tools that were never designed with injection containment as a first-class concern can't offer that same guarantee, no matter how carefully it's configured after the fact. The infrastructure a model reads through, in the end, decides far more about its exposure than the model itself ever will.
Sources
- How Prompt Injection Attacks Compromise AI Agents in 2026
- Prompt Injection Attacks on AI Agents: How to Detect and Prevent Them
- Prismata: Confining Cross-Site Prompt Injection in Web Agents
- Fooling AI Agents: Web-Based Indirect Prompt Injection Observed in the Wild
- EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
- Prompt injection still drives most agentic AI security failures in production - Help Net Security


