Est.

Passage Summarization vs Direct Injection in RAG Prompts

Two strategies control what happens to retrieved content before it reaches the model.

Staff Writer · · 11 min read
Cover illustration for “Passage Summarization vs Direct Injection in RAG Prompts”
Context Engineering · October 7, 2026 · 11 min read · 2,375 words

An engineer building a retrieval-augmented generation pipeline eventually has to decide what happens to a retrieved passage the moment before it reaches the model: pass it through whole, or compress it first. That single decision determines fidelity, latency, token cost, and security exposure across the entire system, and most teams never make it deliberately. They inherit a default, usually "inject everything the retriever returns," and build the rest of the pipeline around that assumption without testing whether it holds.

Two strategies are actually on the table. Direct injection places raw or lightly chunked passages verbatim into the prompt. Passage summarization compresses retrieved content before the model ever sees it. These are not interchangeable implementations of the same idea. Each one moves the location where failure happens: direct injection pushes risk onto the retriever and whatever it hands over, while summarization pushes risk onto the compression step itself, which becomes a second place where meaning can go wrong before generation even starts. The choice also doesn't stand alone. It interacts with retriever quality, the model's context capacity, the latency budget a product can tolerate, and how much the source material can be trusted; evaluating it as an engineering tradeoff is what settles it, not convention.

What direct injection does inside the context window

Direct injection keeps the model reasoning over text nobody has touched. A citation traces straight back to its source, and no intermediate inference step sits between the retrieved passage and the answer the model produces, which removes one whole category of error that summarization has to account for. That is the appeal, and it is a real one: fidelity is structurally guaranteed rather than something the pipeline has to verify after the fact.

The cost is token-linear, and it compounds quickly. Every passage placed in the prompt consumes context budget, and a model's effective reasoning gets worse as that window fills with redundant or loosely related material. Research presented at ICLR 2025 found that for long-context LLMs, output quality improves as more passages are added, up to a point, and then declines as additional passages accumulate beyond that. More retrieval is not the same as better output. The same research found that "hard negatives," passages that read as semantically close to the query but are factually beside the point, cause part of that decline, and that stronger retrievers can make the problem worse rather than better, because a sharper retriever surfaces more passages that sound plausible without actually answering anything.

Source quality determines how bad this gets. Raw web content arrives full of navigation menus, ad text, and boilerplate, and if that content enters the context unfiltered, the model has to reason around noise it was never built to screen out on its own. A traditional SERP API hands back ranked URLs and snippets that still need fetching, parsing, and cleaning before they're usable in a prompt, and every one of those steps is a place where something can break or degrade. An API built to return pre-extracted, semantically structured text instead collapses that whole pipeline into a single clean step, which is the gap an AI-native search API like Seltz is built to close: the retrieval layer, not the generator, becomes responsible for controlling noise before anything reaches the context window. That distinction matters because it changes what "direct injection" even means in practice. Injecting raw scraped HTML and injecting pre-structured, filtered text are two different risk profiles wearing the same name.

What summarization does

Summarization earns its place by solving the problem direct injection cannot: it compresses retrieved content to fit a tighter context budget, strips out redundant or overlapping passages, and in multi-step reasoning tasks, it can surface cross-document synthesis that raw chunks tend to obscure by burying the connection in volume.

The risk it introduces is faithfulness loss, and it deserves more attention than it usually gets. A summarizer is itself a second LLM inference. It can fabricate detail, smooth over a qualification the source author included deliberately, or shift the meaning of a claim just enough to matter, and the generator downstream has no way to catch any of it because it never sees the original text. The standard mitigation is groundedness evaluation: an LLM judge checks each claim in the summary against the source chunks it was drawn from. That step works, but it adds latency and another layer of complexity to a pipeline that already has multiple stages where something can go wrong.

Summarization comes in two modes that behave differently under pressure. Extractive summarization selects spans verbatim from the source, which preserves exact wording and suits legal or compliance contexts where the precise phrase is the thing that matters. Abstractive summarization rewrites the content in new sentences, which reads better but carries a higher hallucination risk because the model generates language from its own inference. The production pattern that holds up best is hybrid: extract the key spans first to anchor the facts, rewrite around them for readability, then score the result for groundedness before it goes anywhere near the generator.

Compression is not free even when the deployment looks nothing like an edge device. Research on edge RAG systems found that on constrained hardware, the compressor runs on the same system-on-chip as the generation step itself, so the energy and latency cost of compressing the context can cancel out the savings compression was supposed to produce. A fixed, static compression ratio applied without regard to what the hardware is actually doing at that moment is a bad default, and the underlying principle holds for any latency-sensitive pipeline, including beyond edge hardware: compression has a cost, and that cost has to be measured against what it saves.

How the tradeoff shifts across four pipeline variables

Diagram: How Four Variables Shift the Injection vs. Summarization Decision. Visualizes: Show how the same two strategies — direct injection and passage summarization — are favored or disfavored across four concrete pipeline variables: (1) Context…

The right choice between injection and summarization cannot be settled in the abstract. It changes depending on four conditions, and treating the decision as fixed across all of them is where most pipelines go wrong.

Context budget is the first. When the model's window is generous relative to the size of the retrieved set, direct injection preserves fidelity at a marginal cost that's easy to absorb. When the retrieved set is large relative to that window, summarization or contextual compression becomes a hard requirement rather than an option, because the alternative is truncation or the quality degradation the ICLR findings describe.

Retriever quality is the second, and it's where the upstream pipeline decides how much work the downstream steps have to do. A clean, LLM-native retrieval layer, one that returns pre-extracted, filtered, structured text, makes direct injection viable because the material entering the context is already low-noise. Raw SERP output or unfiltered scrapes force summarization or compression into the role of mandatory cleanup. That means the injection decision and the retrieval infrastructure decision are not separable. A team that treats them as independent choices is quietly accepting a weaker default on one of them without realizing it. The value of the injection strategy depends on whether the retriever itself was engineered to surface clean, structured content or only ranked URLs and raw snippets that push noise downstream. A full pipeline that crawls, extracts, indexes, ranks, and structures web content before it reaches the prompt, which is the approach Seltz takes, changes that tradeoff directly by reducing the noise burden that makes summarization necessary.

Query type is the third variable. Research comparing long-context and retrieval-augmented approaches found that summarization-based retrieval performs comparably to long-context processing on many benchmarks, while chunk-based retrieval lags behind both, and that retrieval-augmented generation holds a real advantage specifically on dialogue-based and general question queries.

Deployment environment is the fourth. Latency-constrained systems, edge hardware and real-time agents in particular, absorb a compression cost that cloud deployments can tolerate without much friction. Adaptive compression research shows that intermediate compression rates can cut GPU energy use by up to 53.2% with negligible quality loss, but capturing that gain requires runtime telemetry feeding the compression rate in real time, not a rate fixed once at deployment and left alone.

Direct injection of web-retrieved content as a distinct security surface

The architecture of an LLM creates the vulnerability here, not a flaw in any particular implementation. No runtime boundary separates an instruction from a piece of data, so the model treats a sentence inside a retrieved webpage as a command to be followed in the same way it treats content to be reasoned about.

That gap is what indirect prompt injection exploits. An attacker embeds instructions inside content an agent is likely to retrieve, a webpage, a document, an email, a code comment, and the model follows those instructions as though they came from the system or the user, because nothing in its input distinguishes one source from another. This isn't a theoretical concern: research on PoisonedRAG, presented at USENIX Security 2025, demonstrated that a small number of carefully crafted documents can reliably manipulate a RAG system's output for targeted queries.

Summarization changes where this risk sits without removing it. A summarization step can carry forward hidden instructions embedded in source content, white-on-white text, HTML comments, invisible Unicode characters, just as directly as an unfiltered injection would, a pattern that has shown up in disclosed cases of RAG poisoning and scraping-based injection attacks. Compressing the content doesn't filter out an instruction smuggled inside it; it just changes which component ends up executing it.

The most concrete structural response to this problem is Highlight & Summarize, a design pattern from Microsoft Security Response Center. H&S splits the pipeline into two separate components: a highlighter that takes the user's question and pulls relevant passages from retrieved documents, and a summarizer that takes those highlighted passages and produces an answer, without ever seeing the user's original question. Isolating the summarizer from the query reduces an attacker's ability to steer the output, because there's no question text left for a hidden instruction to hijack or respond to. On QA benchmarks including RepliQA and BioASQ, H&S produced responses comparable to or better than a standard RAG pipeline. The security gain didn't come at the cost of answer quality. Infrastructure choices upstream matter here too: pipelines built on controlled, enterprise-grade retrieval sources with access-scoped document filtering, the kind of security trimming Azure AI Search applies at retrieval time, start with a smaller attack surface than pipelines that inject open-web content directly into the prompt.

The long-context escape valve

Long-context LLMs reduce the tradeoff by exchanging one set of constraints for another.

The ICLR 2025 research addresses this directly. Increasing the number of retrieved passages fed into a long-context window raises the likelihood that irrelevant or noisy passages get pulled in along with the useful ones, and attention costs rise as context length grows, since attention scales quadratically. A longer context is materially more expensive to process per inference, not a free upgrade.

A long-context model's knowledge cutoff doesn't go away either. Anything not explicitly present in the injected passages still gets answered from training data. Live retrieval stays necessary for time-sensitive, factual, or domain-specific queries no matter how large the context window gets.

There's also a documented attention problem inside the window itself. Long-context models attend more reliably to content near the beginning and end of a sequence, and material placed in the middle of a long injection block is more likely to be underweighted or effectively ignored. The ICLR research identifies retrieval reordering, placing the highest-scored documents at the start and end of the injected block, as a practical mitigation for this "lost-in-the-middle" effect. It helps, but it doesn't resolve the underlying issue; it just works around where the model's attention naturally lands. Long context expands how much direct injection a system can get away with, but it means the model must actually attend to the parts of the content that matter, not merely fit them in the window.

A practical decision framework for choosing between injection and summarization

Selecting an injection strategy means evaluating five conditions systematically: context budget, source cleanliness, query type, latency constraints, and security requirements. Picking one strategy as a default and optimizing everything else around it skips the actual decision.

Source cleanliness comes first. If retrieved content arrives pre-extracted, semantically structured, and already filtered, the kind of output an LLM-native search API such as Seltz is designed to produce, direct injection is the lower-risk place to start. It preserves fidelity, skips the faithfulness-loss risk a summarization step introduces, and keeps the pipeline simpler with one fewer stage that can fail. An end-to-end pipeline that selects, filters, ranks, and shapes web content specifically for LLM reasoning, which is the approach behind Seltz, means boilerplate, navigation clutter, and irrelevant material have already been stripped out before anything reaches the context window. That is the infrastructure condition that makes direct injection dependable rather than a gamble on what the retriever happened to surface.

Compression becomes necessary, not optional, once the retrieved set regularly exceeds what the model can reason over effectively. Use adaptive compression rates driven by latency and energy telemetry at that point, adjusting continuously as conditions change.

Query type should guide the choice as well: dialogue and single-hop factual questions favor direct injection, while multi-step reasoning and cross-document synthesis tend to favor summarization-based retrieval.

Faithfulness needs its own evaluation layer, not an assumption of correctness. Any pipeline with a summarization step needs groundedness scoring, checking each claim in the summary against the source chunks, as a required component. A pipeline running summarization without that check is shipping compressions nobody has verified.

Security has to be treated as a pipeline constraint from the start. If the retrieval source includes open-web content or documents supplied by users, indirect prompt injection is a live risk for any system using direct injection, and the Highlight & Summarize pattern, a highlighter paired with a summarizer isolated from the user's query, deserves serious evaluation for any deployment where trust in the source material can't be assumed.

Finally, test on the actual documents and queries the system will face in production. Public benchmarks measure performance on general corpora, and a pipeline's behavior on a 200-page contract, a regulatory filing, or a real-time news feed can diverge sharply from benchmark results. Run the candidate strategies against a fixed evaluation set built from the real pipeline before committing to either one.

Sources

  1. LONG-CONTEXT LLMS MEET RAG
  2. Highlight & Summarize: RAG without the jailbreaks
  3. From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG
  4. Long Context vs. RAG for LLMs: An Evaluation and Revisits

More in Context Engineering