Deduplication Strategies for Web-Scale AI Training and Retrieval Corpora
Layered deduplication pipelines catch different types of web redundancy at scale.

The web repeats itself at a scale that defeats any simple rule like "remove the obvious copies and move on." Modern AI training and retrieval corpora are built from crawls so large and so structurally redundant that naive deduplication leaves most of the problem untouched. The repetition comes from several distinct sources at once: the same site gets crawled multiple times, a single wire article gets syndicated across dozens of domains, boilerplate like nav bars, copyright notices, and cookie banners repeats on every page of a given site, and near-copy SEO placeholder text spreads across huge numbers of unrelated domains. Each of these mechanisms produces a different shape of duplicate, and none of them responds to the same detection method.
Scale turns this from an annoyance into a structural risk. The Scale Dependent Data Duplication paper frames the trajectory from Llama 1, trained on roughly 1 trillion tokens, to Llama 4, trained on up to 40 trillion tokens by that paper's account, to show how far token volumes have grown. At that volume, if even a small percentage of data is duplicated, the corpus loses a large absolute number of distinct training examples. Duplicate supervision does more than waste compute: it narrows the effective diversity of what a model sees and raises the risk of overfitting to repeated patterns. That is a measurable cost, not a cosmetic one, and it sets the terms for everything that follows: deduplication has to be built as a layered pipeline, with each layer catching a kind of redundancy the others miss, rather than treated as a single filtering pass run once before training begins.
Exact and near-duplicate deduplication: what hashing and MinHash can and cannot do
The first two layers of any serious deduplication pipeline handle exact copies and near copies, and together they form the floor everything else sits on. Exact deduplication relies on content hashing, typically something like SHA-256 computed at ingestion, so that any byte-for-byte identical document can be flagged and dropped cheaply. The Nucleus-Image data pipeline used exact-hash deduplication as one stage among several in its multi-stage filtering process ahead of training on a large-scale image-text corpus, which shows how standard this layer has become even outside pure text applications. In practice, exact-hash deduplication tends to remove a modest share of records from a typical web-crawl subset: a meaningful cut, but nowhere near the bulk of what a crawl actually repeats.
Near-duplicate deduplication picks up where exact matching stops. Building on MinHash or Locality Sensitive Hashing involves computing a compact sketch of each document, then estimating similarity between sketches. MinHash is particularly good at catching templated content: software licenses that differ only in the name of the licensed entity, or placeholder SEO text copied with minor edits across hundreds of sites. The Falcon RefinedWeb team ran MinHash over 5-grams at a scale few prior pipelines had attempted, processing an enormous CommonCrawl snapshot under both heavy filtering and strict deduplication; the public release holds 600 billion tokens. RefinedWeb's broader finding, that web data cleaned and deduplicated this way can outperform models trained on curated corpora like The Pile, is the clearest evidence available that this baseline layer carries real weight in determining model quality.
Both methods share a limit that matters more as corpora grow. Hashing and MinHash operate on surface form: the actual sequence of characters or tokens on the page. Neither can recognize a translation of the same article, a paraphrase that rewrites every sentence while keeping the same content, or two independently written pieces that happen to cover identical ground. These documents pass every check this layer can run, so the redundancy they carry moves forward into training untouched.
Subdocument deduplication: why document-level detection is too coarse
A third layer exists because duplication often happens inside documents, not just between them. A page can share a template, a navigation bar, a copyright footer, a quoted passage, or a copied code fragment with thousands of other pages while still containing substantial unique content of its own. Document-level deduplication, whether exact or fuzzy, treats the document as the unit of comparison and so cannot act on redundancy that occupies only part of it.
Lowering the similarity threshold to catch more of this local overlap creates its own problem: documents with only a small amount of shared text alongside large amounts of unique material start getting discarded wholesale, and document-level methods have no way to resolve that tension on their own. The Tencent Hunyuan team built a framework around a different design principle: retain more copies of repetitions that are short or infrequent, and prune more aggressively when repetitions are long or frequent. Copy retention, in this framework, is not a flat yes-or-no decision but a quantity calibrated to how often and how extensively something repeats. Suffix-array-based methods offer an alternative route to the same goal: they identify variable-length exact repetitions without needing a predefined segmentation scheme, but building a single suffix array across an entire corpus costs too much at scale. Real implementations shard the corpus and match within each shard. Duplicates spanning two shards go undetected, and the results depend heavily on how the sharding was done. The Hunyuan hashing-based method sidesteps this by mapping naturally onto distributed key-value aggregation, so identical units get counted globally no matter how the data happens to be partitioned across machines.
OLMo 3, built by AI2, applied a three-stage pipeline covering exact, fuzzy (MinHash), and substring-level removal, using a purpose-built toolkit called Duplodocus for the exact and MinHash stages and a tool called bsade for the substring and suffix-array stage, both written in native Rust for large-scale, high-performance execution. The result cut the web corpus by 75% in document count, and total text bytes also fell substantially. Separately, experiments on FineWeb-Edu and a code-containing web corpus found that models trained on data processed with the Hunyuan frequency- and length-aware method achieved the best overall performance among the settings tested, confirming that explicit control over copy retention produces measurable gains over fixed, one-size-fits-all rules.
Scale-dependent semantic duplication: the layer that standard pipelines miss entirely
A fourth layer of redundancy exists that none of the first three can touch, because it has nothing to do with surface form. Two documents can pass every exact, fuzzy, and subdocument check while still delivering the same training signal to the model, and this kind of redundancy gets worse, not better, as the model grows more capable.
Researchers affiliated with Stanford and EPFL published the Scale Dependent Data Duplication paper, and in it they name two separate mechanisms behind this. First, as a model's capability increases, the cross-entropy loss gradients it produces for semantically equivalent documents become more aligned with each other. Small models produce gradients that mostly track surface-level similarity, like shared tokens, but large models produce gradients that track meaning. A translation and its original, in other words, start to look like near-duplicates to a large model's training dynamics even though they share almost no tokens. Second, as corpus size grows into the hundreds of billions of tokens, the nearest-neighbor cosine similarities between documents deviate sharply from the isotropic power-law pattern that holds at moderate corpus sizes, a sign of accelerating semantic collisions as more documents pack the same conceptual space. To show this, the paper embeds 192 million FineWeb-Edu-Dedup documents with a compact embedding model, and it finds that the scaling laws predicting model behavior at moderate corpus sizes stop holding once corpora get large enough, so the predictability researchers normally rely on starts to break down. Controlled pretraining experiments back this up directly: sampling with replacement from a finite pool of documents causes only mild degradation in small models, but the loss penalty grows rapidly as model size increases, undercutting naive extrapolation from small-scale results. The paper goes further still: it derives explicit scaling laws that quantify how far actual performance deviates from expectation given a corpus's limited semantic uniqueness, so you can estimate how diverse a corpus really is before you commit compute to training on it.
Meta FAIR's SemDeDup put the embedding-based approach into practice ahead of this theoretical account, using embeddings from pre-trained models to find and remove semantically similar data pairs that are not textually identical. It showed efficiency gains on both a largely uncurated dataset (LAION) and a partially curated one (C4), and it preserved downstream model performance too, so the method works outside a purely theoretical setting. The two findings compound in a way surface-level methods cannot address: capable models get trained on ever-larger corpora, and larger corpora produce more semantic collisions, so the exact conditions that make semantic duplication dangerous are the same conditions driving the field's push toward bigger models and bigger datasets.
The deduplication-quality trade-off: why aggressive removal is not always optimal
Having established that all four layers matter, a complication follows: more deduplication is not always better. Duplicate count is, on its own, a weak signal of quality, and it can point in either direction. A document that appears many times in a corpus might be a low-value SEO page copied across spam domains, or it might be a canonical reference document or a widely cited code pattern that earns its repetition through genuine usefulness. Stripping all repeated instances indiscriminately removes real signal along with noise.
The damage from repetition is not linear starting from the first copy. Repeating a document a handful of times carries a different cost profile than repeating it hundreds of times, and returns from additional copies diminish quickly after the first few. That argues for a retention policy that keeps one or a small number of copies. The Hunyuan frequency- and length-aware retention policy turns this principle into a concrete mechanism: short or infrequent repetitions get a larger copy budget, while long or frequent repetitions get pruned harder. This gives the non-monotonic relationship between deduplication and quality an operational rule. The Scale Dependent Data Duplication paper's scaling-law formulation adds a second tool on top of this, letting practitioners estimate a corpus's effective diversity and predict how far performance will deviate from expectation, turning what used to be a judgment call into something closer to a measurable engineering parameter.
A natural objection follows: given a non-monotonic relationship, a practitioner might prefer to keep more data to avoid cutting too much. The answer comes down to opportunity cost. Compute at web scale is finite, and every token spent training on a near-duplicate is a token not spent on a genuinely novel example. At corpora running toward 40 trillion tokens, that opportunity cost becomes significant enough to shape how much a given training run can actually learn.
How deduplication requirements diverge between training corpora and retrieval corpora
The same four-layer framework applies to retrieval corpora, but the stakes take a different shape. In pretraining, redundancy degrades gradient quality and wastes compute. In retrieval, redundancy degrades retrieval accuracy, narrows the diversity of what gets returned to a user, and, in enterprise deployments, can produce grounded answers that contradict each other even when every individual system component is working as designed.
MS MARCO V2's document collection contains substantial near-duplicate overlap, and left in place, that overlap drags down retrieval accuracy and shrinks the variety of documents a retrieval-augmented system actually surfaces. The TREC RAG Track's MS MARCO V2.1 addressed this directly with a two-stage approach: building equivalence classes of documents using LSH with MinHash and 9-gram shingles, essentially the same fuzzy-deduplication tooling used in training pipelines, applied instead to a retrieval index. A 2026 empirical study of a clean academic retrieval corpus found something different: byte-exact deduplication at query time cut the corpus by only 0.16%, which suggests that deduplication's payoff in retrieval depends heavily on how messy the underlying corpus already is. A corpus like MS MARCO V2 gains enormously from aggressive deduplication, but a corpus that starts out clean has far less to gain.
Enterprise retrieval systems face a failure mode that pretraining does not: a knowledge base holding multiple versions of the same policy document can produce contradictory grounded answers even when the retrieval system is functioning exactly as intended. The fault sits in corpus hygiene, not in model capability. A related failure comes from staleness rather than duplication in the usual sense: an index that isn't refreshed after the source information changes keeps grounding answers in content that used to be accurate and no longer is. Removing superseded versions, a form of temporal deduplication, matters here just as much as removing redundant copies. Semantic deduplication in a retrieval context has a different goal than it does in pretraining: the aim is to make sure the passages handed to a language model's context window each carry distinct informational content, not that they merely use different words. For any RAG pipeline that ingests live web content or an enterprise knowledge base that gets updated over time, deduplication cannot be a single preprocessing pass run once before deployment. It has to run continuously, because the underlying corpus keeps changing after the system goes live.
Federated and continuous deduplication: when preprocessing cannot happen before training
Federated learning breaks the standard "deduplicate before training" model at a structural level. Data in a federated setting cannot be centralized. A duplicate sitting on one client's device and an identical document sitting on another client's device are invisible to any deduplication process that only looks at local data. The conventional response, running a global deduplication pass before training starts, carries its own costs in this setting: the serialized sequence of preprocessing followed by training is expensive to make fault-tolerant, and it offers no good way to handle clients joining the system dynamically after training has already begun.
Researchers at Nankai University proposed Deduplication-while-Training, or DwT, as a structural response to this problem. DwT turns cross-client deduplication from a one-time, globally synchronized preprocessing operation into a continuous online service, one that manages state, allows concurrent claiming of data across clients, and handles failure recovery as an ongoing concern. Measured against state-of-the-art schemes, DwT cuts failure-recovery time overhead by up to 93.04%. That figure marks deduplication's shift from something done to a corpus once, before training starts, toward something built into the infrastructure a training system runs on for as long as it keeps running.
Sources
- Nucleus-Image: Sparse MoE for Image Generation
- Deduplication-while-Training: A Resilient Paradigm for Privacy-Preserving Cross-Client Deduplication in Federated Learning
- Scale Dependent Data Duplication
- Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining
- The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
- Olmo 3


