What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Adding AI-generated web text does not have one fixed effect on language-model training. A September 2026 preprint reports that its value depends on how much human text a model has already seen, how much AI text is added, and whether the model is evaluated on human or AI-generated text. For some data-starved models, AI text initially helped held-out human-text performance before the benefit saturated and reversed. With larger human-text budgets, the authors report that adding AI text raised human-text loss sooner, while repeating human text continued to lower it.
This is a conditional result about unlabeled, AI-generated text found in web corpora—not proof that all synthetic training data is harmful, that models inevitably collapse, or that the original Chinchilla result is wrong.
What is the AI data satiation point?
“AI data satiation” describes a point at which additional AI-generated web tokens stop improving a model’s performance on human text and may make it worse. It is not a single universal token count. In the experiments reported by Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, and Bradley Emi, the point varied with the model’s existing human-text budget, model size, and ratio of added AI to human tokens.
The authors pretrained 800 language models across different added-AI-token ratios, then fitted scaling laws to held-out loss on human and AI text. Their proposed law has separate terms for benefit and harm, allowing the modeled value of an AI token to change sign as more is added. These are the authors’ findings from their tested models and evaluation sets; they do not establish the same threshold for every architecture, dataset, or training recipe.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Why the human-text budget matters
In a data-starved setting, AI text could provide an initial benefit to held-out human-text loss. As the added AI-to-human ratio grew, that benefit saturated and could turn into harm. When models had larger human-text budgets, the reported human-text loss rose with AI additions sooner, while repeating available human text continued to lower it.
So “more data” is not a sufficient description of the choice. The composition of the data and the model’s existing exposure matter. The results do not imply that AI text has no value in every regime; they imply that its marginal value depends on the training and evaluation conditions.
Rank #2
What does “wild” AI text mean?
Here, “wild” means unlabeled AI-generated text that has entered ordinary web material and is collected as part of a web corpus. It is different from a dataset deliberately generated, filtered, and curated for a particular training task. It is also different from experiments that repeatedly train models on outputs from earlier generations of models to study recursive model collapse.
The distinction matters because those setups ask different questions. Russell and colleagues study the effect of adding AI-generated web text to pretraining data and measure performance on held-out text. Their result should not be generalized to every use of synthetic examples, such as carefully designed training data, without separate evidence.
How much web text did the study label as AI-generated?
In Russell and colleagues’ sampled web crawl, after FineWeb quality filters were applied, Pangram labeled 27.5% of tokens as AI-generated in the June 2026 sample and 31.1% in the August 2026 sample. These percentages describe tokens in those particular filtered samples and the detector’s labels. They are not verified estimates of the share of all internet content, nor of the share of every model’s training data.
Detection labels are a measurement aid, not ground truth. The study also reports releasing WildAI, an 83-billion-token corpus with AI, topic, and format labels, alongside models and code. The authors’ reported corpus size does not by itself establish that the dataset is currently accessible under terms suitable for reuse.
Does this disprove the Chinchilla scaling law?
No. Hoffmann and colleagues’ 2022 Chinchilla work studied compute-optimal training under its setup. It trained more than 400 language models, ranging from 70 million to over 16 billion parameters, on 5 billion to 500 billion tokens. Its central result was that model size and training-token count should scale together: for every doubling of model size, double the training tokens. The Chinchilla test model had 70 billion parameters and used four times Gopher’s training data at the same compute budget.
The 2026 preprint addresses a narrower issue: how adding unlabeled AI-generated web text changes the relationship between training data and held-out loss. The authors say their proposed law reduces to Chinchilla when no AI text is added. In their tests, fitting on smaller models predicted held-out human-text loss for models up to 3.6 times larger with 41% lower error than the best existing law over the tested AI ratios. That is a reported benchmark within the paper’s experiments, not proof of performance on all future frontier models.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
Scaling-law conclusions also depend on the objective. For example, a separate inference-aware analysis by Sardana and colleagues argues that ordinary token-to-parameter training ratios can overstate the impact of extra tokens at extreme ratios. That does not negate Chinchilla; it reflects a different consideration in choosing model and data scales.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Why the evaluation target changes the answer
The same AI text can have different value depending on what the model is expected to handle. In the reported experiments, AI-generated text remained useful when the evaluation target was AI-generated text, even where additional AI text harmed performance on human-text evaluation. A combined validation set can conceal that difference if its human and AI portions are not reported separately.
| Evaluation target | What the reported result means |
|---|---|
| Human text | Added wild AI text may help initially in data-starved settings, but the benefit can saturate and reverse; harm appeared sooner with larger human-text budgets. |
| AI-generated text | The authors report that AI text remains useful when AI text itself is the target. |
| Mixed human and AI text | A single aggregate loss can obscure a change on the human-text slice. Report each slice separately to make the trade-off visible. |
What should model trainers do with the result?
The authors recommend “filtering AI text when the target is human text, repeating human text before expanding the training dataset with AI-generated web text, and reporting validation loss on human and AI text separately”. This is their recommendation based on the preprint, not a universal standards-body rule. Its most important practical qualification is to define the intended target distribution: AI text may still be useful when AI-generated text is part of that target.
- Measure the data mix. Track human and AI text separately where possible, and document how AI labels were produced rather than treating detector output as certain.
- Match validation to the intended use. Keep human-text and AI-text validation losses separate; an aggregate score alone may hide a divergence.
- Compare additions against repetition. The authors advise repeating available human text before expanding with wild AI web text when human text is the target.
- Interpret a threshold in context. A satiation point depends on the human-token budget, model size, added AI-to-human ratio, and evaluation set—not simply on a corpus-wide AI percentage.
Will AI run out of human training data?
A 2024 ICML position paper by Pablo Villalobos and coauthors conditionally forecast that training datasets could approach its estimate of the available stock of public human-generated text between 2026 and 2032 if then-current trends continued, or earlier under overtraining. This is a model-based forecast, not a measured exhaustion date or a prediction that model collapse must follow. The paper discusses synthetic data, transfer from data-rich domains, and improved data efficiency as possible responses.
That supply question is related to, but distinct from, the satiation result. A limited supply of public human text may motivate alternatives; it does not show that every alternative works equally well for every target. A separate 2025 SynthLLM preprint reports a performance plateau near 300 billion tokens for its synthetic-data framework and experiments. Because it studies a different setup, it is not direct confirmation of the wild-web-text findings.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

