For the past few years, the question of how much of the scientific literature is written by AI has been stuck at the guessing stage: detectors are unreliable, and surveys depend on self-reporting. A preprint submitted to arXiv on August 11 (arXiv:2608.10715) now offers the hardest answer yet: by the end of 2025, 89% of open-access biomedical papers archived in PubMed Central showed an excess of LLM-associated vocabulary. Nature's coverage put "Staggering 90%" in the headline and noted the figure is far higher than previous estimates of LLM use.

Word-frequency drift, not per-paper detection

The work comes from Lena Holzwarth, Rita González-Márquez, and Dmitry Kobak. Rather than building yet another per-document "AI detector," they propose an unbiased estimation method based on changing word frequencies: treat the whole corpus as an epidemiological sample, measure the systematic drift in word frequencies of the LLM era, and estimate the share of text altered by LLMs at the distribution level. The authors state plainly in the abstract that despite recent progress, no existing method could produce reliable estimates — trustworthy measurement has to come before policy, and that is the starting point of the whole work.

The numbers on three axes

  • Time: across English-language papers published in 2025, LLM usage reached 77%, versus 52% in 2024; papers published in December 2025 showed traces in nearly nine out of ten cases.
  • Section: paragraph-level usage in the Discussion section (68%) is about twice that of the Methods section (32%); yet even inside Methods, overall prevalence exceeds 50%. Abstracts, introductions and discussions show traces more often than methods and results.
  • Self-report: a 2025 survey found 71% of researchers admit using AI-assisted writing, with actual usage likely higher — consistent with the word-frequency estimates.

The real alarm: Methods past the halfway mark

Heavy LLM rewriting of Discussion sections surprises few people — that is where authors "perform." What stings is the Methods section: its entire purpose is to record precisely how the experiment was done, and it is what reproducibility rests on. When paragraphs that describe facts start being rewritten by a model, the reader faces not a style question but a distance question — how far this record sits from what was actually done.

For journals and review systems, this study turns a vague worry into a measurable baseline. The authors also note the double edge: LLM-assisted writing lowers language barriers — a real benefit for non-native English authors — while the risks of misconduct and fraud grow in step. The predictable next step is tighter disclosure requirements and stricter writing norms for Methods sections — once numbers like "more than half of Methods paragraphs" exist, the cost of looking away only goes up.

References: arXiv:2608.10715 · Nature coverage