A preprint posted to arXiv on August 12 by Holzwarth, González-Márquez, and Dmitry Kobak at Ghent University did something simple that almost no one had actually done at scale: they ran the full text of every English biomedical paper archived in PubMed Central for 2024 and 2025 through a more sensitive word-frequency detector of large-language-model writing. The numbers are stark: 77% of English papers archived in 2025 carry detectable traces of LLM-assisted writing, compared with 52% in 2024, and for December 2025 alone the rate climbs to nearly 90% (see https://www.nature.com/articles/d41586-026-02551-z and the Solidot report https://www.solidot.org/story?sid=85173).
What makes 90% feel jarring is that it disagrees with every prior estimate. Kobak himself told Nature his first reaction was "we must have done something wrong." After repeated checking, he conceded the data hold.
Why the numbers are so much higher than earlier studies
For the past few years, the field's best estimates of "AI involvement in academic writing" sat between 10% and 15%. A 2025 Kobak-team study of abstracts put 2024 at 13.5% or higher (https://www.nature.com/articles/d41586-025-02097-6). A 2026 cross-disciplinary study by Kyle Siler estimated 57% for 2025 (PNAS, https://doi.org/10.1073/pnas.2605754123).
The new study's headline numbers are higher for two reasons. First, it analyses full text, not just abstracts. Second, it returns direct estimates instead of lower bounds. Re-running the new method against the 2024 abstract corpus bumps the detection rate from 13.5% to 31% — meaning earlier studies were undercounting 2024 itself by more than half.
An independent signal corroborates this: a 2025 Wiley survey found 71% of researchers self-reported using AI for writing assistance. Given that sensitive behaviors are almost always under-reported, 77% is not just plausible; it may even be conservative.
Not every section is touched equally
The most informative detail in the new study is that AI traces are heavily uneven across sections of a paper. For December 2025:
- Discussion sections: about 78% show AI traces
- Abstracts and introductions: similarly elevated
- Methods and Results: about 58%
This distribution is not accidental. Abstracts, introductions, and discussions are the heaviest "writing work" sections, exactly where rewriting, polishing, and expansion by an LLM feels natural. Methods and Results, by contrast, carry the actual experimental design and data processing — the part the author knows best, and the part where AI assistance has the lowest motive.
That asymmetry is exactly what makes the situation dangerous. The low Methods/Results detection rate does not mean those sections are safe; it means that if LLM contamination reaches them, the consequences are far worse. Kobak flags this directly: LLMs are prone to "hallucinating" data, so any AI痕迹 in Methods/Results carries a non-trivial risk of fabricated or invented data.
Three under-appreciated risks
What makes this worth more than a viral statistic is the three structural things it implies.
First, peer review's detection capability has been overestimated. If 90% of papers pass review without triggering alarm, then reviewers are either unaware or the criteria simply do not cover this dimension.
Second, literature is being slowly biased. Kobak's quote in the Nature piece is worth lifting: "Whatever bias the LLM may have will just suddenly permeate the literature." Unlike outright data fabrication, this permeation is gradual and distributed, which makes it harder to attribute and easier to ignore.
Third, a tacit "use equals comply" consensus is forming. When 71% of researchers self-report use, 77% of papers show traces, and there is no wave of retractions or disputes, "AI-assisted writing" has quietly migrated from "potentially non-compliant" to "default behavior." The rules did not change; the baseline drifted.
How peers are reading it
The Nature piece also carries reactions from independent scholars. Kyle Siler of the University of Toronto — author of the 57% cross-disciplinary estimate — is blunt: "The toothpaste is out of the tube, and it's not going back." Other quoted researchers broadly agree that the high figure may not generalise to every discipline, but that LLM writing is now a default component of research output and is unlikely to reverse.
A useful boundary: PubMed Central is biomedical-only. The preprint's dataset does not cover physics, computer science, mathematics, or engineering. Discipline-by-discipline writing habits and AI-motivation differ substantially; 90% does not automatically extrapolate across fields.
What to actually do now
For researchers, the takeaway is not "ban AI from papers" — that moment has passed. The realistic move is two-track.
- Disclose by default: declare in the submission which sections used AI, which tool, and what degree of modification. Journals and conferences are increasingly codifying this.
- Guardrails on critical sections: Methods and Results, which carry the facts, should either be fully human-written or carry traceable prompts and generation logs. If Discussion uses AI, run a fact-check pass at minimum.
For AI-tool builders, this is a clear product signal. A research-grade AI writing assistant that only does "polish and rewrite" will increasingly look suspicious. One that bakes in "immutable provenance for every AI-touched paragraph" by default can carve out position before the new rules arrive.
For research policy and ethics bodies, the signal is that more concrete standards are overdue. Mark AI-touched sections in the PDF or in metadata, require tool/version disclosure, define what counts as authorship-grade AI contribution. Editorial consortia, funders, and ethics committees all need firmer positions inside the next 12 months — otherwise the rules will only fall further behind the actual behavior.
Personal take
The 77% number will travel. What matters more is that it forces a new question onto every paper, every reviewer, and every future AI training corpus: who actually wrote this paragraph, and who actually modified it?
That is not a question about whether AI can write a paper. It is about whether academia can record what happened, honestly, every time.