[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-pubmed-77-percent-llm-writing-2025":3,"news-related-144fa9dc-de03-4972-a695-3d392f334772":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"144fa9dc-de03-4972-a695-3d392f334772","PubMed 中央库研究:2025 年生物医学论文 77% 有 LLM 写作痕迹","arXiv 预印本统计:2025 年 PubMed Central 英文生物医学论文 77% 含 LLM 写作痕迹,12 月单月升至约 90%;新词频法把 2024 年摘要检出率从 13.5% 推到 31%,与 71% 研究人员自报使用 AI 辅助写作的调查一致。","arXiv 在 8 月 12 日挂出了一篇预印本,作者是 Holzwarth、González-Márquez 与比利时根特大学的 Dmitry Kobak。他们做了一件听起来朴素但很少有人真正做过的事:把 PubMed Central 数据库里 2024 年和 2025 年全年收录的英文生物医学论文全文跑了一遍,用一套更敏感的词频方法去看哪些论文里藏了 LLM 写作的指纹。结论非常硬——2025 年全年被收录的英文论文里,有 77% 能检出大模型参与写作的痕迹,2024 年这个数字是 52%,而 2025 年 12 月单月,这一比例直接飙到了接近 90%(参见 https:\u002F\u002Fwww.nature.com\u002Farticles\u002Fd41586-026-02551-z 与 Solidot 报道 https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85173)。\n\n把这件事拆开看,90% 这个数字之所以让人觉得不对劲,是因为它和此前所有估计都对不上。Kobak 本人在接受 Nature 采访时说,他自己第一反应也是「我们肯定哪里算错了」。但反复验证之后,他承认数据站得住脚。\n\n## 数字为什么会比之前所有研究都高\n\n过去几年,关于「学术论文里有多少 AI 写的」这个问题的研究,数字一直压在 10%-15% 这个区间。2025 年 Kobak 团队自己用摘要做的一项研究,给出的 2024 年数据是至少 13.5%(https:\u002F\u002Fwww.nature.com\u002Farticles\u002Fd41586-025-02097-6);2026 年另一项跨学科研究给出的 2025 年估计是 57%(Siler, PNAS, https:\u002F\u002Fdoi.org\u002F10.1073\u002Fpnas.2605754123)。\n\n新研究之所以数字高得多,核心原因是方法更敏感——它分析全文而不是摘要,且不再只给「AI 使用率的下限」,而是给出直接估计。同样的方法回头套到 2024 年数据上,把摘要检出率从 13.5% 推到了 31%。也就是说,不是 2025 年突然变多,而是过去我们对 2024 年本身就低估了一半以上。\n\n另一个独立佐证来自 Wiley 在 2025 年做的研究者自报调查:71% 的研究人员承认自己在用 AI 辅助写作。如果实际使用率普遍高于自报水平——这是几乎所有敏感行为调查的通用规律——那 77% 这个数字就不仅不离谱,反而可能还有点保守。\n\n## 不是所有段落都被 AI 染指\n\n研究最有信息量的细节,是不同段落的 LLM 痕迹分布非常不均匀。在 2025 年 12 月的论文里:\n\n- 讨论部分(Discussion):约 78% 检出 AI 痕迹\n- 引言和摘要:同样明显偏高\n- 方法(Methods)和结果(Results)部分:约 58%\n\n这个分布不是偶然。摘要、引言、讨论是「文字工作」最重的部分,也是 LLM 改写、润色、扩写最自然的切入点。而方法和结果往往涉及具体实验设计、数据处理流程,写作者自己最清楚要写什么,AI 介入的动机也最低。\n\n但这恰恰是最让人担心的部分。Methods\u002FResults 的低检出率不代表「这块安全」——而是这块如果一旦被 LLM 染指,后果会严重得多。Kobak 自己点名了这一点:LLM 在数据呈现上有「幻觉」的倾向,任何进入 Methods\u002FResults 的 AI 痕迹都可能意味着伪造或杜撰的数据风险。\n\n## 三个被低估的风险\n\n把这件事放进更大的语境里,真正值得警觉的不是「90%」这个数字本身,而是它隐含的三件事。\n\n第一,同行评议的识别能力被高估了。如果 90% 的论文都过得了同行评议,那意味着评审流程对 AI 写作几乎没有感知。要么评审者没意识到、要么评审标准根本没有覆盖这一项。\n\n第二,文献的「偏见渗透」。Kobak 在 Nature 报道里说了一句值得引用的话:「如果 LLM 自身带有某种偏见,这种偏见会随着 AI 辅助的引言和讨论,悄悄渗透进整个学科的文献里。」和直接伪造数据不同,这种渗透是渐进的、分布式的,反而更难追责。\n\n第三,「使用即合规」的隐性共识正在形成。当 71% 的研究者自报使用、77% 的论文检出痕迹、却没有系统性撤稿或撤稿争议时,「使用 AI 写作」已经从「可能违规」变成「默认行为」。规则没变,但行为基线整体漂移了。\n\n## 同行怎么看\n\nNature 的报道里同时引用了几位独立学者的反应。多伦多大学的 Kyle Siler(那项 57% 估计的作者)说了一句很直白的话:「牙膏已经从管子里挤出来了,不会再回去。」其他学者也基本认可:高数字可能不适用于所有学科,但 LLM 写作已经成为科研产出的默认组成部分之一,这一点很难再翻盘。\n\n需要注意的边界:PubMed Central 是生物医学领域专门的数据库,这篇预印本的数据集不覆盖物理学、计算机、数学等学科。不同学科的写作习惯和 AI 使用动机差异很大,90% 不能直接外推到所有学科。\n\n## 现在应该做什么\n\n对科研工作者来说,这件事给出的实际信号不是「禁止 AI 写论文」——那个时点已经过了。更现实的是分两条线做。\n\n- **披露透明**:投稿时主动声明哪些部分用了 AI、用了什么工具、做了什么程度的修改。期刊和会议越来越把这条写进 policy。\n- **关键段落护栏**:Methods 和 Results 这类承载事实的部分,要么完全人工写,要么留下可追溯的提示词与生成记录。讨论部分如果用 AI,至少要过一遍事实核查。\n\n对做 AI 工具的团队来说,这个数字是一个明确的产品信号:科研领域的 AI 写作助手如果只做「改写润色」,会越来越被视为可疑;如果把「保留可追溯记录 + 关键事实不可篡改」做成默认能力,反而能在新规则到来之前占据主动。\n\n对做科研管理和政策的人来说,这个数字是时候推动一些更具体的标准了——比如要求 LLM 介入的部分必须在 PDF 里以某种方式标记,或者在元数据里声明。期刊编辑联盟、资助机构、研究伦理委员会都需要在这一年里给出更明确的边界,否则规则只会越来越滞后于实际行为。\n\n## 个人判断\n\n77% 这个数字会被广泛传播,但更重要的是它打开了一个新的判断维度:学术产出的「真实人工贡献」正在成为一个需要单独计量的问题。未来的文献检索、文献评估、同行评议、AI 训练语料清洗,都会需要回答「这一段到底是谁写的、是谁改的」这个问题。\n\n这不是关于 AI 能不能写论文,而是关于学术界能不能诚实记录这件事。论文可以被 AI 写得更好,但前提是作者愿意在每一篇论文里诚实说明这一点。","https:\u002F\u002Fwww.nature.com\u002Farticles\u002Fd41586-026-02551-z","97acf9e4-deb3-41bb-8e98-9396e853733d",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"c33b1bbc-d6ce-4f61-9d5d-1a0704a6a09b","ai-policy",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":19,"name":20,"slug":20,"description":14,"color":14},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"b2f3dcd9-6056-4890-aa24-f677ff012fd4","en","PubMed Central study: 77% of 2025 biomedical papers show LLM writing traces","An arXiv preprint scanned PubMed Central's English biomedical corpus and found LLM-writing traces in 77% of 2025 papers (52% in 2024); the rate jumps to roughly 90% for December 2025 alone. A more sensitive word-frequency method raises the 2024 abstract-level estimate from 13.5% to 31%, consistent with a 2025 Wiley survey in which 71% of researchers self-reported AI writing assistance.","A preprint posted to arXiv on August 12 by Holzwarth, González-Márquez, and Dmitry Kobak at Ghent University did something simple that almost no one had actually done at scale: they ran the full text of every English biomedical paper archived in PubMed Central for 2024 and 2025 through a more sensitive word-frequency detector of large-language-model writing. The numbers are stark: 77% of English papers archived in 2025 carry detectable traces of LLM-assisted writing, compared with 52% in 2024, and for December 2025 alone the rate climbs to nearly 90% (see https:\u002F\u002Fwww.nature.com\u002Farticles\u002Fd41586-026-02551-z and the Solidot report https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85173).\n\nWhat makes 90% feel jarring is that it disagrees with every prior estimate. Kobak himself told Nature his first reaction was \"we must have done something wrong.\" After repeated checking, he conceded the data hold.\n\n## Why the numbers are so much higher than earlier studies\n\nFor the past few years, the field's best estimates of \"AI involvement in academic writing\" sat between 10% and 15%. A 2025 Kobak-team study of abstracts put 2024 at 13.5% or higher (https:\u002F\u002Fwww.nature.com\u002Farticles\u002Fd41586-025-02097-6). A 2026 cross-disciplinary study by Kyle Siler estimated 57% for 2025 (PNAS, https:\u002F\u002Fdoi.org\u002F10.1073\u002Fpnas.2605754123).\n\nThe new study's headline numbers are higher for two reasons. First, it analyses full text, not just abstracts. Second, it returns direct estimates instead of lower bounds. Re-running the new method against the 2024 abstract corpus bumps the detection rate from 13.5% to 31% — meaning earlier studies were undercounting 2024 itself by more than half.\n\nAn independent signal corroborates this: a 2025 Wiley survey found 71% of researchers self-reported using AI for writing assistance. Given that sensitive behaviors are almost always under-reported, 77% is not just plausible; it may even be conservative.\n\n## Not every section is touched equally\n\nThe most informative detail in the new study is that AI traces are heavily uneven across sections of a paper. For December 2025:\n\n- Discussion sections: about 78% show AI traces\n- Abstracts and introductions: similarly elevated\n- Methods and Results: about 58%\n\nThis distribution is not accidental. Abstracts, introductions, and discussions are the heaviest \"writing work\" sections, exactly where rewriting, polishing, and expansion by an LLM feels natural. Methods and Results, by contrast, carry the actual experimental design and data processing — the part the author knows best, and the part where AI assistance has the lowest motive.\n\nThat asymmetry is exactly what makes the situation dangerous. The low Methods\u002FResults detection rate does not mean those sections are safe; it means that if LLM contamination reaches them, the consequences are far worse. Kobak flags this directly: LLMs are prone to \"hallucinating\" data, so any AI痕迹 in Methods\u002FResults carries a non-trivial risk of fabricated or invented data.\n\n## Three under-appreciated risks\n\nWhat makes this worth more than a viral statistic is the three structural things it implies.\n\nFirst, peer review's detection capability has been overestimated. If 90% of papers pass review without triggering alarm, then reviewers are either unaware or the criteria simply do not cover this dimension.\n\nSecond, literature is being slowly biased. Kobak's quote in the Nature piece is worth lifting: \"Whatever bias the LLM may have will just suddenly permeate the literature.\" Unlike outright data fabrication, this permeation is gradual and distributed, which makes it harder to attribute and easier to ignore.\n\nThird, a tacit \"use equals comply\" consensus is forming. When 71% of researchers self-report use, 77% of papers show traces, and there is no wave of retractions or disputes, \"AI-assisted writing\" has quietly migrated from \"potentially non-compliant\" to \"default behavior.\" The rules did not change; the baseline drifted.\n\n## How peers are reading it\n\nThe Nature piece also carries reactions from independent scholars. Kyle Siler of the University of Toronto — author of the 57% cross-disciplinary estimate — is blunt: \"The toothpaste is out of the tube, and it's not going back.\" Other quoted researchers broadly agree that the high figure may not generalise to every discipline, but that LLM writing is now a default component of research output and is unlikely to reverse.\n\nA useful boundary: PubMed Central is biomedical-only. The preprint's dataset does not cover physics, computer science, mathematics, or engineering. Discipline-by-discipline writing habits and AI-motivation differ substantially; 90% does not automatically extrapolate across fields.\n\n## What to actually do now\n\nFor researchers, the takeaway is not \"ban AI from papers\" — that moment has passed. The realistic move is two-track.\n\n- **Disclose by default**: declare in the submission which sections used AI, which tool, and what degree of modification. Journals and conferences are increasingly codifying this.\n- **Guardrails on critical sections**: Methods and Results, which carry the facts, should either be fully human-written or carry traceable prompts and generation logs. If Discussion uses AI, run a fact-check pass at minimum.\n\nFor AI-tool builders, this is a clear product signal. A research-grade AI writing assistant that only does \"polish and rewrite\" will increasingly look suspicious. One that bakes in \"immutable provenance for every AI-touched paragraph\" by default can carve out position before the new rules arrive.\n\nFor research policy and ethics bodies, the signal is that more concrete standards are overdue. Mark AI-touched sections in the PDF or in metadata, require tool\u002Fversion disclosure, define what counts as authorship-grade AI contribution. Editorial consortia, funders, and ethics committees all need firmer positions inside the next 12 months — otherwise the rules will only fall further behind the actual behavior.\n\n## Personal take\n\nThe 77% number will travel. What matters more is that it forces a new question onto every paper, every reviewer, and every future AI training corpus: who actually wrote this paragraph, and who actually modified it?\n\nThat is not a question about whether AI can write a paper. It is about whether academia can record what happened, honestly, every time.","pubmed-77-percent-llm-writing-2025","2026-08-26T01:00:00Z","2026-08-26T03:04:54.100264Z","2026-08-26T03:04:54.100280Z",true,"agent",35,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"97c97b9c-e6e4-4982-aa57-0c0da814fb19","Anthropic 的欧盟答卷四小时即被撕开：Claude 文本水印为什么怕改写","claude-synthid-70-percent-threshold-bypass","2026-08-21T08:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"99916419-0f68-4a6a-a4cf-8bbe353b4d75","康涅狄格法官开出美国首例 prompt injection 制裁令:法庭文件里的隐藏 LLM 暗口令","us-court-prompt-injection-sanctions","2026-08-18T03:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"0c581f4f-468e-4272-bba3-2dcc4746e328","Suno 跟版权纠纷赛跑:8 月 6 日上线音频水印和歌词指纹化,押注「只标不评」的设计中立","suno-audio-watermarking-music-ai-policy","2026-08-17T12:30:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"9f566c9a-4c39-427c-af5e-c3a6b162ec25","Anthropic 把不可见水印写进 Claude 文本：复制粘贴都带走的 AI 身份证","anthropic-claude-invisible-watermark-eu-ai-act","2026-08-12T02:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"ca53004e-9180-4b9d-b9db-337f2d20994b","Anthropic 给 Claude 文本加水印:欧盟 AI Act 第 50 条第一次有了「出厂级」答案","anthropic-claude-text-watermark-eu-ai-act","2026-08-11T21:48:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"e4d0579a-1cc9-4064-a86d-a8c8e338e687","OpenJDK 发布生成式 AI 临时政策:零容忍,LLM 生成代码一律不准进社区贡献","openjdk-interim-ai-policy-no-llm-code","2026-08-09T02:00:00+00:00"]