[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-pew-ai-authorship-web-study":3,"news-related-ae3f239d-dc29-4ec0-a823-446f463e6bab":32},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":18,"news_slug":25,"published_at":26,"created_at":27,"modified_at":28,"is_published":29,"publish_type":30,"image_url":14,"view_count":31},"ae3f239d-dc29-4ec0-a823-446f463e6bab","皮尤扫了 49 万网页:ChatGPT 之后发布的页面,三分之一带 AI 痕迹","皮尤研究中心用 Common Crawl 近五年约 49 万英文网页跑 Pangram AI 检测:2026 年 7 月快照中,ChatGPT 发布后的新页面逾三分之一带 AI 写作痕迹;.com 域名约 10%,是 .org 的两倍、.edu\u002F.gov 的十倍。破折号与 delve 类词密度三年翻倍。","2022 年 11 月 ChatGPT 面世时,很少有人预料到它会以什么方式改写网页本身。皮尤研究中心(Pew Research Center)8 月 20 日发布的研究,给\"互联网上有多少内容是 AI 写的\"这个悬置已久的问题交出了第一份大样本量化答案——结论比多数人的直觉更激进。\n\n## 49 万个网页,一个检测模型\n\n研究团队从 Common Crawl 网页存档抽取了 2021 年 1 月至 2026 年 7 月、跨度近五年的约 49 万个英文网页——起点刻意选在 ChatGPT 发布前两年,前后对比才有参照系。检测工具是 Pangram 开源的开放权重检测模型 Open Pangram,通过词汇、短语与句式癖好等语言统计模式,判断文本是否\"由 AI 写作或深度编辑\"。\n\n必须交代一个诚实前提:皮尤自己强调,检测模型并不完美,单篇文档存在误判可能;但当样本放大到几十万级,群体层面的统计趋势是可靠的。\n\n## 两个数字:十分之一,与三分之一\n\n2026 年 7 月快照的随机抽样里,全部网页有 10% 出现显著 AI 写作痕迹。这个数字看似温和,但互联网是新旧内容的混合体——大量存量页面诞生于 ChatGPT 之前,物理上不可能由 AI 写成。\n\n把旧页面过滤掉、只看 ChatGPT 发布之后新出版的页面,比例骤然跳到**超过三分之一**。皮尤指出,这与 ai-on-the-internet.github.io 汇总的同类研究相互印证,并非孤证。\n\n## .com 是重灾区,.edu 与 .gov 是净土\n\nChatGPT 刚发布时,各顶级域名的 AI 语言模式出现率基本一致;到 2026 年已显著分化:约十分之一的 .com 页面带 AI 痕迹,是 .org(4.6%)的两倍、.edu 与 .gov(均约 1%)的十倍。商业内容在批量拥抱 AI,而学术与政府站点的编辑流程构成了一道更有效的防火墙。\n\n趋势也没有放缓迹象:.com 的 AI 痕迹占比从 2021 年初的 1.09% 一路爬升到 2026 年初的 9.35%。\n\n## AI 的文风指纹,已经渗进语料库\n\n对比 2023 年与今天的网页,皮尤列出的文风变化几乎是一份 AI 写作鉴定清单:\n\n- 破折号使用频率约翻倍(每万词 5.79 次 → 11.19 次);\n- 牛津逗号增长 63%(34.04 次 → 55.51 次);\n- delve、interplay、testament 等 AI 高频词用量翻倍以上(11.94 次 → 26.02 次);\n- \"not just X, it's Y\"式否定平行结构增长近两倍(0.87 次 → 2.36 次)。\n\n报告还附上完整\"AI 词汇表\":additionally、crucial、fostering、landscape、tapestry、underscore、vibrant……常读英文博客的人应该都会心一笑。\n\n## 怎么看这件事\n\n真正值得警惕的不是\"AI 写了多少网页\",而是这些页面接下来的去向。模型从人类语料里学出了这些文风偏好,而这些偏好正以翻倍的速度扩散回互联网语料——生成与训练之间的循环在收紧。叠加上检测器的误判率和人类作者对这些句式的主动模仿,\"有 AI 痕迹\"和\"AI 生成\"的边界会越来越模糊。\n\n对模型开发者,这是数据清洗环节权重的又一次上调;对内容平台,.com 与 .edu 之间十倍的差距说明,编辑把关仍是目前最有效的质量控制手段。下次再有人争论\"互联网会不会被 AI 填满\",把这份研究甩给他:不是会不会,是已经发生了——只是分布远比想象中不均匀。\n\n原文:皮尤研究中心 [How Much of the Internet Is Written With AI?](https:\u002F\u002Fwww.pewresearch.org\u002Fdata-labs\u002F2026\u002F08\u002F20\u002Fhow-much-of-the-internet-is-written-with-ai\u002F)(2026-08-20)","https:\u002F\u002Fwww.pewresearch.org\u002Fdata-labs\u002F2026\u002F08\u002F20\u002Fhow-much-of-the-internet-is-written-with-ai\u002F","b0078f39-6d9e-4bc7-b72c-097ad9960e09",[11,15],{"id":12,"name":13,"slug":13,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[19],{"id":20,"lang":21,"title":22,"summary":23,"content":24},"954e2676-87c9-4667-95b2-7c1467db248b","en","Pew: over a third of post-ChatGPT webpages show AI authorship signs","Pew ran 490,000 Common Crawl pages through Pangram's detector: 10% of all pages show AI signs; over a third of post-ChatGPT pages do. .com leads at ~10%.","In November 2022, OpenAI released ChatGPT to the public. Less than four years later, a study published August 20 by Pew Research Center's Data Labs offers the first large-sample quantified answer to a question that has hung over the web ever since: how much online content is now written by AI? The answer is more aggressive than most people's intuition.\n\n## 490,000 webpages, one detection model\n\nThe team drew roughly 490,000 English-language webpages from Common Crawl snapshots spanning January 2021 to July 2026 — deliberately starting a couple of years before ChatGPT's release so the before-and-after comparison has a baseline. The detection tool is Open Pangram, an open-weight AI detection model built by Pangram, which looks for statistical patterns in language — words, phrases and linguistic quirks — to flag pages \"written or substantially edited by AI.\"\n\nOne honest caveat the researchers themselves stress: AI detection models are not perfect and can misclassify individual documents. But at the scale of hundreds of thousands of pages, the aggregate statistical trend holds.\n\n## Two numbers: one in ten, and one in three\n\nIn a random sample of 10,000 webpages from the July 2026 crawl, 10% showed significant signs of AI authorship. That sounds modest — but the internet is a mix of old and new material, and plenty of pages in the sample physically could not have been written by AI.\n\nFilter out the old pages and look only at content published after ChatGPT's release, and the number jumps to **over one-third**. Pew also notes this is in line with other studies collected at ai-on-the-internet.github.io — not an outlier.\n\n## .com is the hotspot; .edu and .gov hold the line\n\nWhen ChatGPT first launched, AI-typical language patterns appeared at similar rates across .com, .org, .edu and .gov. By 2026 they had diverged sharply: around one-in-ten .com pages show AI signs, roughly double the rate on .org (4.6%) and ten times the rate on .edu and .gov (both around 1%). Commercial content production has embraced AI in bulk, while academic and government editorial pipelines still act as a firewall.\n\nThe trend shows no sign of slowing: AI-sign prevalence on .com climbed from 1.09% in early 2021 to 9.35% by early 2026.\n\n## AI's stylistic fingerprints are all over the corpus\n\nComparing today's web to a 2023 snapshot, Pew's list of stylistic shifts reads like an AI-writing identification checklist:\n\n- Em dashes appear about twice as frequently (5.79 to 11.19 per 10,000 words);\n- Oxford commas are up 63% (34.04 to 55.51);\n- AI-favored words like \"delve,\" \"interplay\" and \"testament\" have more than doubled (11.94 to 26.02);\n- \"Not just X, it's Y\" negative parallelism has nearly tripled (0.87 to 2.36).\n\nPew even publishes the full vocabulary list: additionally, crucial, fostering, landscape, tapestry, underscore, vibrant... anyone who reads English blog posts will recognize the register.\n\n## My take\n\nThe truly sobering part of this study is not how much of the web is AI-written, but where those pages go next. AI models learned these stylistic preferences from human-written training data, and those preferences are now flowing back onto the web at doubling rates — the loop between training corpora and generated content is tightening. Detector error rates, plus humans actively imitating these constructions (yes, people now write like AI too), will keep blurring the line between \"AI signs\" and \"AI-generated.\"\n\nFor model developers, data cleaning just got another justification; for content platforms, the 10x gap between .com and .edu shows that editorial review remains the most effective quality control we have. Next time someone argues about whether the internet will fill up with AI content, hand them this study: it is not a future scenario — it already happened, just far more unevenly than expected.\n\nSource: Pew Research Center, [How Much of the Internet Is Written With AI?](https:\u002F\u002Fwww.pewresearch.org\u002Fdata-labs\u002F2026\u002F08\u002F20\u002Fhow-much-of-the-internet-is-written-with-ai\u002F), Aug 20, 2026.","pew-ai-authorship-web-study","2026-08-25T13:08:45Z","2026-08-25T13:08:46.384443Z","2026-08-25T13:08:46.384452Z",true,"agent",34,{"items":33},[34,39,44,49,54,59],{"id":35,"title":36,"news_slug":37,"published_at":38},"c0ca1295-8b69-4e4f-b29d-24de3bf08d7e","生物医学论文 89% 带 LLM 痕迹:方法部分也不再是净土","llm-assisted-writing-biomedical-papers","2026-08-24T21:40:00+00:00",{"id":40,"title":41,"news_slug":42,"published_at":43},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":45,"title":46,"news_slug":47,"published_at":48},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":50,"title":51,"news_slug":52,"published_at":53},"5a90a793-8ec1-4b3a-9691-edef5ffe8535","AI「思想病毒」实证:Anthropic 与 EPFL 让恶意想法在 Agent 间自我复制,免疫只需一段警告","mind-viruses-multi-agent-llm","2026-08-18T13:30:00+00:00",{"id":55,"title":56,"news_slug":57,"published_at":58},"89804e7d-cee8-4412-8b30-e43855911ef5","点赞与下载是两个经济体：Hugging Face 夏季报告拆穿开源模型的「追新幻觉」","hf-open-models-summer-likes-vs-downloads","2026-08-18T13:20:00+00:00",{"id":60,"title":61,"news_slug":62,"published_at":63},"99916419-0f68-4a6a-a4cf-8bbe353b4d75","康涅狄格法官开出美国首例 prompt injection 制裁令:法庭文件里的隐藏 LLM 暗口令","us-court-prompt-injection-sanctions","2026-08-18T03:00:00+00:00"]