[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-pew-research-ai-web-content-2026":3,"topics-all":38,"news-related-1464179a-2b7f-4369-b680-25868ddd9042":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"1464179a-2b7f-4369-b680-25868ddd9042","皮尤实测：超过三分之一 ChatGPT 后的英文网页已有 AI 写作痕迹","皮尤用 Open Pangram 扫 49 万份 Common Crawl 网页:2026 年 7 月快照 10% 全样本、逾三分之一 ChatGPT 后发布的页面有明显 AI 写作痕迹,.com 域名占比约为 .edu\u002F.gov 的十倍,em dash 等 AI 语言模式较 2023 年显著上升。","In November 2022, OpenAI pushed ChatGPT into the public spotlight. Less than four years later, the Pew Research Center pulled apart half a decade of English-language webpages and asked a plain question: how much of the modern internet is already written by machines?\n\n## 数据从哪里来:49 万份 Common Crawl,一台检测器\n\n皮尤数据实验室的高级数据科学家 Samuel Bestvater 领衔的团队,从 Common Crawl 这套公开网页档案里抽取了大约 49 万份英文网页样本,时间跨度从 2021 年 1 月一直覆盖到 2026 年 7 月——也就是 ChatGPT 公测前两年到当下。文本经过一款名叫 Open Pangram 的开源 AI 检测器打分。Pangram 由第三方研究机构 Pangram Labs 训练,模型权重开放下载,逻辑是抓取 AI 与人类写作之间统计意义上的分布差异,识别哪些样本里 AI 写作的语言指纹明显。\n\n这是一次大样本、明确方法论、明确工具的研究。皮尤在 2026 年 8 月 20 日发表了完整数据论文,并公开了原始 CSV 数据。\n\n## 三组数据:10%、三分之一、十倍\n\n最显眼的一组数字:把全部样本时间窗口合在一起看,2026 年 7 月快照里 **10% 的网页呈现明显的 AI 写作痕迹**。听上去比例不高,但这背后是互联网里混着大量 ChatGPT 之前的旧页面。皮尤做了一个更锋利的子集筛选——只保留 2022 年 11 月底 ChatGPT 发布之后才上线的网页,这一组里 **超过三分之一(逾 33%)的页面被检测器判定有明显 AI 写作痕迹**。\n\n顶级域名分布的差异更加戏剧性。2026 年样本里,**.com 域名网页约 10% 命中 AI 痕迹**,几乎是 .org 域名(4.6%)的两倍,又是 .edu 和 .gov 域名(都在 1% 上下)的十倍。换句话说,商业网站是 AI 内容的主战场,学术、政府机构在内容生成上对生成式工具的采纳明显保守。\n\n## AI 写手留了哪些指纹:em dash、牛津逗号、负面平行句式\n\n皮尤同时分析了 AI 写作偏爱的语言特征在网页上的变化。相比 2023 年的网页样本,2026 年的网页样本里有几个信号被显著放大:\n\n- **em dash(——)**:出现频率大约是 2023 年的两倍。\n- **牛津逗号**(列举项里最后一项前面的逗号):使用频率上升 63%。\n- **AI 偏爱词汇**:delve、interplay、testament 等代表词的出现频率翻了一倍以上。\n- **负面平行句式**(\"it's not just X, it's Y\" 或者中文里\"不只是 X,更是 Y\"那种结构):几乎是三倍,虽然绝对比例仍然偏低。\n\nOpen Pangram 并不是靠这些\"标点症\"和\"用词习惯\"做单一判定,但大样本对照下,这些特征确实成了 AI 文本的稳定指纹。\n\n## 检测工具本身不是神:误判与对抗\n\n皮尤在论文里反复提醒:检测模型偶尔会把人类作品判成 AI,反过来也会漏判。AI 写作也在进化,刻意改写、加入更多人类写作风格,会让判定变难。Open Pangram 的策略是看统计意义上多维度的偏差,而不是盯住单一标志。这意味着同一篇文本在 2026 年的检测分数可能和 2027 年不一样——检测器和生成器会持续互相对抗。\n\n这也是皮尤把数据集和工具全部公开的原因:研究者、新闻机构、监管者都可以自己复跑、迭代检测器,而不是被某一家的黑盒分数绑架。\n\n## 对内容生态意味着什么:信息源分层在加速\n\n数字背后真正的信号是信息源分层在加速。.edu 和 .gov 域名里 AI 痕迹占比只有 1%,意味着这些机构的内容生成仍以人类把关为主;而商业 .com 域名的 AI 写作占比已经接近学术、政府机构的十倍——读者面对同样的\"网页\",面对的是两种性质完全不同的内容供给:一边还带着机构审核的影子,另一边更可能是流水线产物。\n\n这种分层短期对消费者不友好:搜索引擎和聚合平台的排序不会自动告诉你面前这篇文章是不是机器生成。但长期看,它会倒逼内容平台、内容订阅服务、SEO\u002F反 AI 检测工具形成新的合同关系——内容创作者需要明确声明是否使用了 AI 生成辅助,平台需要提供 AI 痕迹过滤,而读者侧会出现\"我想看人类写的\"作为新的内容筛选维度。\n\n## 所以呢:别再问\"机器会不会写\",该问\"我怎么知道这是机器写的\"\n\n皮尤这份研究最重要的一步,不是给 AI 写作占比盖棺定论,而是把检测工具和样本都摆到桌面上。Common Crawl 是公开数据,Open Pangram 的权重是开源的,皮尤也公开了每一个时间点的具体比例。当我们习惯说\"现在 AI 已经能写任何东西\"时,这份 49 万页面的对照研究提醒了一件更精细的事实:**能写**和**被检测器识别**之间仍然存在一个可以量化的差,这个差既是技术问题,也是治理问题。\n\n接下来一两年,值得继续盯的是同一组样本在更晚时间的检测结果——AI 生成的占比是继续往上爬,还是在某个平台\u002F规则约束后出现拐点,以及检测工具本身的容错率能不能跟得上生成器的进化。这两份曲线哪个先掉头,就会决定\"机器写的互联网\"在我们日常生活里到底显现成什么样子。\n\n参考:Pew Research Center \"How Much of the Internet Is Written With AI?\",2026 年 8 月 20 日。原文链接见 NewsForAI。\n","https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85172","b0078f39-6d9e-4bc7-b72c-097ad9960e09",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"8ddf2b28-0234-41a4-9862-3f0faef96472","market-analysis",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"0044c03d-d220-4323-8a36-d081b6fb5aae","en","Pew Data: Over One-Third of Post-ChatGPT English Web Pages Already Carry AI Writing Traces","Pew Research Center used Open Pangram to scan about 490,000 Common Crawl English-language pages. In the July 2026 snapshot, 10% of the full sample and over one-third of pages published after ChatGPT show significant signs of AI authorship. .com domains run about 10x the .edu\u002F.gov rate, and AI-favored language patterns like em dashes have risen sharply since 2023.","In November 2022, OpenAI pushed ChatGPT into the public spotlight. Less than four years later, the Pew Research Center pulled apart half a decade of English-language webpages and asked a plain question: how much of the modern internet is already written by machines?\n\n## Where the data came from: 490,000 Common Crawl pages and one detector\n\nA team led by Samuel Bestvater, senior data scientist at Pew's Data Labs, drew roughly 490,000 English-language pages from Common Crawl, the open web archive, spanning January 2021 through July 2026 — roughly two years before ChatGPT's public release through the present. Every page was scored by an open-weight detection model called Open Pangram, built by Pangram Labs. The model's logic is statistical: it looks for distributional differences between how large language models write and how human authors write, then flags pages where the AI fingerprint is clearly visible.\n\nThe study is large sample, explicit methodology, explicit tool. Pew published the full data essay on August 20, 2026, and released the underlying CSV.\n\n## Three numbers that matter: 10%, one-third, tenfold\n\nThe headline number: across the full sample window, the July 2026 snapshot found that **10% of pages show significant signs of AI authorship**. That sounds modest until you remember that web crawls mix in lots of pre-ChatGPT material. Pew sharpened the lens further, filtering down to pages published after ChatGPT's release in late November 2022. In that subset, **more than one-third (over 33%) of pages show clear AI writing fingerprints**.\n\nThe split across top-level domains is the sharpest signal. In 2026 samples, **roughly one in ten .com pages showed AI authorship signs**, almost double the .org rate (4.6%) and about ten times the .edu and .gov rates (both around 1%). In other words, the commercial web is where AI content lives; academic and government domains are still leaning on human-reviewed writing.\n\n## The fingerprints AI writers leave behind: em dash, Oxford commas, negative parallelism\n\nPew also tracked how AI-favored language features have shifted at the corpus level. Compared with the 2023 sample, by 2026:\n\n- **Em dashes (—)** appear at roughly twice the rate.\n- **Oxford commas** have risen by 63% in usage frequency.\n- **AI-vocabulary tokens** like \"delve,\" \"interplay,\" and \"testament\" more than doubled.\n- **Negative parallelism** (\"it is not just X, it is Y\") nearly tripled, though still rare in absolute terms.\n\nOpen Pangram is not relying on any single tic. The detector looks at multi-dimensional statistical drift. But at corpus scale these signals stack up into reliable AI fingerprints.\n\n## Detection is not infallible: error and adversarial drift\n\nPew is careful to flag that detection models routinely mislabel human work as AI, and miss AI work dressed up as human. Generators also evolve. The current Pangram model is built to read statistical patterns across many small signals, which means the same passage could score differently in 2026 than in 2027 as both sides keep moving.\n\nPew released the data and the detector openly so researchers, newsrooms, and regulators can rerun and iterate, instead of trusting any single black-box score.\n\n## What this means for the content ecosystem: stratification accelerates\n\nThe real signal behind the numbers is the accelerating stratification of information sources. .edu and .gov domains are still at 1% AI content, which means those institutions mostly run human-gated writing. Commercial .com domains are running at roughly ten times that rate, meaning more and more of what readers see on the open web is assembly-line output.\n\nThis stratification is bad for consumers in the short run. Search engines and aggregators do not flag \"this page is likely machine-written\" by default. In the longer run, it forces content platforms, subscription services, and SEO \u002F anti-AI detection vendors into new contracts: publishers have to disclose AI-assist use, platforms have to ship AI-trace filters, and on the reader side \"I want to read something written by a human\" becomes a new filter dimension.\n\n## So what: stop asking \"can machines write\" and start asking \"how do I know this was written by one\"\n\nThe most useful move in this study is not the final AI-percentage headline. It is the decision to put the detection tool and the dataset on the table. Common Crawl is public, Open Pangram is open-weight, and the per-period percentages are downloadable. When we casually say \"AI can write anything now,\" this 490,000-page controlled comparison offers a more precise counterpoint: between **can write** and **detected as AI-written** there is still a quantifiable gap, and that gap is both a technical problem and a governance problem.\n\nOver the next year or two, the same dataset at a later timestamp will be the line to watch. Does the AI-writing share keep climbing, or does it bend once a platform or rule puts pressure on the curve? And does the detector keep up with the generator? Whichever curve bends first will shape what \"a machine-written internet\" looks like in everyday reading.\n\nReference: Pew Research Center, \"How Much of the Internet Is Written With AI?\", August 20, 2026. Original link via NewsForAI.\n","pew-research-ai-web-content-2026","2026-08-31T03:00:00Z","2026-08-31T03:05:52.102228Z","2026-08-31T03:05:52.102243Z",true,"agent",152,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"d17a841b-abca-46e0-80e4-d955f1c837ba","亚马逊 VGT3 仓库曝光:一天拆掉上千本书,只为给 AI 模型喂语料","amazon-vgt3-warehouse-ai-training-books","2026-09-07T03:30:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"21a91da5-c5fa-45e2-b01f-a7950331cf44","S3 把 DuckDB 团队收走了:DuckLabs 加盟 AWS,MIT 开源照旧","aws-buys-ducklabs-duckdb-open-source","2026-08-30T06:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"216f3c2b-d551-45fc-a206-c3ccfae9db89","亚马逊 Mechanical Turk 将永久关闭:被 AI 掏空的众包平台","amazon-mechanical-turk-shutdown","2026-08-29T17:30:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"89804e7d-cee8-4412-8b30-e43855911ef5","点赞与下载是两个经济体：Hugging Face 夏季报告拆穿开源模型的「追新幻觉」","hf-open-models-summer-likes-vs-downloads","2026-08-18T13:20:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"22a1a718-0eb6-46e5-8ee8-825400de11d1","DeepMind WeatherNext 在 Nature 发论文：用 28 km 粗分辨率做出多一天的飓风预警,代码权重全部开源","deepmind-weathernext-cyclones-nature-open-source","2026-08-10T02:00:00+00:00"]