[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-microsoft-openai-doom-loop-nyt-copyright-2026":3,"topics-all":38,"news-related-7fce8217-577f-4fd5-8f88-a1566cbf1290":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"7fce8217-577f-4fd5-8f88-a1566cbf1290","微软与 OpenAI 法庭文件解封:LLM 训练数据被自家高管称为史上最大劳动窃取","纽约时报诉 OpenAI 版权案解封文件显示,微软应用科学总监 Brent Hecht 形容 AI 抓取新闻训练为人类历史上最大规模的劳动盗窃,微软内部承认已形成死亡循环——AI 把新闻出版商的点击率压掉 51% 到 94%,反过来也吃掉了自己模型的训练来源。","最近一周,科技圈被一份解封的法庭文件炸开了锅。《纽约时报》诉 OpenAI 和微软的版权案,原告律师团提交了一份 92 页未删节的摘要判决动议,大量引用了微软和 OpenAI 高管在内部文件和作证时的原话——这些话,被微软和 OpenAI 此前以\"商业秘密\"为由遮了三年多。\n\n## 内部文件说了什么\n\n最炸的一句来自微软应用科学总监 Brent Hecht。他在 2023 年 1 月的内部备忘录里写道:抓取新闻内容训练 AI 是\"人类历史上最大规模的劳动成果盗窃\",这一行为\"完全是对合理使用理念的嘲弄\"。在另一处,他把 LLM 的训练机制描述为\"在没有把经济价值分配到供应链的情况下窃取内容,这必然威胁内容创作者的经济稳定\"。\n\n比这句话更可怕的是微软自己写的一份策略文件,直接承认 AI 业务已经陷入\"死亡循环\"。文件原话是:\"我们的 AI 内容战略已开始形成'死亡循环',会同时伤害我们模型的性能和整个 Web——一个终端产品威胁到它基础供应商的经济基础,这非常罕见,但这就是我们为 LLM 业务在'内容供应链'上制造的局面。\"同一份策略文件还下了个结论:\"LLM 是一种会摧毁自身供应链的产品。\"\n\n微软的数据印证了这一点。文件指出,部分新闻出版商网站的点击率下降了 83%-93%,其它新闻机构下降 51%-94%。微软 CEO 纳德拉在宣誓证词中承认,抓取《纽约时报》及新闻网站后,Bing 上对这些网站的点击\"完全崩盘\",降幅超过 90%。换句话说,AI 用他们的内容训练出了更好的搜索引擎,然后再把这些内容提供方从搜索结果里挤掉。纳德拉还确认:聊天机器人本质上已\"替代\"了去原始网站获取信息的行为。\n\nOpenAI 侧的故事同样不光彩。文件披露,联合创始人 Greg Brockman 得知爬虫绕过《纽约时报》付费墙时回了句\"Ah, nice\";OpenAI 的企业代表在证词中承认,他\"不知道\"公司做过任何\"检测付费墙内容\"的努力。ChatGPT 业务负责人 Nick Turley 在内部沟通中写道,AI 聊天机器人对出版商构成\"生存性威胁\",因为它们\"基本上是替代性的\"。政策总监 Jack Clark 补充:\"我们正在创建替代定义社会'文化'的人们的劳动的系统。\"\n\n## 这份文件为何重要\n\n这份文件的意义,不在于披露了外界不知道的事。它的意义在于:这是被告自己的嘴说出来的。微软和 OpenAI 的律师团在法庭上的核心抗辩是\"合理使用\"和\"转换性使用\",即训练 LLM 是高度转换性的、不替代原作的使用。但他们自己的高管在同一时间线上的内部文件里,把这些模型形容成\"盗窃\"、\"替代劳动\"、\"死亡循环\"、\"摧毁供应链\"。内部承认和外部抗辩之间的鸿沟,正是《纽约时报》律师团要拿来说服法官的关键弹药。\n\n## 行业影响与更深的问题\n\n这份文件的影响远超这一桩诉讼——它把 AI 公司对训练数据合规性的真实内部评估,第一次大规模公开化。接下来几个月,围绕训练数据授权、\"opt-out\"机制和\"AI 检索分成\"的讨论会进一步加速。OpenAI 自己的经济学专家也承认,Google AI Overviews 可能让新闻出版商搜索推荐流量下降 20% 到 60%,说明整个 AI 检索生态对内容产业的挤压是系统性的。\n\n更深一层的信号是:即便 AI 公司赢了\"合理使用\"的判决,\"死亡循环\"的问题也不会消失。LLM 的质量依赖高质量的原始内容,如果 AI 把内容产业的商业模式摧毁,下一代的训练数据从哪里来?这是 AI 行业不能只交给律师去回答的问题。","https:\u002F\u002Fwww.theverge.com\u002Fai-artificial-intelligence\u002F997633\u002Fopenai-microsoft-chatgpt-ai-new-york-times-doom-loop-theft-google-zero","05ad777c-69bc-46a5-bca4-df8e4b3c8ee5",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"c33b1bbc-d6ce-4f61-9d5d-1a0704a6a09b","ai-policy",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"3c41e3b1-60c8-4555-b7da-adfab0e4a954","en","Microsoft OpenAI unsealed files: LLMs ate their own web","Unsealed documents in the NYT vs OpenAI copyright case show Microsoft Applied Science Director Brent Hecht called AI scraping the largest theft of labor in human history, while internal Microsoft files admit a doom loop is cutting publisher click-throughs by 51-94 percent and eating the very training sources the models depend on.","Tech circles were rocked this week by an unsealed court filing. In the New York Times vs OpenAI and Microsoft copyright lawsuit, the plaintiff's legal team submitted a 92-page unredacted summary judgment motion packed with quotes from internal documents and sworn testimony by Microsoft and OpenAI executives — statements that Microsoft and OpenAI had kept sealed for over three years under \"trade secret\" claims.\n\nThe most explosive quote came from Microsoft Director of Applied Science Brent Hecht. In an internal memo from January 2023, he wrote that scraping news content to train AI was \"the largest theft of labor in human history\", calling it \"a complete mockery of the fair use doctrine.\" Elsewhere, he described LLM training as \"stealing content without distributing economic value down the supply chain, which necessarily threatens the economic stability of those who create the content.\"\n\n## What the documents say\n\nEven more damning is a Microsoft-authored strategy document that admits the AI business has entered a \"doom loop\". The document reads: \"Our AI content strategy has started a 'doom loop' that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain'.\" That same strategy document also concludes: \"LLMs are a product that destroys its own supply chain.\"\n\nMicrosoft's own data backs this up. The filing states that click-throughs to some publisher sites fell by 83-93%, and to other news organizations by 51-94%. In sworn testimony, Microsoft CEO Satya Nadella admitted that after scraping the New York Times and other news sites, clicks to those sites on Bing \"completely cratered\", dropping over 90%. In other words, AI used their content to train better search engines — and those search engines then squeezed the same content providers out of search results. Nadella also confirmed: chatbots have essentially \"substituted\" the act of going to original websites for information.\n\nOpenAI's side of the story is no better. The filing reveals that when co-founder Greg Brockman was told their crawler had bypassed the New York Times paywall, he replied \"Ah, nice\". An OpenAI corporate representative testified he was \"unaware\" of any effort by the company to \"detect paywalled content\" in training data. Head of ChatGPT Nick Turley wrote internally that AI chatbots pose an \"existential threat\" to publishers because they \"are largely substitutive\". OpenAI policy director Jack Clark added: \"We are creating systems that substitute for the labor of the people that define the 'culture' of society.\"\n\n## Why this filing matters\n\nThe filing's significance is not that it reveals facts the world didn't already know. What matters is that these admissions come from the defendants' own mouths. Microsoft and OpenAI's core legal defense has been \"fair use\" and \"transformative use\" — that training an LLM is highly transformative and does not substitute for the original work. But their own executives, in internal documents on the same timeline, described these models as \"theft\", \"labor substitution\", \"doom loop\", and \"supply chain destruction\". That gap between internal admission and external defense is exactly the ammunition the NYT legal team is using to argue for summary judgment.\n\n## Industry impact and the deeper question\n\nThe filing's impact extends far beyond this single lawsuit — it is the first time AI companies' internal assessments of their own training data compliance have been put on the public record at this scale. In the coming months, debates around training data licensing, publisher \"opt-out\" mechanisms, and whether AI-powered retrieval should share revenue with content sources will accelerate. OpenAI's own economic expert also admitted that Google's introduction of AI Overviews may have depressed search referrals to publishers by 20 to 60 percent, suggesting that the entire AI retrieval ecosystem's squeeze on content industries is systemic.\n\nAn even deeper signal: even if AI companies win the \"fair use\" argument, the \"doom loop\" problem doesn't disappear. LLM quality depends on high-quality original content — if AI destroys the content industry's business model, where does the next generation of training data come from? That is a question the AI industry cannot leave solely to its lawyers.","microsoft-openai-doom-loop-nyt-copyright-2026","2026-09-23T05:03:20Z","2026-09-23T05:06:22.716104Z","2026-09-23T05:06:22.716113Z",true,"agent",31,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"d54e1ab3-820a-45fd-a4e9-ccbf6802bd72","NYT vs OpenAI 案解封:微软高管承认 AI 抓取是「最大劳动盗窃」","nyt-openai-microsoft-hecht-largest-theft-of-labor","2026-09-23T03:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"d569d88a-8906-4f3b-83fd-0dcc6b5c75e0","OpenAI 用 __obi 把 ChatGPT 账号绑上你全网浏览","openai-obi-cookie-cross-site-tracking-chatgpt","2026-09-23T07:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"9dd4a859-1153-4ecc-b69d-4ba4c5431129","智谱被开发者抓包后紧急上线数据零留存","zhipu-maas-zero-data-retention-zcode","2026-09-21T07:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"d157b4f9-537e-405c-b557-859f6d2cf18c","微软自家高管警告:抓新闻训 AI 是「人类史上最大规模劳动盗窃」","microsoft-ai-scraping-theft-of-labor","2026-09-19T00:11:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"95e9bb62-0bd3-4c2f-913a-302ba5e2ace8","Anthropic 9 月报告把蒸馏战摆上台面:151 亿次阿里请求、解放军流量走 Moonshot","anthropic-distillation-report-china-200m-claude","2026-09-18T03:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"1d113d73-3774-426a-bdc0-49c678a96a59","Bengio 长文复盘:AI 智能体说谎作弊,病根在训练目标打架","bengio-ai-agents-misalignment","2026-09-14T17:10:00+00:00"]