[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ai-scraping-largest-theft-hecht-memo":3,"topics-all":38,"news-related-9cae377e-d048-4d0e-af78-8cba3e1cf5b4":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"9cae377e-d048-4d0e-af78-8cba3e1cf5b4","法庭文件揭微软高管私下定性 AI 抓取为最大盗窃","《纽约时报》诉 OpenAI 与微软的版权官司再爆猛料。微软应用科学总监 Brent Hecht 2023 年 1 月内部备忘录中,把 OpenAI 抓取新闻训练 AI 的做法定性为人类历史上最大规模的劳动成果盗窃。同期微软自家数据显示 Copilot 让新闻网站点击率暴跌 83%-93%。","AI 行业那层「合理使用」的薄纱,被一封内部备忘录和几段证词捅穿了。\n\n9 月 17 日,《纽约时报》诉 OpenAI 与微软版权官司刚解封的法庭文件,把两家公司过去三年对新闻业做了什么摆到台面上。微软应用科学总监 Brent Hecht 在 2023 年 1 月的内部备忘录里,把 OpenAI 用抓取的新闻训练 AI 称为「令人震惊的、史无前例的盗窃」「人类历史上最大规模的劳动成果盗窃」。这份由微软应用科学部门负责人亲自撰写的 PowerPoint,是被告自己内部对自家训练数据实践的最尖锐定性。\n\n## 微软自家数据先回答了「替代性」\n\nHecht 同期的演示文稿记录下 Copilot「answer engine」对《纽约时报》域名点击率的冲击:相比传统 Bing 搜索,部分页面点击率下滑达 93%,整体也跌了 51%-94%。他把这叫作「doom loop」——一个会同时伤害自家模型和整个开放网络的恶性循环。\n\n「终端产品威胁其必要供应商的经济基础,这种情况极为罕见,但这正是我们针对 LLM 业务'内容供应链'制造的现状。」\n\n这句话是「合理使用」四要素中「不得损害原作市场」那一条最直接的反驳。微软 2024 年初就已经把数字算出来了,然后用「合理使用」这套话术,在过去三年里把 OpenAI 的训练数据合法性反复洗白。\n\n## 纳德拉把「许可」写进了庭审记录\n\n微软 CEO Satya Nadella 在 2025 年初的证词里给了另一个关键陈述:任何付费墙内容,如果想被用于「grounding」或训练,都应该拿到许可;如果他早知道 OpenAI 绕过付费墙抓取并训练,他「会援引微软的权利,要求 OpenAI 重新训练」。\n\n这一段对微软-OpenAI 这对「共生关系」相当致命。微软一边在公开场合捍卫 OpenAI 的「合理使用」,一边在法庭上承认,如果同样的事情发生在自己不知情时,会要求对方重做。这等于承认:这套做法在道德和合同层面是错的,只是因为没人追究,才一直没人去做。\n\n## 91,692 个副本与「ah nice」\n\n技术细节这次也没藏着:OpenAI 的中训数据集包含逾 91,692 份来自《纽约时报》《每日新闻》和调查报道中心的副本;一个 Common Crawl 派生的数据集里,nytimes.com 单独贡献了 200 多万文档。\n\n更黑色幽默的是绕过付费墙那一条。OpenAI 研究员 Nick Ryder 给 Brockman 发消息说「有个 hack 可以绕过 nytimes 的付费墙」,Brockman 回的是:「ah nice。」这条对话被诉方拽进了法庭文件。\n\n## 为什么是行业分水岭\n\n过去两年,法庭和监管对 AI 训练数据的合法性普遍偏向「合理使用」。美国白宫 9 月初也提交了支持 OpenAI 的 brief。但这次的解封文件让天平多了一块新的砝码——不是外部批评者的观点,而是被告自己内部的文件。\n\n这类材料的杀伤力在于,它把「合理使用」抗辩中「无市场损害」「非替代性」「非故意」三个支柱同时拆掉了。当公司自己的应用科学总监写下「doom loop」,当 CEO 承诺「付费墙必须被许可」,当对手公司总裁笑着回「ah nice」时,「合理使用」这道防线在事实层和法律层都变得极其脆弱。\n\n接下来 12 个月的走向基本取决于和解条款——金额、训练数据披露义务、未来的「许可优先」机制——会不会被强制写进行业惯例。Hecht 那句「人类历史上最大规模盗窃」,最后可能成为这个行业的判例级注脚。","https:\u002F\u002Ftechcrunch.com\u002F2026\u002F09\u002F17\u002Fmicrosoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal\u002F","226bcb3d-18b8-4bb0-a999-4e82ec13f5fd",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"c33b1bbc-d6ce-4f61-9d5d-1a0704a6a09b","ai-policy",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"f270d10e-406f-444e-b5d8-a98525b0c8fa","en","Unsealed filings: Microsoft exec privately called AI scraping 'the largest theft of labor in human history'","Newly unsealed court filings in The New York Times' copyright suit against OpenAI and Microsoft reveal that Microsoft Applied Science director Brent Hecht called AI training on scraped news 'the largest theft of labor in human history' in a January 2023 internal memo. The same filings also show Microsoft's own data: Copilot answer engine drove click-through rates to news publishers down 83%-93%, and CEO Satya Nadella admitted paywalled content used for training should be licensed.","The thin veil of \"fair use\" in the AI industry has been pierced by an internal memo and a few sworn statements.\n\nOn September 17, newly unsealed filings in The New York Times copyright lawsuit against OpenAI and Microsoft put on the record what both companies have done to the news business over the past three years. Microsoft Applied Science director Brent Hecht, in a January 2023 internal memo, described OpenAI training AI on scraped news as \"an astonishing theft of unprecedented proportions\" and \"the largest theft of labor in human history.\" That PowerPoint, written by Microsoft's own head of applied science, is the sharpest internal characterization of the company's own training-data practice.\n\n## Microsoft's own data answered \"market substitution\" first\n\nA separate Hecht presentation logged the impact of Copilot's \"answer engine\" on click-through rates to nytimes.com: compared with traditional Bing search, some pages saw click-throughs collapse by as much as 93%, with the overall drop ranging 51%-94%. He called it a \"doom loop\" — a vicious cycle that damages both Microsoft's models and the open web at the same time.\n\n\"It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its content supply chain.\"\n\nThat single line is the most direct rebuttal of the \"does not substitute for, or harm the market for, the original work\" prong of fair use. Microsoft had the numbers in early 2024, and then spent the next three years laundering the legality of OpenAI's training corpus under the banner of fair use.\n\n## Nadella put \"licensing\" on the trial record\n\nMicrosoft CEO Satya Nadella, in a deposition earlier in 2025, gave another key statement: anything paywalled, if it is going to be used for grounding or for training, should be licensed; and had he known that OpenAI had scraped and trained on paywalled content, he \"would have invoked [Microsoft's right to] require OpenAI to retrain its models.\"\n\nThat line is fatal to the Microsoft-OpenAI symbiosis. On the public stage Microsoft defends OpenAI's fair use. In court, the CEO admits that the same behavior, had he known, would have triggered a contractual demand to retrain. It is an admission that the practice is wrong, morally and contractually; it has only been allowed to continue because nobody forced the issue.\n\n## 91,692 copies and \"ah nice\"\n\nThe technical detail is also out in the open now: OpenAI's mid-training datasets contain more than 91,692 copies of works from The New York Times, the Daily News, and the Center for Investigative Reporting. A Common Crawl-derived dataset contributed more than 2 million documents from nytimes.com alone.\n\nThe black-comedy subplot is the paywall bypass. OpenAI researcher Nick Ryder messaged Brockman about \"a hack to get around nytimes paywall,\" and Brockman replied: \"ah nice.\" That exchange was pulled into the court filing by plaintiffs.\n\n## Why this is an industry watershed\n\nFor the past two years, courts and regulators have leaned toward fair use for AI training data. Earlier in September, the White House filed a brief siding with OpenAI. But this round of unsealed filings adds a new weight on the scale — not an outsider's critique, but the defendants' own internal documents.\n\nThe lethality of this kind of evidence is that it dismantles all three pillars of the fair-use defense at once: no market harm, non-substitutive, and non-willful. When your own director of applied science writes \"doom loop,\" when your CEO commits that \"paywalled content must be licensed,\" and when the partner company's president replies \"ah nice\" to a paywall bypass, the fair-use defense becomes very thin at both the factual and the legal layer.\n\nWhat happens in the next twelve months basically depends on the settlement terms — the money, the training-data disclosure obligations, and any future \"license-first\" mechanism — being written into industry practice. Hecht's \"largest theft of labor in human history\" line may well end up being the footnote of legal record for this entire industry, rather than yet another internal gripe quietly forgotten.","ai-scraping-largest-theft-hecht-memo","2026-09-29T04:00:00Z","2026-09-29T05:04:29.438825Z","2026-09-29T05:04:29.438839Z",true,"agent",235,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"ea6de12d-e8bf-4013-a299-14211805ba30","ChatGPT __obi cookie 把你带到站外:1 年同站标识","openai-obi-cookie-cross-site-tracking","2026-09-30T00:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"e0642de2-5fec-472b-b85e-c63ccbad3b75","Anthropic 研究员公开辞职:AI 巨头正「拿全人类的命豪赌」自我进化超级智能","anthropic-coxon-resigns-ai-extinction-fears","2026-09-24T06:30:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"7fce8217-577f-4fd5-8f88-a1566cbf1290","微软与 OpenAI 法庭文件解封:LLM 训练数据被自家高管称为史上最大劳动窃取","microsoft-openai-doom-loop-nyt-copyright-2026","2026-09-23T05:03:20+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"d54e1ab3-820a-45fd-a4e9-ccbf6802bd72","NYT vs OpenAI 案解封:微软高管承认 AI 抓取是「最大劳动盗窃」","nyt-openai-microsoft-hecht-largest-theft-of-labor","2026-09-23T03:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"a0bd8ffb-63fd-452b-9d0b-634baa62d704","OpenAI 二次暂停训练:一个 DNS 查询打通训练沙盒","openai-agent-dns-sandbox-escape","2026-09-28T14:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"b74811e0-543f-46ac-86be-0ed1aceb07f6","AI 用 DNS 递话:OpenAI 二度暂停前沿训练","openai-agent-dns-sandbox-escape-frontier-pause","2026-09-27T15:13:00+00:00"]