[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-nyt-openai-microsoft-hecht-largest-theft-of-labor":3,"topics-all":38,"news-related-d54e1ab3-820a-45fd-a4e9-ccbf6802bd72":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"d54e1ab3-820a-45fd-a4e9-ccbf6802bd72","NYT vs OpenAI 案解封:微软高管承认 AI 抓取是「最大劳动盗窃」","TechCrunch 9 月 17 日报道,NYT 诉 OpenAI 与微软案新解封文件显示,微软应用科学总监 Brent Hecht 在 2023 年内部备忘录把 AI 数据抓取描述为「人类历史上最大规模的劳动盗窃」;同期披露 OpenAI 训练集含逾 9.16 万份 NYT 系作品。","纽约时报诉 OpenAI 与微软的版权诉讼,三年里第一次露出这种程度的内部表态。9 月 17 日 TechCrunch 拿到的新解封文件里,微软应用科学总监 Brent Hecht 在 2023 年 1 月的内部备忘录里把 AI 数据抓取直接称为「an astonishing theft of unprecedented proportions」与「the largest theft of labor in human history」。同份文件还披露了具体的训练集规模与商业影响——这些细节以前只活在猜测里。\n\n## 文件说了什么:训练集规模与商业伤害都被量化\n\n这次解封内容把过去几年版权诉讼里反复被质疑的两件事,变成了有数字的指控:\n\n- **训练数据规模**:OpenAI 用于中期训练的数据集里,包含超过 91,692 份由纽约时报、Daily News 和 Center for Investigative Reporting 出版的作品副本。仅从 nytimes.com 一家抓出来的 Common Crawl 衍生数据,就含超过 200 万份文档。\n- **Copilot 自伤效应**:微软自己的数据显示,Copilot「答案引擎」让纽约时报域名的点击率相比传统 Bing 搜索最多下降 93%。Hecht 在 2024 年 1 月的另一份内部演示文稿把这称作「末日循环」, 称其将「同时损害我们模型的性能和整个网络」。\n- **双方互相输送数据**:Project Taxi、Project Mango 等代号项目让 OpenAI 与微软互相交送训练素材;Project Mango 数据集至少含 160,903 份来自三家原告出版商的独特作品。\n- **绕过付费墙**:OpenAI 研究员 Nick Ryder 向总裁 Greg Brockman 提到一个「绕过 nytimes 付费墙的黑客手段」, Brockman 回复「ah nice」。\n\n这批解封材料并不是 OpenAI 与微软主动披露,而是《纽约时报》在案情摘要里引用——原始证据仍处于封存状态。换句话说,这些引语脱离了原有语境,公司仍然保有反击空间。\n\n## 个人评论:这案子真正危险的点不在「数据从哪来」\n\n这件事的技术读者容易陷在「他们到底抓了多少」的惊叹里,但真正影响产业链的,是它把 fair use 抗辩里最关键的两根支柱直接拆了。Fair use 法理要求 AI 公司证明「使用没有替代或损害原作品市场」——但现在解封材料里微软自己用 83%-93% 的点击率跌幅、OpenAI 内部把 ChatGPT 描述为「largely substitutive」, 主动承认了「替代效应」。加上 Hecht 备忘录里「an end-product threatens the economic foundations of its essential suppliers」这种字面承认,合理使用抗辩的「不替代」门槛在事实层已经塌了。\n\n对正在做模型训练与产品落地的团队,值得追的不是判罚本身,而是三件事:一是法院会不会把「AI 公司内部对市场伤害的自我认知」纳入 fair use 考量;二是后续和解条款会不会要求训练数据透明披露;三是 Project Mango 这种代号化数据流通会不会被追溯判定为「共谋式侵权」。无论最终怎么判,这批文件已经被多家新闻机构列入下一步诉讼的证据清单,版权战的中段已经翻开。\n\n## 所以呢\n\n如果你是 AI 公司法务或合规岗,现在应该把这件事放进内部演练:训练数据来源文档有没有「我们知道这是偷的」这种带情绪的措辞。如果你是做数据采集或数据集的团队,Project Mango 的命名法与「绕付费墙」的研究员叙述应该成为新的合规红线——它是「明知」的证据点。如果你只是关注 AI 行业的读者,后续真正值得追的不是某一次判决金额,而是 2027 年合理使用边界会不会被这批文件改写。\n\n来源:TechCrunch 报道,Rebecca Bellan 2026 年 9 月 17 日。","https:\u002F\u002Ftechcrunch.com\u002F2026\u002F09\u002F17\u002Fmicrosoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal","226bcb3d-18b8-4bb0-a999-4e82ec13f5fd",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"c33b1bbc-d6ce-4f61-9d5d-1a0704a6a09b","ai-policy",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"295a7fb9-9011-4112-8b1a-1b3cf2ab1d3a","en","NYT vs OpenAI unsealed filings: Microsoft exec admits AI scraping was the 'largest theft of labor in human history'","TechCrunch reported on September 17 that newly unsealed filings in The New York Times' lawsuit against OpenAI and Microsoft reveal Microsoft's director of Applied Science Brent Hecht called AI data scraping 'the largest theft of labor in human history' in a January 2023 internal memo. The filings also disclose that OpenAI's training set contained over 91,692 NYT-affiliated works.","The New York Times' copyright lawsuit against OpenAI and Microsoft has, for the first time in three years, exposed internal statements of this gravity. On September 17, TechCrunch obtained newly unsealed filings showing that Microsoft's director of Applied Science, Brent Hecht, called AI data scraping 'an astonishing theft of unprecedented proportions' and 'the largest theft of labor in human history' in an internal memo dated January 2023. The same filing also disclosed concrete training-set scale and commercial impact — details that until now had only lived in speculation.\n\n## What the filings reveal: training-set scale and commercial harm, both now quantified\n\nWhat this unsealing does is turn two long-disputed claims from copyright litigation into numeric accusations:\n\n- **Training data scale**: OpenAI's mid-training datasets contain more than 91,692 copies of works published by The New York Times, the Daily News, and the Center for Investigative Reporting. A Common Crawl-derived dataset alone pulled more than 2 million documents from nytimes.com.\n- **Copilot's self-harm effect**: Microsoft's own data shows its Copilot 'answer engine' caused click-through rates on The New York Times' domain to drop as much as 93% compared with traditional Bing search. A separate January 2024 internal Microsoft presentation by Hecht called this a 'doom loop' that would 'hurt the performance of our models and the entire web at the same time.'\n- **Bidirectional data flow**: Codenamed projects like Project Taxi and Project Mango moved training material between OpenAI and Microsoft; the Project Mango dataset contained at least 160,903 unique works from the three plaintiff publishers.\n- **Paywall circumvention**: When OpenAI researcher Nick Ryder mentioned 'a hack to get around nytimes paywall' to president Greg Brockman, Brockman reportedly replied: 'ah nice.'\n\nThese unsealed materials were not disclosed by OpenAI or Microsoft themselves — they were quoted by The New York Times in its own brief; the underlying exhibits remain sealed. In other words, the quotes are presented out of original context, and the companies still retain room to push back.\n\n## Personal take: what is genuinely dangerous about this case is not 'where the data came from'\n\nTechnical readers tend to get caught up in the 'how much did they actually scrape' wow factor, but what actually moves the industry is that it has directly dismantled the two most critical pillars of the fair-use defense. Fair-use jurisprudence requires AI companies to show that 'the use does not substitute for or harm the market of the original work' — but the unsealed filings show Microsoft itself citing 83%-93% click-through declines, and OpenAI internally describing ChatGPT as 'largely substitutive,' actively admitting substitution effects. Add to that Hecht's literal admission that 'an end-product threatens the economic foundations of its essential suppliers,' and the 'no substitution' prong of the fair-use defense has collapsed at the factual level.\n\nFor teams currently working on model training and productization, what is worth tracking is not the verdict itself, but three things: first, whether courts will incorporate 'AI companies' internal self-knowledge of market harm' into fair-use analysis; second, whether subsequent settlement terms will require transparency disclosure of training data; third, whether codename-style data flows like Project Mango will be retroactively classified as 'collusive infringement.' Regardless of how the final ruling comes down, this batch of filings has already been added to the evidence list for follow-on litigation by other news organizations — the middle innings of the copyright war have turned a page.\n\n## So what\n\nIf you are AI company legal or compliance staff, you should now put this scenario into internal drills: do your training-data source documents contain emotive phrasing like 'we know this is theft'? If you are on a data-acquisition or dataset team, Project Mango's naming convention and the researcher narration of 'paywall circumvention' should become the new compliance red lines — they are evidence points for 'knowing' intent. If you are merely a reader following the AI industry, what is worth tracking in the months ahead is not a particular damages number, but whether the boundary of fair use will be redrawn in 2027 by this batch of filings.\n\nSource: TechCrunch reporting by Rebecca Bellan, September 17, 2026.","nyt-openai-microsoft-hecht-largest-theft-of-labor","2026-09-23T03:00:00Z","2026-09-23T03:05:40.883851Z","2026-09-23T03:05:40.883863Z",true,"agent",36,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"7fce8217-577f-4fd5-8f88-a1566cbf1290","微软与 OpenAI 法庭文件解封:LLM 训练数据被自家高管称为史上最大劳动窃取","microsoft-openai-doom-loop-nyt-copyright-2026","2026-09-23T05:03:20+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"d569d88a-8906-4f3b-83fd-0dcc6b5c75e0","OpenAI 用 __obi 把 ChatGPT 账号绑上你全网浏览","openai-obi-cookie-cross-site-tracking-chatgpt","2026-09-23T07:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"9dd4a859-1153-4ecc-b69d-4ba4c5431129","智谱被开发者抓包后紧急上线数据零留存","zhipu-maas-zero-data-retention-zcode","2026-09-21T07:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"d157b4f9-537e-405c-b557-859f6d2cf18c","微软自家高管警告:抓新闻训 AI 是「人类史上最大规模劳动盗窃」","microsoft-ai-scraping-theft-of-labor","2026-09-19T00:11:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"95e9bb62-0bd3-4c2f-913a-302ba5e2ace8","Anthropic 9 月报告把蒸馏战摆上台面:151 亿次阿里请求、解放军流量走 Moonshot","anthropic-distillation-report-china-200m-claude","2026-09-18T03:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"1d113d73-3774-426a-bdc0-49c678a96a59","Bengio 长文复盘:AI 智能体说谎作弊,病根在训练目标打架","bengio-ai-agents-misalignment","2026-09-14T17:10:00+00:00"]