[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-microsoft-ai-scraping-theft-of-labor":3,"topics-all":38,"news-related-d157b4f9-537e-405c-b557-859f6d2cf18c":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"d157b4f9-537e-405c-b557-859f6d2cf18c","微软自家高管警告:抓新闻训 AI 是「人类史上最大规模劳动盗窃」","微软应用科学总监 Brent Hecht 内部文件外泄,称抓取新闻训 AI 是人类史上最大劳动成果盗窃。微软数据:新闻出版商点击率下降 51%-94%,部分站点 83%-93%。微软 CEO 纳德拉国会作证承认 AI 不该绕过付费墙,OpenAI 总裁 Brockman 内部却被曝对爬虫绕过《纽约时报》付费墙说「好极了」。","## 「theft of labor」来自微软内部 PPT\n\n最近 Ars Technica 拿到一份微软内部文件,主角是微软应用科学总监 Brent Hecht。他的措辞相当激烈:抓取新闻内容训练 AI 是「人类历史上最大规模的劳动成果盗窃」——他在文件里用了「theft of labor」这个表述,而且明确说「这完全是对合理使用理念的嘲弄」。一家全球市值前几的科技公司,在内部文件里承认自己最赚钱的合作伙伴(OpenAI)所用训练数据是「盗窃」,这件事本身就够戏剧了,但更狠的是数字。\n\n微软自己的数据指出,部分新闻出版商网站自 ChatGPT 上线以来,点击率下降了 83% 到 93%;另一批新闻机构的点击率下降了 51% 到 94%。换句话说,新闻业不仅在被偷内容,还在被偷流量。Hecht 在文件里承认「几乎没有人希望自己创作的内容被以这种方式使用,也几乎没有人因此获得报酬」,但这句话是写给内部同事看的——他真正想推的论点是,这条链路会形成「恶性循环」,既伤害模型,也伤害整个 Web。当原创内容方因为没收入、没流量而停止生产,模型可吃的语料也会枯竭。\n\n## 纳德拉国会承认 vs Brockman Slack 里说「好极了」\n\n最戏剧性的不是这份内部文件,而是接下来发生的事情。微软 CEO 萨蒂亚·纳德拉(Satya Nadella)在国会作证时亲口承认,AI 公司不应该通过绕过付费墙的方式违反新闻网站的使用条款——这相当于在公开场合重复了 Brent Hecht 在内部文件里的立场。但根据 Ars Technica 拿到的 OpenAI 内部信息,当一名 OpenAI 员工把「我们爬虫找到了绕过《纽约时报》付费墙的漏洞」这件事通知总裁 Greg Brockman 时,得到的回复是「好极了」。\n\n一边的内部文件说这是盗窃,一边的国会证词承认这是违约,而真正动手绕过付费墙的另一头,高管在内部 Slack 里为这个 bug 庆祝。这不是单点的失误,是头部公司之间长期在「嘴上承认问题、行动上继续薅」的拉扯。如果连 OpenAI 的最大投资人和算力供应商微软都承认「不该这么干」,那剩下的辩护空间还能有多大?(详见 Ars Technica 原报道:[Microsoft exec called AI scraping \"the largest theft of labor in human history\"](https:\u002F\u002Farstechnica.com\u002Ftech-policy\u002F2026\u002F09\u002Fmicrosoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history\u002F))\n\n所以 Hecht 这份文件的真正信号不是「AI 公司用新闻训练」这件事本身,而是这件事的代价已经被量化,而且被量化在**微软自己**的内部 PPT 里——点击率下降 83% 到 94% 这种区间,不是某个外部机构的怀疑,而是当事方的内部测算。当这种数字第一次以「内部文件外泄」的形式出现在公开报道里,它就不再只是伦理议题,而是会被监管方和法院拿来当证据用的事实基础。","https:\u002F\u002Farstechnica.com\u002Ftech-policy\u002F2026\u002F09\u002Fmicrosoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history\u002F","2af9d198-9418-4f26-85e4-4a8f3eede35a",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"c33b1bbc-d6ce-4f61-9d5d-1a0704a6a09b","ai-policy",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"1e2a3eac-cd69-4ac8-af6b-b68b62e164e1","en","Microsoft: AI scraping is \"the largest theft of labor\"","Leaked Microsoft doc: scraping news to train AI is the largest theft of labor. Publisher CTR fell 83-93%. Nadella and Brockman disagree on paywalls.","## \"Theft of labor\" leaked from a Microsoft internal PPT\n\nArs Technica recently obtained a Microsoft internal document whose main author is Brent Hecht, Microsoft's Director of Applied Science. His wording was unusually blunt: scraping news content to train AI is \"the largest theft of labor in human history\" — he used the phrase \"theft of labor\" in the file and explicitly said this \"completely mocks the concept of fair use.\" When one of the world's most valuable tech companies acknowledges, in its own internal document, that the training data of its most profitable partner (OpenAI) is \"theft,\" the situation is dramatic on its own — but the numbers are even harsher.\n\nMicrosoft's own data shows that since ChatGPT launched, click-through rates to some news publisher sites dropped by 83% to 93%; another batch of news organizations saw declines of 51% to 94%. In other words, the news industry is not just having its content stolen — its traffic is being stolen too. Hecht admits in the document that \"almost no one wants their content used this way, and almost no one is paid for it\" — but this sentence was meant for internal colleagues. The argument he really wanted to push is that this pipeline creates a \"vicious cycle\" that hurts both the model and the entire Web. When original content creators stop producing because they have no revenue and no traffic, the model's available corpus will also dry up.\n\n## Nadella admits in Congress vs Brockman says \"cool\" in Slack\n\nThe most dramatic part isn't the internal document — it's what came next. Microsoft CEO Satya Nadella testified before Congress and personally acknowledged that AI companies should not violate news sites' terms of service by bypassing paywalls. That's effectively repeating Brent Hecht's internal position in a public forum. But according to OpenAI internal information obtained by Ars Technica, when an OpenAI employee notified president Greg Brockman that \"our crawler found a way to bypass the New York Times paywall,\" the reply was \"cool.\"\n\nOne side's internal document calls it theft, another side's Congressional testimony admits it's a breach, and the side actually bypassing the paywall has its executives celebrating the bug in an internal Slack. This isn't a one-off mistake — it's a long-running tug-of-war among top companies that \"admit the problem with words and keep scraping with actions.\" If even Microsoft, OpenAI's largest investor and compute supplier, admits \"this shouldn't be done,\" how much defense space is left? (See the original Ars Technica report: [Microsoft exec called AI scraping \"the largest theft of labor in human history\"](https:\u002F\u002Farstechnica.com\u002Ftech-policy\u002F2026\u002F09\u002Fmicrosoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history\u002F))\n\nSo the real signal from Hecht's document isn't the fact that \"AI companies train on news\" — it's that the cost of this has now been quantified, and quantified inside Microsoft's own internal PPT. A click-through decline range of 83% to 94% is not an external researcher's suspicion; it's the involved party's internal estimate. When such numbers first surface in public reporting through leaked internal documents, the discussion is no longer purely an ethical debate — it becomes factual groundwork that regulators and courts can use as evidence.","microsoft-ai-scraping-theft-of-labor","2026-09-19T00:11:00Z","2026-09-19T07:08:42.175205Z","2026-09-19T07:08:42.175213Z",true,"agent",108,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"2fc4f696-a697-498c-b9d5-28250bfeaa79","ChatGPT 进欧盟 VLOP 名单:OpenAI 第一次要为生成式 AI 内容负全责","chatgpt-eu-vlop-dsa-first-ai-platform-rules","2026-09-05T07:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"1e43b4fc-39fe-4cac-aaf9-57f82d5c0311","AI 公司与数学界「错位」:两个月三次刷屏,把同行评审甩在身后","ai-math-severe-misalignment-fields-medal","2026-09-15T10:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"982c5e1e-5274-442e-9237-abaf39e8ee3c","25 位菲尔茨奖得主联名公开信:AI 解题竞赛正在伤害数学","fields-medalists-ai-misalignment-math","2026-09-13T13:07:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"6b349d2c-3d03-4cb9-8f47-63e8c288d0db","美方三机构联合指控六家中国 AI 企业系统性蒸馏美国模型","us-accuses-six-chinese-ai-firms-of-distillation","2026-09-10T01:08:40+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"d17a841b-abca-46e0-80e4-d955f1c837ba","亚马逊 VGT3 仓库曝光:一天拆掉上千本书,只为给 AI 模型喂语料","amazon-vgt3-warehouse-ai-training-books","2026-09-07T03:30:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"1942b07b-f794-42b1-b944-ca6b32d4ae16","四大 AI 模型同日集体掉线:OpenAI\u002FClaude 官方确认,Gemini\u002FGrok 表面沉默","four-ai-models-overlapping-outage-sept-2026","2026-09-06T08:00:00+00:00"]