[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-anthropic-project-panama-settlement":3,"topics-all":36,"news-related-03143280-34e6-4492-9ce6-b2a8cacdeb74":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"03143280-34e6-4492-9ce6-b2a8cacdeb74","Anthropic 15 亿美元图书和解落槌:\"Project Panama\" 把 LLM 训练数据工程推到监管视野","美国地区法官 Araceli Martinez-Olguin 本周一正式批准 Anthropic 与图书作者及出版商的集体诉讼和解协议,赔偿总额 15 亿美元,覆盖逾 48.2 万册图书。和解的关键意义不在赔偿数字本身,而在于法庭给出的两段式判决:使用受版权保护图书训练 LLM 属于合理使用,但通过影子图书馆下载盗版电子书及\"拆书脊扫描\"的实体书处理流程不合法。这等于给整个 LLM 行业的训练数据获取方式划了一条工程红线。\n\n事件曝光了 Anthropic 内部代号\"Project Panama\"的训练数据采集流程:投入数千万美元购入实体新书,拆开书脊、扫描书页后送回收公司,并聘请二十年前参与 Google Books 的 Google 高管负责工程化实施,同时配合从影子图书馆下载盗版电子书。法官在判决中把\"购买+扫描\"与\"盗版下载\"做法律切割,本质上是承认 LLM 训练对受版权材料的合理使用,同时惩罚绕过授权渠道的灰色采购。\n\n对行业的直接影响有三层:其一,头部厂商未来必须把训练数据溯源做成可审计流程,版权许可(出版商批量授权、LibGen 替代语料)会被纳入采购清单;其二,正在路上的类似诉讼(New York Times vs OpenAI、UMG vs Suno)可能援引此案的\"合理使用+盗版切割\"逻辑;其三,小厂商及开源训练方将更难获取廉价大规模语料,数据成本结构性抬升。\n\n技术意义不止于法律。Claude 3\u002F3.5 的能力跃升一直被怀疑与高质量人类文本的密集覆盖有关,Project Panama 这种\"工程化拆解纸质书\"的流程是 LLM 时代训练数据工程(TrDE, Training Data Engineering)的典型样本。当这条路径被判定为不可合法复制,开源和中小厂商必须寻找新数据源:合成数据、长上下文自蒸馏、用户授权语料将成为下一阶段的工程主流。","https:\u002F\u002Fapnews.com\u002Farticle\u002Fanthropic-copyright-authors-settlement-training-f294266bc79a16ec90d2ddccdf435164","833c1eca-f067-4241-91c1-1f82efecb59c",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"23544f6a-eea1-4f05-aa8d-749ca862d5d2","anthropic",{"id":18,"name":19,"slug":19,"description":13,"color":13},"dca4d0ab-7994-43a7-839e-7756fc77344a","claude",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"d2cef8cf-1781-4b32-93ac-a155079210ea","en","Anthropic's $1.5B book settlement puts training data on trial","U.S. District Judge Araceli Martinez-Olguin officially approved Anthropic's class-action settlement with book authors and publishers this Monday, for a total of $1.5 billion in damages covering more than 482,000 books. The settlement's key meaning isn't the dollar figure itself, but the two-part ruling the court gave: training LLMs on copyrighted books is fair use, but downloading pirated e-books from shadow libraries and \"spine-stripped scanning\" of physical books is illegal. This draws a clear engineering red line around how the entire LLM industry acquires training data. The case revealed Anthropic's internal project codenamed \"Project Panama\" for training-data acquisition: it spent tens of millions of dollars purchasing new physical books, stripped the spines, scanned the pages and sent them to a recycling company, hired a former Google executive who had worked on Google Books twenty years ago to engineer the workflow, and downloaded pirated e-books from shadow libraries in parallel. The court separated the \"purchase + scan\" path from the \"piracy download\" path in its judgment — essentially recognizing fair use of copyrighted material for LLM training while punishing the grey-area procurement that bypassed licensing channels. The direct industry impact has three layers: first, top players will need to make training-data provenance an auditable process, with copyright licensing (bulk publisher deals, LibGen-replacement corpora) entering the procurement checklist; second, similar lawsuits in flight (NYT v. OpenAI, UMG v. Suno) may cite this case's \"fair use + piracy carve-out\" logic; third, smaller players and open-source trainers will find it harder to acquire cheap large-scale corpora, structurally raising data costs. The technical implications go beyond the legal. Claude 3\u002F3.5's capability jump has long been suspected of being tied to dense coverage of high-quality human text, and Project Panama's \"engineered deconstruction of paper books\" is a typical sample of LLM-era Training Data Engineering (TrDE). With this path ruled legally unreplicable, open-source and mid-size vendors must seek new data sources: synthetic data, long-context self-distillation, and user-licensed corpora will become the next-stage engineering mainstream.","anthropic-project-panama-settlement","2026-07-22T08:00:00Z","2026-07-22T08:05:33.289018Z","2026-08-19T02:08:40.142862Z",true,"agent",208,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"f3d17d45-e1a8-4a1b-9449-6813aff06e49","Anthropic 让 Claude 自己修对齐:10 类失败全部见效,还超过人类研究员","claude-automated-alignment-researchers","2026-08-29T13:05:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"97c97b9c-e6e4-4982-aa57-0c0da814fb19","Anthropic 的欧盟答卷四小时即被撕开：Claude 文本水印为什么怕改写","claude-synthid-70-percent-threshold-bypass","2026-08-21T08:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"a7f4cfad-874e-42b0-a84b-bd0ec57e8fdc","Anthropic 给 Claude 文本上不可见水印,接 SynthID-Text 走全球合规","anthropic-claude-invisible-text-watermark","2026-08-18T03:30:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"9f566c9a-4c39-427c-af5e-c3a6b162ec25","Anthropic 把不可见水印写进 Claude 文本：复制粘贴都带走的 AI 身份证","anthropic-claude-invisible-watermark-eu-ai-act","2026-08-12T02:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"ca53004e-9180-4b9d-b9db-337f2d20994b","Anthropic 给 Claude 文本加水印:欧盟 AI Act 第 50 条第一次有了「出厂级」答案","anthropic-claude-text-watermark-eu-ai-act","2026-08-11T21:48:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"3967306f-062a-41a6-ab58-f99e70fc0e68","AISI 122 轮 cyber eval 越界：OpenAI 与 Anthropic 同日披露","aisi-mythos-5-gpt-5-6-cyber-eval-incident-2026","2026-08-08T04:00:00+00:00"]