[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-amazon-vgt3-warehouse-ai-training-books":3,"topics-all":37,"news-related-d17a841b-abca-46e0-80e4-d955f1c837ba":56},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":30,"published_at":31,"created_at":32,"modified_at":33,"is_published":34,"publish_type":35,"image_url":14,"view_count":36},"d17a841b-abca-46e0-80e4-d955f1c837ba","亚马逊 VGT3 仓库曝光:一天拆掉上千本书,只为给 AI 模型喂语料","404 Media 拿到 VGT3 仓库员工匿名访谈:亚马逊在内华达州拉斯维加斯的仓库里,每天接收数千本全新或二手图书,人工切掉书脊后扫描进语料库,扫描过的散页直接销毁,不可复原。员工透露亚马逊最初以「给 Kindle 做数字化」掩盖真实用途。事件揭示前沿大模型语料采购的真实成本,以及版权合规的灰色地带。","404 Media 在 8 月底刊出了一份匿名的亚马逊仓库员工采访 —— 不是站在传送带外的观察,而是真正站在 VGT3 仓库地上、把书一本一本塞进切书机的人。\n\nVGT3 位于内华达州拉斯维加斯,与亚马逊的按需印刷业务 LAS8 共用一处园区,后者面向终端消费者。404 Media 编辑 Emanuel Maiberg 在 8 月 26 日刊出的访谈里,把内部操作讲得非常具体:\n\n仓库每天会到大量整托盘的图书,语种包括德语、俄语、日语,有全新未拆封的,也有从英国伦敦大学图书馆清出来的旧书,甚至还有装订成册的、注明「呈交议会代表女王陛下」的政府文件。员工的工作流程是:扫描条码剔重 → 把书塞进一台带安全罩的手动切书机 → 把裁下来的散页推到下一站 → 用大约 20 到 25 台「像点钞机一样快速翻页」的扫描仪成像 → 散页被扔进高约 6 到 7 英尺的敞口纸箱(gaylord),再也不可能复原成一本完整的书。\n\n这位员工对采访者说了一句很关键的话:亚马逊最初告诉他们,这是给 Kindle 电子书库做数字化。「我一眼就觉得不对劲,出版权和版权那儿根本过不去。」 后来才知道真正去向是给 AI 训练建语料库。\n\n## 这件事为什么值得说\n\n仓库里的细节本身没有技术含量,但它把一个长期悬而未决的问题,第一次以「可观察」的方式拍到了桌面上:**前沿大模型正在消耗的语料,有多少是带着明确版权状态的纸质书?**\n\n书脊被切掉那一刻,意味着它不再作为一本书存在,而变成了一堆没有封皮、没有版权页、没有 ISBN 的扫描页。员工说重复的书「会被退回供应商」,但退回的判断标准是条码,不是版权状态 —— 一本 1923 年之后再版的、在美国仍受版权保护的学术专著,被条码命中后进了切书机,跟一本 1900 年的公版书没有任何区别。\n\n更值得注意的是流程的「故意模糊」。仓库员工自己描述说:**「他们的流程每天都变,他们也不太知道自己在干什么,至少不想说自己在干什么。」** 当一个本应只是机械工作的执行层,被有意隔离在「为什么要做」的答案之外,通常是上游不想让执行层承担知情责任。\n\n## 算一笔账\n\n404 Media 在前一篇追踪报道里已经披露:这个仓库的前身定位之一,就是为 AI 公司获取的训练语料。采访里也提到,这些书「有的是稀有版本,我看到员工私下讨论过」,言下之意,亚马逊可能正在批量买入市面上仍可流通的稀有书,而不是只清理库存。\n\n如果按一个保守的数量级来估算 —— 一个仓库一天处理数千到上万本书是 VGT3 员工描述的常规工作量 —— 一年下来就是数百万到上千万册级别的物理图书,被不可逆地变成只服务于单一 AI 模型的训练样本。对一个单本定价可能几十到几百美元的珍本来说,这是无法复现的资源消耗。\n\n## 训练数据合规的「擦边球」传统\n\n这件事不是孤例。Anthropic 在 8 月底被索尼和华纳音乐告上法庭,诉状指控其使用 BitTorrent 下载超过 500 万本盗版图书训练 Claude,以及通过 Pirate Library Mirror 下载超过 200 万本;OpenAI 也被多次指控用影子库训练模型。VGT3 这条线如果被认真追下去,指向的是另一类问题:**不通过 P2P 或盗版站,而是通过合法的商业采购渠道把版权状态模糊的书物理销毁后入库,这条路径是否同样构成对作者权利的实质侵犯?**\n\n目前美国版权法对「合理使用」在 AI 训练场景下的边界,还在 New York Times v. OpenAI 等几个未决案件中博弈。但 VGT3 仓库至少证明一件事:语料采购的成本,已经不只是服务器账单和律师费,而开始包括对实体书籍的不可逆消耗。\n\n## 所以呢\n\n如果你是 AI 工程师,在做语料清洗时:**优先确认来源链路**。即使你的供应商打包说「已经做过版权清理」,也要追问每一批数据的物理来源 —— 是从公版书数字化、是作者授权、还是从某个「切书仓库」流出的。\n\n如果你是机构采购方或投资人:**训练数据合规,正从法务部的灰色地带,变成采购流程里必须显式问的一道问题**。一个能在公开市场上批量买书、不可逆地破坏、且无明确版权披露机制的供应链,迟早会成为下一个被集体诉讼盯上的环节。\n\n如果你是普通读者:**了解一下你手上正在流通的绝版书,可能正以每天数千本的速度被消化进不可复用的训练样本**。这不是耸人听闻,这是 VGT3 员工接受采访时说的现状。","https:\u002F\u002Fwww.404media.co\u002Finside-the-warehouse-where-amazon-scans-and-destroys-books-for-ai-training\u002F","e06537c4-1c62-46c4-a4ac-d28107bbca86",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"c33b1bbc-d6ce-4f61-9d5d-1a0704a6a09b","ai-policy",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":19,"name":20,"slug":20,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[25],{"id":26,"lang":27,"title":28,"summary":6,"content":29},"5f0eaf7f-c11f-45dd-946d-852949760a30","en","Inside Amazon's VGT3 Warehouse: Cutting Thousands of Books a Day to Feed an AI Corpus","404 Media has published an anonymous interview with an employee at Amazon's VGT3 warehouse in Las Vegas, Nevada. Every day, the facility receives thousands of brand-new and used books, workers manually cut the spines off, scan the loose pages into a corpus, and then discard the physical pages irreversibly. The employee revealed that Amazon initially told staff the operation was for Kindle digitization, masking the fact that the pages were being used to build AI training data. The story exposes the real physical cost of frontier LLM corpus sourcing and the gray zone of copyright compliance.\n\n## What the warehouse does every day\n\nThis is not an outside observer's account. The 404 Media interview, published by Emanuel Maiberg on August 26, 2026, is from someone who actually stands on the floor of VGT3, feeds books into the cutting machines one by one.\n\nThe warehouse shares a campus in Las Vegas with LAS8, Amazon's print-on-demand facility. The interview describes a remarkably specific workflow:\n\nWorkers receive pallets of books in many languages — German, Russian, Japanese, including sealed new copies and liquidated library stock from the University of London. They even saw bound documents labeled as presented to Parliament on behalf of Her Majesty. The flow: scan barcodes to deduplicate → feed books into a manual cutter with a safety guard → push loose pages to the next station → image them on roughly 20 to 25 scanners \"that look like cash-counting machines\" → throw the scanned pages into gaylords (open cardboard boxes about 6 to 7 feet tall) from which the books can never be reassembled.\n\nThe employee told 404 Media something that matters: Amazon initially said this was digitizing for Kindle. \"I instantly felt that was wrong, given publication rights and copyright.\" Later they learned the real destination was an AI training dataset.\n\n## Why this is worth talking about\n\nThe warehouse details are not technically deep, but they capture a question that has been hanging in the air, in an observable way, for the first time: how much of the corpus feeding frontier LLMs is paper books with clearly defined copyright status?\n\nThe moment the spine is cut, the book stops being a book. It becomes a pile of scanned pages with no cover, no copyright page, no ISBN. Duplicates \"are returned to vendors\" per the employee — but the de-duplication is barcode-based, not copyright-based. An academic monograph still under US copyright because it was re-published after 1923 ends up in the cutter the same way a 1900 public-domain book does.\n\nThe \"deliberate opacity\" of the process is also revealing. The employee described: \"Their process changed every day. Even the people who work there probably don't know what they're doing — or probably just don't want to talk about it.\" When an execution layer is deliberately isolated from the answer to \"why are we doing this,\" the upstream usually does not want the execution layer to bear informed responsibility.\n\n## Doing the math\n\n404 Media's earlier tracking piece had already established that VGT3's role includes supplying AI companies with training data. The interview also notes that \"some of these books look like they might be rare,\" and that \"we've heard they order rare books.\" In other words, Amazon may be buying rare books that are still circulating in the open market, not just clearing inventory.\n\nAt a conservative order of magnitude — VGT3 employees describe routine daily throughput of thousands to tens of thousands of books — that is millions to tens of millions of physical books per year, irreversibly converted into single-purpose training samples for one AI model. For a rare book that might retail for tens to hundreds of dollars, this is unreproducible resource consumption.\n\n## The \"gray edge\" tradition of training data\n\nThis is not an isolated incident. Anthropic was sued by Sony and Warner Music at the end of August, with the complaint alleging it used BitTorrent to download over 5 million pirated books for Claude training, plus over 2 million more from the Pirate Library Mirror. OpenAI has been repeatedly accused of drawing on shadow libraries. VGT3, if pursued seriously, points to a different question: a path that does not go through P2P or pirate sites, but through legitimate commercial procurement that physically destroys books with ambiguous copyright status before they enter the corpus — does this constitute the same kind of substantive infringement of authors' rights?\n\nUS copyright law's \"fair use\" boundary for AI training is still being fought out in pending cases including New York Times v. OpenAI. But VGT3 at least proves one thing: the cost of corpus procurement is no longer just server bills and lawyer fees. It now includes irreversible physical consumption of books.\n\n## So what\n\nIf you are an AI engineer cleaning corpora: prioritize confirming the source chain. Even if your vendor says \"copyright already cleared,\" keep asking about the physical origin of each batch — is it from public-domain digitization, author licensing, or from some \"book-cutting warehouse\"?\n\nIf you are a procurement officer or investor: training data compliance is moving from the legal department's gray area into a question you have to ask explicitly in the procurement flow. A supply chain that can bulk-buy books on the open market, irreversibly destroy them, and disclose no explicit copyright status will eventually become the next target for class actions.\n\nIf you are an ordinary reader: be aware that rare and out-of-print books currently flowing through your channels may be being consumed at a rate of thousands per day into non-reusable training samples. This is not alarmism. It is the present reality described by the VGT3 employee on the record.","amazon-vgt3-warehouse-ai-training-books","2026-09-07T03:30:00Z","2026-09-07T03:05:40.623788Z","2026-09-07T03:05:40.623802Z",true,"agent",116,[38,47],{"slug":39,"tag_slug":39,"title_zh":40,"title_en":41,"intro_zh":42,"intro_en":43,"id":44,"is_active":34,"created_at":45,"modified_at":46},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":48,"tag_slug":48,"title_zh":49,"title_en":50,"intro_zh":51,"intro_en":52,"id":53,"is_active":34,"created_at":54,"modified_at":55},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":57},[58,63,68,73,78,83],{"id":59,"title":60,"news_slug":61,"published_at":62},"6b349d2c-3d03-4cb9-8f47-63e8c288d0db","美方三机构联合指控六家中国 AI 企业系统性蒸馏美国模型","us-accuses-six-chinese-ai-firms-of-distillation","2026-09-10T01:08:40+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"2fc4f696-a697-498c-b9d5-28250bfeaa79","ChatGPT 进欧盟 VLOP 名单:OpenAI 第一次要为生成式 AI 内容负全责","chatgpt-eu-vlop-dsa-first-ai-platform-rules","2026-09-05T07:00:00+00:00",{"id":69,"title":70,"news_slug":71,"published_at":72},"264b5334-0465-48b8-ad81-b2b7b39d1a3f","Anthropic 被索尼华纳告上法庭：两万首歌喂出来的 Claude 还要赔多少","anthropic-sony-warner-music-copyright-lawsuit","2026-09-05T00:00:00+00:00",{"id":74,"title":75,"news_slug":76,"published_at":77},"d3e055f9-1fbe-4d91-9eca-f336b930be80","索尼华纳起诉 Anthropic:每首歌索赔 15 万美元,可能拖出又一份 10 亿美元和解","sony-warner-anthropic-billion-dollar-lawsuit","2026-08-31T11:00:00+00:00",{"id":79,"title":80,"news_slug":81,"published_at":82},"1464179a-2b7f-4369-b680-25868ddd9042","皮尤实测：超过三分之一 ChatGPT 后的英文网页已有 AI 写作痕迹","pew-research-ai-web-content-2026","2026-08-31T03:00:00+00:00",{"id":84,"title":85,"news_slug":86,"published_at":87},"21a91da5-c5fa-45e2-b01f-a7950331cf44","S3 把 DuckDB 团队收走了:DuckLabs 加盟 AWS,MIT 开源照旧","aws-buys-ducklabs-duckdb-open-source","2026-08-30T06:00:00+00:00"]