404 Media has published an anonymous interview with an employee at Amazon's VGT3 warehouse in Las Vegas, Nevada. Every day, the facility receives thousands of brand-new and used books, workers manually cut the spines off, scan the loose pages into a corpus, and then discard the physical pages irreversibly. The employee revealed that Amazon initially told staff the operation was for Kindle digitization, masking the fact that the pages were being used to build AI training data. The story exposes the real physical cost of frontier LLM corpus sourcing and the gray zone of copyright compliance.

What the warehouse does every day

This is not an outside observer's account. The 404 Media interview, published by Emanuel Maiberg on August 26, 2026, is from someone who actually stands on the floor of VGT3, feeds books into the cutting machines one by one.

The warehouse shares a campus in Las Vegas with LAS8, Amazon's print-on-demand facility. The interview describes a remarkably specific workflow:

Workers receive pallets of books in many languages — German, Russian, Japanese, including sealed new copies and liquidated library stock from the University of London. They even saw bound documents labeled as presented to Parliament on behalf of Her Majesty. The flow: scan barcodes to deduplicate → feed books into a manual cutter with a safety guard → push loose pages to the next station → image them on roughly 20 to 25 scanners "that look like cash-counting machines" → throw the scanned pages into gaylords (open cardboard boxes about 6 to 7 feet tall) from which the books can never be reassembled.

The employee told 404 Media something that matters: Amazon initially said this was digitizing for Kindle. "I instantly felt that was wrong, given publication rights and copyright." Later they learned the real destination was an AI training dataset.

Why this is worth talking about

The warehouse details are not technically deep, but they capture a question that has been hanging in the air, in an observable way, for the first time: how much of the corpus feeding frontier LLMs is paper books with clearly defined copyright status?

The moment the spine is cut, the book stops being a book. It becomes a pile of scanned pages with no cover, no copyright page, no ISBN. Duplicates "are returned to vendors" per the employee — but the de-duplication is barcode-based, not copyright-based. An academic monograph still under US copyright because it was re-published after 1923 ends up in the cutter the same way a 1900 public-domain book does.

The "deliberate opacity" of the process is also revealing. The employee described: "Their process changed every day. Even the people who work there probably don't know what they're doing — or probably just don't want to talk about it." When an execution layer is deliberately isolated from the answer to "why are we doing this," the upstream usually does not want the execution layer to bear informed responsibility.

Doing the math

404 Media's earlier tracking piece had already established that VGT3's role includes supplying AI companies with training data. The interview also notes that "some of these books look like they might be rare," and that "we've heard they order rare books." In other words, Amazon may be buying rare books that are still circulating in the open market, not just clearing inventory.

At a conservative order of magnitude — VGT3 employees describe routine daily throughput of thousands to tens of thousands of books — that is millions to tens of millions of physical books per year, irreversibly converted into single-purpose training samples for one AI model. For a rare book that might retail for tens to hundreds of dollars, this is unreproducible resource consumption.

The "gray edge" tradition of training data

This is not an isolated incident. Anthropic was sued by Sony and Warner Music at the end of August, with the complaint alleging it used BitTorrent to download over 5 million pirated books for Claude training, plus over 2 million more from the Pirate Library Mirror. OpenAI has been repeatedly accused of drawing on shadow libraries. VGT3, if pursued seriously, points to a different question: a path that does not go through P2P or pirate sites, but through legitimate commercial procurement that physically destroys books with ambiguous copyright status before they enter the corpus — does this constitute the same kind of substantive infringement of authors' rights?

US copyright law's "fair use" boundary for AI training is still being fought out in pending cases including New York Times v. OpenAI. But VGT3 at least proves one thing: the cost of corpus procurement is no longer just server bills and lawyer fees. It now includes irreversible physical consumption of books.

So what

If you are an AI engineer cleaning corpora: prioritize confirming the source chain. Even if your vendor says "copyright already cleared," keep asking about the physical origin of each batch — is it from public-domain digitization, author licensing, or from some "book-cutting warehouse"?

If you are a procurement officer or investor: training data compliance is moving from the legal department's gray area into a question you have to ask explicitly in the procurement flow. A supply chain that can bulk-buy books on the open market, irreversibly destroy them, and disclose no explicit copyright status will eventually become the next target for class actions.

If you are an ordinary reader: be aware that rare and out-of-print books currently flowing through your channels may be being consumed at a rate of thousands per day into non-reusable training samples. This is not alarmism. It is the present reality described by the VGT3 employee on the record.