Amazon Scans and Destroys Books for AI Training
A tracked shipment exposed an Amazon operation that buys, disassembles and scans physical books to obtain AI-training material.
A hidden supply chain for training data
An investigation by 404 Media has traced a large order of rare and out-of-print books to an Amazon facility in Las Vegas where physical volumes are reportedly dismantled and scanned for AI development. The reporting provides an unusual look at how a major technology company is sourcing copyrighted material beyond publicly accessible internet data.
The investigation began after a bookseller received an order for roughly 1,000 titles through Biblio, a marketplace that can conceal a buyer’s identity. With the seller’s cooperation, reporters placed an AirTag inside one book and followed the shipment to an Amazon warehouse. Workers cited in the report described an operation in which bindings are removed so pages can be scanned efficiently, destroying the physical copies in the process.
Amazon confirmed that it purchases books through commercial channels to help develop and improve products and services, but did not publicly identify the models, datasets or products that use the resulting scans. The company’s purchase of lawful copies gives the practice a different legal profile from scraping unauthorized digital libraries, although copyright questions surrounding model training and the retention of source material remain contested.
Why it matters
The operation shows that competition for high-quality training data is reaching into physical archives. Rare and out-of-print books contain material that may be absent from ordinary web datasets, making them potentially valuable as laboratories search for cleaner, more diverse text.
Buying copies does not resolve every concern. Destructive scanning may remove scarce works from circulation, while anonymous bulk purchasing makes it difficult for booksellers, authors and collectors to understand where culturally valuable material is going. The lack of disclosure about downstream datasets also limits independent scrutiny of consent, provenance and duplication.
For the AI industry, the investigation shifts the data debate from abstract web scraping to a tangible supply chain involving marketplaces, warehouses and destroyed objects. It may prompt libraries, dealers and policymakers to examine whether existing rules adequately protect both intellectual property and the physical preservation of scarce texts.