IA · 18 August 2026 · 5 min read

From Online Bookseller to AI Shredder: AirTag Uncovers Amazon Destroying Rare Books for Model Training

In brief: An investigative report using an Apple AirTag planted in a commercial shipment has traced bulk purchases of rare and out-of-print books directly to an Amazon data processing facility in Las Vegas. The company cuts off the book spines to feed pages into high-speed scanners at scale, hunting for pristine pre-2022 human text to avoid model collapse. The findings have ignited severe ethical backlash over the destruction of physical literary heritage.

by Team Mocchi's

From Online Bookseller to AI Shredder: AirTag Uncovers Amazon Destroying Rare Books for Model Training

For over a year, antiquarian booksellers and archivists suspected that major tech players were quietly purchasing large batches of out-of-print, niche, and rare volumes through commercial intermediaries. Concrete evidence has finally emerged thanks to an investigative effort: an Apple AirTag concealed within a rare book shipment tracked the consignment straight to an Amazon-operated facility in Las Vegas.

As reported by Ars Technica, the tracker led 404 Media journalists to the facility designated internally as VGT3. Inside the warehouse, dedicated teams guillotine the spines off physical books to feed loose pages into high-speed commercial scanners, destroying the physical volumes in the process. The unit's unofficial logo on the warehouse door features a Tyrannosaurus rex preparing to devour an open book.

When questioned regarding the findings, Amazon did not deny the operation or the use of destructive scanning methods, providing a boilerplate statement stating that it purchases books through commercial channels to help develop and improve products and services for its customers.

The scramble for pre-2022 pristine tokens

Amazon's focus on physical books reflects a structural bottleneck confronting frontier AI research: data exhaustion. As highlighted by TechCrunch, large parts of the public web are now heavily contaminated with synthetic text generated by earlier language models. Training new models on synthetic or recursive web data risks model collapse, a degradation of factual accuracy and reasoning capabilities.

Physical books printed prior to 2022 represent a coveted reservoir of pristine, highly structured human text free from synthetic pollution. Discussions on internal and worker forums revealed that the VGT3 facility had faced a critical supply drought earlier this year, prompting the company to accelerate secondary-market bulk acquisitions to keep its scanning lines running.

Industrial throughput meets cultural preservation

In standard library conservation, rare works are digitized using non-destructive overhead planetary scanners that preserve delicate bindings, a process that is deliberately slow and labor-intensive. For an AI developer aiming to ingest billions of tokens rapidly, non-destructive scanning is an operational bottleneck. Slicing off book bindings turns bound volumes into loose-leaf stacks that document feeders can digitize in seconds.

The revelation has sparked fierce criticism from bibliophiles, historians, and market commentators, who view the physical destruction of rare and out-of-print editions as an irreversible cultural cost. While competitors such as Anthropic and xAI have stated they do not train on rare physical books, Amazon's industrial scanning pipelines demonstrate the extreme measures frontier AI developers are willing to adopt in their quest for high-quality training text.

Mocchi's take

This development makes one reality crystal clear: the true bottleneck of frontier artificial intelligence is no longer raw compute, but high-quality, uncontaminated human data. For Italian and European enterprises, this reinforces why generic web scraping and unverified synthetic datasets yield diminishing returns. Sustainable competitive advantage will not come from brute-force model size, but from how effectively organizations curate, structure, and safeguard their proprietary vertical archives and domain-specific knowledge.

Further reading

All articles on the Mocchi's blog