Amazon Trashes Rare Books for AI Training, AirTag Shows

An investigation using hidden AirTag trackers has revealed Amazon is destroying rare and collectible books after scanning them for artificial intelligence training purposes. The discovery raises ethical questions about the company's AI development practices and treatment of cultural artifacts used as training data.
Investigation Method and Discovery
Researchers embedded Apple AirTags inside rare books sold to Amazon to track their post-purchase journey. The tracking devices revealed that after Amazon's team scanned the volumes for digitization, the physical books were sent directly to disposal facilities rather than resold, donated, or archived. According to the investigation published by Ars Technica, this practice appears systematic rather than isolated.
The team responsible for this acquisition and scanning process reportedly uses a logo depicting a Tyrannosaurus rex preparing to devour a book, an image that investigators now view as darkly appropriate given the fate of these literary works. The disposal of rare books contradicts assumptions that digitized volumes would be preserved or returned to circulation.
AI Training Data Collection
Amazon has been aggressively building text datasets to train large language models and other AI systems. Rare books offer unique language patterns, historical vocabularies, and diverse writing styles that enhance model performance across specialized domains. The company's approach involves:
- Purchasing collectible and out-of-print volumes from individual sellers
- High-resolution scanning to create digital training corpus
- Extracting text for machine learning datasets
- Disposing of physical copies after digitization
This methodology allows rapid dataset expansion without maintaining large physical archives. However, the practice has drawn criticism from bibliophiles, archivists, and cultural preservation advocates who argue that irreplaceable editions are being destroyed for commercial AI development.
Release Date
The investigation findings were officially released on August 2026 through Ars Technica's reporting. The tracked books were purchased and processed over preceding months, with disposal occurring between scanning completion and the investigation's publication.
Industry Implications
Amazon's book destruction practice highlights broader tensions in AI training data acquisition. Technology companies face mounting pressure to build comprehensive datasets while navigating copyright concerns, ethical sourcing questions, and cultural preservation responsibilities. The revelation that rare books end up in landfills after scanning contradicts public perception that digitization efforts preserve literary heritage.
Competitors developing similar AI capabilities may employ different acquisition strategies, including partnerships with libraries and archives that maintain physical collections post-scanning. The controversy underscores growing scrutiny of how tech giants source training data and what happens to original materials afterward.
What This Means
This investigation exposes a previously hidden aspect of AI training infrastructure where cultural artifacts become disposable raw materials. For book collectors and sellers, the findings raise questions about Amazon's purchasing intentions and whether rare volumes are valued beyond their digital utility. As AI development accelerates, companies will face increased demands for transparency about training data sources and ethical acquisition practices. The practice of destroying rare books after scanning may prompt regulatory discussions about cultural preservation requirements in AI dataset creation, particularly when irreplaceable editions are involved.
on Emergent today






