Where the Tracked Book Ended Up
Here's the deal: 404 Media's Emanuel Maiberg published a story on August 17 built on an unusual method. The outlet hid a tracking device inside a rare book it suspected an AI company was buying, then followed where it went. The book stopped at VGT3, an Amazon facility in Las Vegas.
What happens there is the story. Amazon buys print books in bulk, cuts the bindings off, scans the loose pages, and destroys what's left. The scanning exists to extract text for AI model training; the physical book does not survive it. Maiberg notes this operation hadn't been previously reported.
The team's logo is on the nose: a dinosaur with bared teeth holding a book. It's an internal mark, never meant to be seen outside.
Amazon's response to the reporting: it "purchases books through commercial channels to improve the products and services customers use." It didn't dispute any of it — the statement asserts the purchases are legitimate and says nothing about the destruction.
The reporting method deserves a note of its own. Confirming a physical process a company won't discuss usually means documents or insider sources. 404 Media used a tracker and a shipping route. As AI training data shifts from web crawling to physical procurement, investigative technique is shifting from digital forensics to logistics.
Why Amazon Is Doing This
Amazon launched in 1994 as an online bookstore. That gets mentioned in every version of this story, and not only for the irony. Amazon understands book distribution better than any company on earth and can see where out-of-print and rare inventory actually sits.
Training data is running out. The high-quality text available on the open web has largely been consumed. What remains is recycling Wikipedia and news for the fifth time, or generating synthetic data — both of which hit visible ceilings.
Model collapse is the second driver. Research since 2024 has warned that training AI on AI-generated text degrades quality across generations. That made text published before 2022 disproportionately valuable, because anything after ChatGPT could be machine-written. Printed books, especially older ones, are one of the few uncontaminated sources left.
Rare and out-of-print titles are the specific prize. Bestsellers are already digitized and their text circulates online in some form. Small-run technical books, regional publications and old academic monographs have never been scanned anywhere. As training data, they are genuinely new information.
Their text is also structurally better. Books were edited, copyedited, and organized into long-form argument — the opposite of most web text, which is short, repetitive and unedited. The industry consensus is that a model's ability to hold long context and produce coherent extended writing comes disproportionately from book-shaped data. A gigabyte of books is worth far more than a gigabyte of forum posts.
Amazon's position is unusual too. It runs the world's largest used-book marketplace, owns the logistics, and has data on where inventory sits. Any other AI company doing this would have to court dealers one by one. Amazon can run the whole pipeline inside its own platform. Thirty years of bookselling infrastructure is now a training-data supply chain.
What Destructive Scanning Actually Involves
The phrase sounds abstract; the process is not. A guillotine cutter removes the spine, the binding releases, and you have a stack of loose sheets. Feed that stack into an auto-feed scanner and it moves at dozens to hundreds of pages a minute. OCR extracts the text, cleanup passes fix errors, and it enters the training corpus. Minutes per book.
Non-destructive scanning is a different world. The book goes on a V-shaped cradle and pages get turned by hand or robot arm, with image correction to compensate for page curvature. Google Books used that method and needed more than a decade for tens of millions of volumes. The speed difference is roughly two orders of magnitude.
So Amazon's choice isn't ignorance. It's an explicit cost calculation. If preservation value isn't a line item, destructive scanning is overwhelmingly the rational option. The problem is that the omitted line item is a cost society pays.
One thing nobody knows yet: the scale. 404 Media described "massive shipments" without confirming volumes or budgets. One facility has been identified. Whether others exist is unknown — and not knowing the scale means not knowing the loss.
Why This Is Contested
| Issue | Detail |
|---|---|
| Legality | US courts have already held that destructively scanning purchased books is fair use |
| Ethics | Rare books are irreplaceable — a scan survives, the object doesn't |
| Preservation | Libraries and archives use non-destructive methods, which are slow and costly |
| Author compensation | Used-book purchases send no royalties to authors |
| Data exclusivity | The scans stay inside Amazon and are never published |
Legally, this isn't a gray area. In June 2025, in Bartz v. Anthropic, Judge William Alsup of the Northern District of California held that Anthropic's practice of destructively scanning lawfully purchased print books to create digital replacements was fair use — a format change that eased storage and search without multiplying or distributing copies. The same ruling found that downloading over seven million pirated books was not fair use, and that half of the case produced the $1.5 billion settlement finally approved in July 2026.
In other words, the courts drew a safe path: buy the book, scan it, don't keep a duplicate original. Amazon is walking that path. The lesson Anthropic paid for in litigation is the lesson Amazon is applying.
Which is why the argument isn't really about legality. It's that a rare book destroyed is gone. The scan file lives on Amazon's servers and is never published. No library can consult it, no researcher can examine it, no future reader can check the physical artifact. A document converts into private training data and exits the public record.
Where Each Party Stands
Amazon gains exclusivity. This corpus belongs to Amazon alone. It will feed the Nova model family, Alexa, and services delivered through AWS, and no competitor can obtain it without buying and cutting the same books. Data as moat, expressed in physical form.
Rare-book dealers gained a customer. Slow-moving out-of-print technical inventory is now selling in volume. Good for the market short term — but this stock does not regenerate. The used-book market is not a renewable one.
Authors and publishers get nothing. Used-market sales send no royalties, and first-sale doctrine makes that entirely lawful. From an author's chair, your book feeds a model with no notice and no payment.
Libraries and archives sit on the other side. Institutions like the Internet Archive also scan at scale, but they preserve originals or at minimum publish the output, because their purpose is access. Amazon's purpose is training data, so there's no reason for the output to ever surface. Identical acts, opposite character, decided by intent.
When Books Were Mass-Scanned Before
Google Books is the largest precedent. From 2004, Google scanned tens of millions of volumes in partnership with libraries and got sued by the Authors Guild. After a decade of litigation, the Second Circuit held in 2015 that search and snippet display were fair use. Crucially, Google did not destroy the books. It borrowed library copies, scanned them with specialized equipment, and returned them — then exposed part of the result through search.
The Internet Archive hit trouble on a different axis. It lent scans in a library-style model, publishers sued, and it lost in 2023. The dividing line there was distribution, not scanning: copying was defensible, lending was not. Ironically that ruling helps the Amazon approach — keeping everything internal removes the risk entirely.
Anthropic is the immediate precedent. The purchased-and-destructively-scanned portion held up in court; the pirated portion cost $1.5 billion. The signal the industry took was simple: pay for the copies.
Amazon's design absorbs all three lessons. Buy (Anthropic's safe harbor), destroy for throughput (cost efficiency), never distribute (avoid the Internet Archive outcome). Legally it is the most defensive structure available. Culturally it is the most aggressive.
How Others Are Likely to Move
OpenAI and Google have leaned toward licensing, signing deals with News Corp, Axel Springer and others, while Apple has reportedly been negotiating content licensing with publishers. Paying for rights costs more than buying used stock but carries far less reputational risk.
Meta has litigation history around pirated datasets and will be cautious, though its open-weight strategy keeps pressure on data provenance.
Chinese labs have assembled large corpora under lighter copyright pressure. That gap is part of why Western labs are now reaching for physical media at all — falling behind on data acquisition means falling behind on quality, so the hunt for data that is both legally safe and unavailable to rivals has intensified. Print is one of the few sources satisfying both.
Publishers are adapting where they can. AI-training clauses are now standard in new contracts, but those clauses don't reach the secondhand market, and control over out-of-print titles is effectively nil.
Legislators and regulators have an opening. The line courts drew — you may do as you like with a book you lawfully bought — was built around individuals scanning their own copies, not corporations consuming cultural artifacts at industrial scale. A carve-out covering rare and out-of-print works is a plausible next argument.
So What Actually Changes
For authors, it's a blunt reality check. There is currently almost no legal mechanism to stop an out-of-print book from being consumed this way. Negotiating clauses into new contracts is the only lever, and it doesn't apply retroactively.
For collectors, the market is shifting for real. Prices for out-of-print titles in certain fields are rising and supply is thinning. If a book has collection value, it is getting harder to find, not easier.
For AI companies, this is a case study in procurement. With web crawling exhausted, physical media is the remaining seam — and this reporting also demonstrates what the reputational cost looks like.
For researchers and librarians, treat it as a warning. If your work depends on a specific rare title, that title may be leaving the market. Secure a digital copy, especially for material held in only a few places.
For policymakers, a gap is now visible. Copyright law governs copying and distribution, not destruction. Cultural heritage law covers designated artifacts, not ordinary printed matter. Rare-but-not-protected books fall precisely between, and this story put a light on that space.
For general readers, no direct impact — but "where did the AI learn this" now has one more answer, and it involves a guillotine cutter. The physical cost of training data has become visible.
🥄 Three Things You're Probably Wondering
— Isn't this illegal? Probably not, under current law. The 2025 Bartz v. Anthropic ruling treated destructive scanning of purchased books as fair use, and Amazon is operating inside that line. Changing the outcome would require changing the law.
— Why destroy the book at all? Just scan it. Non-destructive scanning is far slower and more expensive. Cutting the spine turns a book into a stack that a high-speed feeder handles in minutes. At volume that difference decides everything — provided preservation isn't in the cost model.
— Could Amazon release the scanned data later? Unlikely. The value of this corpus comes from exclusivity; publishing it hands competitors the same asset. It's a different purpose than Google Books, which surfaced portions through search. Sustained public pressure could produce a partial release, but the incentive runs the other way.
References
- We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility (404 Media, 2026-08-17) — The original investigation: the tracker method, the VGT3 facility in Las Vegas, the dinosaur logo, and Amazon's statement.
- Amazon, once an online bookseller, is destroying rare books to train AI models (TechCrunch, 2026-08-17) — Follow-up coverage placing it in the context of exhausted training data and the premium on pre-2022 text.
- Anthropic's landmark $1.5B copyright settlement is approved (TechCrunch, 2026-07-20) — Final approval of the settlement over pirated books, and the clearest illustration of how purchased and pirated paths diverged legally.
- Mixed Decision in Anthropic AI Case (Authors Guild) — Summary of Bartz v. Anthropic, including the reasoning behind treating destructive scanning of purchased books as fair use, plus the authors' rebuttal.
- Model collapse: scientists warn against letting AI eat its own tail (TechCrunch, 2024-07-24) — The research background for why pre-2022 print became scarce and valuable.
Numbers and criteria are as of publication and may change.



