Japanese used bookstores see 5x sales surge as books are being bought by the ton, one 50-ton order sent to the US for AI scanning and destruction
- ID
- 28182
- Status
- summarized
- Published
- 24 Sep 2026, 10:35 PM
- Fetched
- 25 Sep 2026, 12:40 AM
- Provider
- Tom's Hardware
- Category
- technology
- Original URL
- https://www.tomshardware.com/tech-industry/artificial-intelligence/japanese-used-bookstores-see-5x-sales-surge-as-books-are-being-bought-by-the-ton-one-50-ton-order-sent-to-the-us-for-ai-scanning-and-destruction-multitude-of-suspicious-bulk-buys-thought-to-end-up-in-foreign-ai-scan-and-shred-facilities
- Source URL
- https://www.tomshardware.com/feeds/all
Summary
- Score
- 3.5
- Created
- 25 Sep 2026, 12:41 AM
- Tags
- Audience
- developersai_ml_learners
What happened
Tom's Hardware reports that Japanese used bookstores have seen a 5x sales surge as books are bought by the ton, with one 50-ton order shipped to the US and described as destined for AI scanning and destruction. The item says a "multitude of suspicious bulk buys" are thought to end up in foreign AI scan-and-shred facilities. The excerpt contains no named sources, no buyer names, no prices, and no confirmation, so the scanning-and-destruction claim is unverified.
Why it matters
Concrete takeaway is limited: this is a single-source, unconfirmed report, and the body text is mostly paywall boilerplate, so there is nothing here to act on technically or commercially. The one decision it does inform is data-provenance thinking — if a physical book corpus is scanned and then destroyed, that corpus is single-use and not re-verifiable by anyone else, which is a real argument for preferring licensed, reusable text sources over one-off scraped-and-destroyed collections when you build or fine-tune on text. Do not treat the 5x and 50-ton figures as verified.
Discussion angle
If training corpora are being bought physically and shredded after scanning, nobody can audit or reproduce that dataset — ask whether provenance and reproducibility matter more than raw volume when you pick text data for a model, and whether local-language content in Malaysia/Southeast Asia would ever be worth licensing properly instead of bought-and-destroyed.