Хиймэл оюун ухааны технологи хөгжүүлэгч томоохон компаниуд өөрсдийн LLM загваруудыг сургах чанартай текстэн мэдээлэл олж авахын тулд хэвлэмэл номнуудыг их хэмжээгээр худалдан авч, сканнердсаны дараа устгадаг нь тодорхой болжээ.
Anthropic зэрэг компаниуд олон сая номыг задалж, сканнердаж устгасан нь шүүхийн баримт бичгүүдээр нотлогдсон бөгөөд энэ үйлдэл нь зохиогчийн эрхийн хүрээнд “fair use” буюу шударга ашиглалтад тооцогдох болсон байна. Гэсэн хэдий ч олон нийтийн шүүмжлэлээс зайлсхийхийн тулд эдгээр компаниуд ISBNdb зэрэг зуучлагч үйлчилгээг ашиглан ном худалдан авах ажлыг нууцаар гүйцэтгэх болжээ.
Хиймэл оюун ухааны загваруудыг сургахад интернет дэх чанар муутай контентоос илүүтэй хүний бичсэн, эх сурвалж сайтай текстэн мэдээлэл нэн чухал байгаа нь ийнхүү ховор номнуудын эрэлтийг огцом нэмэгдүүлсэн байна. Ном худалдаачид сүүлийн үед эрэлт ихэссэнийг тэмдэглэж байгаа ч, ховор бүтээлүүдийг ийнхүү устгаж байгаад сэтгэл дундуур байгаагаа илэрхийлжээ.
ISBNdb компани нь “AI-д зориулсан ном нийлүүлэлт” гэсэн үйлчилгээнийхээ хуудсыг устгасан бөгөөд энэ чиглэлээс татгалзаж байгаагаа мэдэгдсэн байна. Хэдийгээр зуучлагч компаниуд энэ үйл ажиллагаагаа зогсоож байгаа ч AI салбарын томоохон тоглогчдын өндөр чанартай дата мэдээлэлд тэмүүлэх эрэлт хэрэгцээ хэвээр үлджээ.
Дэлгэрэнгүйг эх сурвалжаас харах
↓Эх сурвалжийг нээх ↓
AI isn’t doing great on the PR front. Whether it’s the exponentially increasing electricity consumption amidst worsening climate change, the data centers leaking toxic pollutants into neighboring communities, or the years-long parade of business leaders insisting the technology will replace humans in “most, if not all” professional fields, the AI industry has an excellent record of inspiring resentment and contempt.
And, it seems, the AI companies know it. New reporting from 404 Media suggests that, in an effort to shield themselves from the PR fallout of their data harvesting, major AI providers have turned to third-party middleman services to anonymously source the physical books they’ve been stripping for training data at an industrial scale—and destroying in the process.
In an information ecosystem increasingly poisoned by regurgitated and recirculated AI slop, untainted samples of quality written text have, ironically, only become more valuable for the companies making them so hard to find. The AI model arms race requires an ever-widening set of training data, making physical books that predate the proliferation of LLM-generated text a precious source of pristine, human-authored material to imitate—particularly if it’s a book rare enough to be absent from your competitor’s datasets.
Unfortunately, a book is only valuable for an AI company until its last page has been scanned, as demonstrated by a court ruling in 2025 revealing that Anthropic had carved up, de-spined, scanned, and ultimately discarded millions of print books in the process of training its Claude models. Destroying books, it turns out, isn’t just cheaper than maintaining them: The presiding judge also ruled that it’s transformative enough to constitute fair use under Section 107 of the Copyright Act.
But while a judge might have declared it legally defensible, The Washington Post reported that internal documents—unsealed during the author-initiated copyright lawsuit that ended in the company paying a $1.5 billion settlement—indicated Anthropic was well aware that ‘arguably legal’ and ‘cool to do’ are very different things.
“We don’t want it to be known that we are working on this,” Anthropic said in internal planning documents, understanding that mulching millions of books just so its chatbot didn’t sound like “low quality internet speak” might not have been a popular move.
That’s why, as 404 Media now reports, Anthropic and other AI providers are hiring middleman companies to buy those books instead—companies like ISBNdb, which until earlier today featured a now-removed page on its site offering printed book sourcing services, “tailored to your LLM training needs, delivered at the scale AI demands.”
In an archived version of the page, ISBNdb boasted about its ability to secure books “scattered across library shelves, used bookstores, and out-of-print catalogs,” with companies able to purchase up to 1 million books per order. One of the website’s blog posts about the service—also updated today—featured a now-removed promise of a “strict NDA on every engagement,” ensuring that “your identity, strategy, and acquisition targets are never disclosed.” After all, the blog post still reads, “‘AI company destroys two million books’ is not a headline that generates sympathy.”
In a news update, ISBNdb says it removed the page describing the AI book-sourcing service because it was simply “part of exploring demand, and we’ve chosen to pivot away from that direction.”
ISBNdb might be pivoting, but there’s clearly plenty of that demand to go around. 404 quotes a bookseller specializing in rare and low-circulation books who said he and other book dealers have seen an unprecedented surge in sales since April. He’s now fulfilling orders for hundreds of books a week when he’d previously hoped to clear 20.
There’s no commonality between the seemingly random books being bulk-ordered, the bookseller said, except the fact that they all have ISBNs, potentially indicating that the books had been identified by the purchaser from an ISBN-based database. Netherlands book dealers experiencing a similar boom in rare book orders noticed the same trend earlier this year.
“I personally have mixed feelings about all of this. It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I’ve been well-suited for these sales with inventory from overseas and foreign language books,” the bookseller said. “On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.”

2026 games: All the upcoming games
Best PC games: Our all-time favorites
Free PC games: Freebie fest
Best FPS games: Finest gunplay
Best RPGs: Grand adventures
Best co-op games: Better together



