Технологийн компаниуд том хэлний загвар (LLM) сургах зорилгоор хуучин номуудыг бөөнөөр нь худалдан авч, агуулгыг нь дижитал хэлбэрт оруулан устгах практик түгээмэл болжээ.
Хиймэл оюун ухаан хөгжүүлэгчид интернет дэх чанар муутай контентоос зайлсхийхийн тулд 2022 оноос өмнөх үеийн хэвлэмэл номуудыг ашиглахыг илүүд үзэж байна. Anthropic зэрэг компаниуд худалдан авсан номуудаа гидравлик төхөөрөмжөөр задалж, үйлдвэрийн зориулалттай дүрс бичлэгийн төхөөрөмжөөр сканнердах замаар өгөгдлөө бэлтгэдэг аж. Шүүхийн шийдвэрээр энэ үйл явцыг “хувиргах шинжтэй” буюу шударга ашиглалтын (fair use) хүрээнд хамаарна гэж үзсэн нь тус салбарт томоохон маргаан дагуулаад байна.
Энэхүү үйл ажиллагааг ISBNdb зэрэг мэдээллийн сангууд зуучилж байгаа бөгөөд нэг захиалгаар нэг сая хүртэлх номыг бөөнөөр нь худалдан авдаг байна. Компаниуд олон нийтийн зүгээс ирэх шүүмжлэлээс зайлсхийхийн тулд худалдан авагчийн нууцыг чандлан хадгалахыг чухалчилдаг.
Ном борлуулагчдын хувьд хуучин номын нөөцөө борлуулах санхүүгийн боломж гарч байгаа ч ховор болон дахин хэвлэгдэхгүй номууд устгагдаж байгаад санаа зовниж буйгаа илэрхийлжээ. Нидерланд зэрэг улсын ховор номын худалдаачид ч мөн адил технологийн компаниудын ийм төрлийн бөөнөөр худалдан авалт ихэссэнийг тэмдэглэв.
Дэлгэрэнгүйг эх сурвалжаас харах
↓Эх сурвалжийг нээх ↓
Traditionally, books have been good for two things: reading, and looking nice on a shelf.
But AI companies are interested in neither.
To those building large language models, books are nothing more than fodder to be devoured en masse before being spit out like fishbone. Often, they’re happy to use digital books — or even better, pirated digital books, as Meta has been accused of doing, and as Anthropic was forced to pay a $1.5 billion settlement to authors for also doing.
But many companies, including Anthropic, have turned to ingesting physical books instead, which they can buy countless used copies of on the cheap. According to the settled lawsuit, Anthropic used a hydraulic powered cutting machine to neatly remove the pages from the books it procured from book resellers and then scanned them using industrial-grade imaging equipment. In other words, it was literally ripping off authors’ books to train its AI.
This process took advantage of a legal concept known as first-sale doctrine, which allows a buyer to do what they want with a purchase without the original copyright holder’s say-so. And since Anthropic was turning the original physical texts into digital ones — rather than redistributing them as new copies — a judge found this to be “transformative,” and therefore protected by fair use.
Now, as 404 Media reports, this practice has become prevalent enough that even well-established book sellers are looking to cash in on the AI boom. One called ISBNdb, which boasts the “world’s largest book database,” extolls that the “world’s best AI training data is setting on a shelf,” upholding these physical texts as uncorrupted by shoddy AI writing that’s already polluted so much of the internet (and indeed, newer books).
“Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage,” it explains in an article on its website, as quoted by 404.“Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools.”
Once focused on helping libraries, distributors, and book shops find and sell books, ISBNdb now helps AI companies bulk-buy anywhere between 1,000 to one million books per order, according to 404.
As an added bonus, it also promises AI companies that it’ll keep their purchases under wraps — nobody wants to end up in the spotlight like Anthropic and Meta, obviously — while clearly sounding aware about how incredibly shady the practice sounds.
“The optics problem is real,” ISBNdb’s site says. “‘AI company destroys two million books’ is not a headline that generates sympathy.”
One small book seller said that in April, he suddenly went from selling no more than 20 books a week to hundreds, and he’s almost certain that the customers are AI labs, noting the random selection of the books and how they all have ISBNs. He added that his inventory is full with rare and out of print books, meaning that an AI company could be destroying some of the few remaining copies that can be found.
“I personally have mixed feelings about all of this,” the bookseller told 404. “It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I’ve been well-suited for these sales with inventory from overseas and foreign language books. On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.”
More antique books may be endangered. 404 noted how rare booksellers in the Netherlands are being inundated with bulk purchases they believe are being made by AI companies, though they can’t know for sure.
Similar suspicions are being felt all throughout the industry. Services like ISBNdb facilitate these bulk orders, keeping the AI buyers anonymous as promised. There may be strong hints thats a bulk order is for an AI lab, but booksellers can, for the most part, merely speculate.
More on AI: Author Invited to Give Speech at OpenAI Headquarters, Uses Opportunity to Trash AI to Their Faces
The post AI Companies Are Buying Antique Books, Ingesting Their Contents to Train Models, and Then Destroying Them at Incredible Scale appeared first on Futurism.

