Хиймэл оюун ухааны компаниуд “Project Panama” төслийн хүрээнд номыг дижитал хэлбэрт оруулахдаа эх хувийг нь устгах аргыг ашиглаж байгаа нь салбарынхны дунд шүүмжлэл дагуулж байна.
Хиймэл оюун ухааны (LLM) загваруудыг сургахад хүний гараар бичигдсэн эх сурвалж нэн чухал болсон тул хуучин номуудын эрэлт хэрэгцээ огцом нэмэгджээ. 2022 оноос өмнөх үеийн номууд нь хиймэл оюун ухааны оролцоогүй, цэвэр хүний бүтээл гэдгээрээ үнэ цэнтэй бүтээгдэхүүн болсон байна. Гэвч технологийн компаниуд эдгээр номыг “устгах шинжтэй сканнердах” (destructive scanning) аргаар дижиталжуулж байгаа нь номын худалдаачдын санааг зовоож байна.
Anthropic компани өмнө нь зохиогчийн эрх зөрчсөн хэргээр 1.5 тэрбум ам.долларын нөхөн төлбөр төлөхөөр тохиролцсон байдаг. Шүүхийн шийдвэрээр номыг сканнердаж сургалтын мэдээлэл болгон ашиглахыг “шударга хэрэглээ” гэж үзсэн ч компаниуд хуулбарлах үйлдлээс зайлсхийхийн тулд номын эх хувийг устгах үүрэг хүлээжээ. Энэ нь номын нурууг тасдах зэрэг аргаар бие махбодийн хувьд номыг үгүй хийх үйл явцыг багтаадаг байна.
Номын худалдаачид нийтийн хэрэглээний номуудыг устгах нь төдийлөн харамсалтай биш ч, түүхэн ач холбогдолтой эсвэл цөөн тоогоор хэвлэгдсэн ховор бүтээлүүд устаж үгүй болох эрсдэлтэйг анхааруулж байна. Ялангуяа хэзээ ч дижиталжуулагдаагүй, техникийн ховор номууд нь хиймэл оюун ухааны компаниудын хувьд “нууц мэдлэг” агуулсан алтны уурхай мэт үнэлэгддэг. Хэрэв эдгээр ховор бүтээл сургалтын өгөгдөл болж ашиглагдаад устгагдвал түүхэн мэдлэгийн сан бүрмөсөн устаж үгүй болох аюултай юм.
Дэлгэрэнгүйг эх сурвалжаас харах
↓Эх сурвалжийг нээх ↓
Another bookseller, David Tobin, who runs north London’s Walden Books—not to be mistaken for Waldenbooks—said he was also in the midst of a business boom. “In some ways it’s very nice to sell some of these titles which haven’t been sold for many years,” he told the BBC. But he added “it would be sad if they are ultimately destroyed.”
Yeah, there’s a pretty good chance they’re being destroyed.
As 404 Media wrote last month, old books are becoming a premium product because anything published after the adoption of LLMs might have LLM outputs in it, which is bad. Up until about 2022, only humans wrote books. And it’s human writing that these companies need in order to train future LLMs.
Anthropic got in big trouble for using illegal copies of books as training data, eventually paying out $1.5 billion to settle a class action lawsuit. Recently unearthed court documents reveal that the company’s new plan for ingesting the data in books, codenamed “Project Panama.” is nice and legal, but a PR nightmare: Destructive scanning, which means taking a physical copy of a book and scanning it, but destroying the physical copy in the process. Scanning books to train AI falls under the fair use doctrine, a court has found. But Anthropic is obligated to destroy books as well, because only one copy of the book exists before the scan, and only one copy exists after, meaning no piracy is occurring.
Perhaps you’ve seen a report recently about this practice featuring this grisly video of books having their spines sliced off. That’s part of destructive scanning.
But in case you think all of this is a big overreaction, or some kind of naive romanticization of physical books, the booksellers in the BBC’s report do note the difference between shredding one of five million copies of the Da Vinci Code, and destroying something of actual value.
As the owner of an Edinburgh bookshop called McNaughtan’s, Derek Walker, points out to the BBC, there are even some relatively rare books the destruction of which wouldn’t be a huge catastrophe. For instance, Walker notes, “A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market — but it is perhaps not such a great loss if one copy is destroyed.” However, Walker says he has been known to sell “the only known surviving example of an edition from the 18th century.”
Rare technical books that have never even been scanned are a sort of holy grail for competing LLMs, since they remain exactly what they are in movies: dusty old artifacts that contain unique, secret knowledge that’s been lost to history. Frontier AI companies seeking to create so-called “superintelligence” will no doubt want this knowledge. And if AI companies really do find, absorb, and destroy such knowledge in pursuit of training data, then I guess by definition we’ll never know it happened.

