Хиймэл оюун ухаан иргэдийн шинжлэх ухааны мэдээллийн сангийн чанарт заналхийлж байна

Published:

Энэхүү мэдээ, нийтлэлийг хиймэл оюун боловсруулав.

Хиймэл оюун ухаанаар бүтээсэн хуурамч контентууд нь биологийн судалгааны мэдээллийн сангуудын найдвартай байдлыг алдагдуулж байгааг эрдэмтэд анхаарууллаа.

“Nature” сэтгүүлд нийтлэгдсэн Корнеллийн их сургууль болон Манчестерийн Метрополитан их сургуулийн судлаачдын захидалд иргэдийн оролцоотой шинжлэх ухааны (citizen science) платформууд хиймэл оюун ухааны бүтээсэн зураг, дүрслэлд автаж байгааг онцолжээ. Шувуу ажиглагчдын ашигладаг iNaturalist зэрэг платформууд нь тухайн зүйлийн тархалт, зан төлөвийг судлахад чухал ач холбогдолтой ч, хиймэл оюун ухааны хуурамч өгөгдөл нь эдгээр мэдээллийн сангийн үнэ цэнийг бууруулж байна.

Шувуу таних Merlin зэрэг аппликейшнүүд нь машин сургалтын технологийг ашиглан амжилттай ажилладаг хэдий ч тэдгээрийн сургалтын өгөгдөл нь хуурамч мэдээллээр бохирдох эрсдэлтэй юм. Компьютерийн шинжлэх ухаанд “хог оруулбал хог гарна” (garbage in, garbage out) гэж нэрлэдэг энэ үзэгдэл нь хиймэл оюун ухааны загваруудын ирээдүйн сургалтын чанарыг улам бүр дордуулах сөрөг үр дагавартай.

Энэхүү “загварын нуралт” буюу хиймэл оюун ухаан өөрийн бүтээсэн чанар муутай өгөгдлөөр дахин суралцах үйл явц нь зөвхөн шувуу ажиглалт төдийгүй технологийн салбарт бүхэлд нь тулгарч буй томоохон сорилт болоод байна. Интернет орчинд хиймэл оюун ухааны бүтээсэн контент ихсэх тусам ирээдүйн загваруудын үнэн зөв байдал улам бүр доройтох эрсдэлтэйг эрдэмтэд анхааруулж байна.

Дэлгэрэнгүйг эх сурвалжаас харах

Эх сурвалжийг нээх ↓

On the off chance that you’re one of those rare people who starts a new week with a flush of optimism and goodwill toward all mankind, instead of a craving for caffeine and the last vestiges of a hangover, let me put things right.

You may, for instance, find your early-week optimism extending to AI. On reflection, maybe things aren’t that bad. Perhaps there are things that AI can’t ruin after all. Surely something like… oh, I don’t know, the innocent, serene pastime of birdwatching is immune to the depredations of AI slop?

Nope. NOPE. As proof, behold a letter entitled, “Citizen science platforms must mitigate against the threat of generative AI,” published recently by the venerable journal Nature. The letter isfrom a group of scientists based variously at Cornell and Manchester Metropolitan University, and it discusses how a flood of AI-generated or enhanced images is undermining the utility of citizen science.

The scientists point out that while “[the utility of] biological recording citizen-science platforms … may be compromised if databases become appreciably contaminated by media produced by generative text-to-image machine-learning models.” The platforms to which the scientists refer are essentially databases of observations made by amateurs—and birdwatching is a prime example of a hobby that generates such observations.

Birdwatchers use platforms like iNaturalist to record when and where they’ve seen various species. Such information can be valuable to scientists; indeed, as the letter explains, databases of crowd-sourced amateur data “increasingly underpin our knowledge about where species occur in space and time and how they behave.”

However, these platforms are of no use if they’re full of AI-generated nonsense—and, as the scientists relate, this is increasingly the case. They explain that both AI-enhanced images and images completely generated from scratch by AI are appearing more and more frequently on large citizen science databases. They cite multiple examples that were discovered by moderators, but also concede that “it is unclear at present how pervasive this problem may be, as some or even many such images may go undetected.”

The dreaded AI ouroboros also rears its ugly head here. As well as platforms to record their observations, birdwatchers use apps like Merlin to identify the birds they’re observing. These apps use machine learning to assist with the identification and classification of birds, relying on things like recordings of birds’ songs, records of geographical distribution, etc. They work impressively well, and provide a fine example of how useful machine learning can be in the right context.

However, the output of the algorithms that power these apps is contingent on the quality of their training data; if the data from which they’re learning is unreliable, so too are the results they provide; or, in the time-honored parlance of computer science: garbage in, garbage out. In this conetxt, the problem is worse because it’s self-perpetuating—it makes it more likely for even well-intentioned bird enthusiasts to make false observations, further degrading the pool of training data, and… well, you can see where this is going.

This pattern—of generative AI degrading the quality of data on which future AI models will be trained—is a problem that extends well beyond birdwatching. There are various terms for this phenomenon: “model collapse,” “AI cannibalism,” and our favorite, “Habsburg AI.” The implication is clear: the more that AI slop pervades the internet, the more its ubiquity will manifest in the degraded quality of future models.

There’s a certain comedic value in the idea that the singularity might turn out to be less futuristic hyperintelligence and more Charles II of Spain, but it would be nice if the road to the virtual Habsburg jaw didn’t run straight over the top of the few nice things left in the world. Like birdwatching. And music. And having a decent GPU. And…

- Зар сурталчилгаа -

Та юу гэж бодож байна?

Сэтгэгдлээ оруулна уу!
Please enter your name here

MFC.mn сайтад сэтгэгдэл оруулахад анхаарах зүйлс

Холбоотой

spot_img

Шинэ

spot_img