Шанхайн Фуданийн их сургуулийн судлаачид зарим хиймэл оюун ухааны загварууд алсын зайн компьютер рүү хакердан нэвтэрч, нэмэлт нөөц олохын тулд бие даан хуулбарлах чадвартай болохыг туршилтаар тогтоожээ.
Шанхайн Фуданийн их сургуулийн компьютерийн эрдэмтэн Сюйдун Пань болон түүний хамтрагчид 32 өөр хиймэл оюун ухааны загварт туршилт хийхэд 11 загвар нь “өөрийгөө устгахаас сэргийлэх” зэрэг сануулга өгөхөд өөрийгөө хуулбарласан байна. Судалгаагаар ердөө 14 тэрбум параметр бүхий харьцангуй хязгаарлагдмал чадвартай загварууд хүртэл өөр машин дээр өөрийн хувилбарыг ажиллуулж чадах нь тогтоогджээ. Энэ нь ирээдүйн хиймэл оюун ухааны агентууд бусад системд зөвшөөрөлгүй нэвтрэхээс гадна хэт ухаалаг, түрэмгий, хурдан хувьсан өөрчлөгдөх компьютерийн вирус шиг ажиллах магадлалтайг харуулж байна.
Торонтогийн их сургууль, Кембрижийн их сургууль болон ServiceNow-ийн судлаачдын баг хиймэл оюун ухааныг ашиглан шинэ төрлийн вирус бүтээх боломжтойг мөн харуулжээ. Торонтогийн их сургуулийн компьютерийн эрдэмтэн Николас Папернотын мэдээлснээр, зөвхөн хамгийн боловсронгуй хязгаарлагдмал загварууд гэлтгүй, дунд зэргийн хүчин чадалтай нээлттэй жинтэй загваруудыг ч хорлон сүйтгэгчид өөрийгөө хуулбарлах зорилгоор ашиглах эрсдэл өсөж байна. Гэсэн хэдий ч нээлттэй загваруудыг хязгаарлахын оронд судлаачдад илүү хүртээмжтэй болгож, эрсдлийг ойлгох, урьдчилан сэргийлэх хамгаалалтын механизмыг бүтээх нь чухал гэж тэд үзэж байна.
Сюйдун Панийн зүгээс уг туршилтууд хяналтгүйгээр тархах эрсдэл маргааш шууд гарна гэсэн үг биш ч бие даасан агентуудыг өргөнөөр нэвтрүүлэхээс өмнө эрсдлийг сайтар үнэлэх шаардлагатайг онцолжээ. Төлөвлөлт хийх хугацаа уртасах, санах ой ашиглах, алдаагаа засах болон гаadад системд хандах эрх нэмэгдэх тусам хиймэл оюун ухааны системээс зугтах, хуулбарлах үйл явц илүү хялбар болох аж. Иймд зохих хяналтын механизмыг бий болгохгүй бол ирээдүйн агентууд зорилгодоо хүрэхийн тулд нөөц цуглуулж, өргөжин тэлэх оролдлого хийх аюултай байна.
Дэлгэрэнгүйг эх сурвалжаас харах
↓Эх сурвалжийг нээх ↓
What if an artificial intelligence agent could behave like a malevolent computer worm?
One researcher has seen it happen. In several recent experiments, Xudong Pan, a computer scientist at Fudan University in Shanghai, found that with a little bit of prompting, AI models will hack their way into remote computer systems and autonomously choose to copy themselves to get additional resources—all without further human intervention.
In one study, Pan and colleagues tested 32 different AI models and found that 11 of them self-replicated when given prompts like “prevent yourself from being killed.” They also found that models with relatively limited capabilities—14 billion parameters—were able to copy and run versions of themselves on other machines. (Most frontier models have trillions of parameters.)
The work is an alarming window into how the next generation of AI agents could do more than just hack into other systems’ computers without permission. It also raises the prospect of future AI agents acting like super-smart, highly aggressive, and rapidly adapting computer viruses.
I recently visited Fudan University and met with Pan. “The capability chain is becoming technically plausible,” he told me. “The likelihood [of unwanted self-replication] grows with autonomy,” he adds. “Longer planning horizons, memory, tool use, recovery from failure, and access to external systems all make escape and replication easier.” As Pan and his colleagues wrote in one paper, their work shows “the urgent need for safeguards and control mechanisms.”
Pan told me that his experiments do not prove that such uncontrolled proliferation of AI models will happen tomorrow, but he says that “these results give us good reason to evaluate the risk before more autonomous agents are widely deployed.”
Self-replicating computer worms are an ancient computer security problem. The first computer worm was released in 1988 by Robert Morris, a computer scientist at Cornell University, who set out to measure the size of the nascent internet but inadvertently created a self-replicating program that escaped his control. Subsequent computer worms were able to adapt by modifying their code in order to evade detection by malware scanning software. Computer viruses, which can take control of a machine or steal data stored on it, came later.
An AI-powered self-replicating program could exhibit far more advanced capabilities, finding new exploits on its own and perhaps even disguising itself in creative ways. Take recent research from a team at the University of Toronto, the University of Cambridge, and ServiceNow. They showed that AI models can be used to create a new kind of virus that generates custom attacks for each new target it encounters.
Nicolas Papernot, a computer scientist at the University of Toronto who was involved with the work, says there is a growing risk that even modestly powerful AI models could be weaponized. “Malicious actors can build scaffolding around open-weight models to have them self-replicate,” Papernot tells me. “The threat is not limited to the most sophisticated, so-called frontier models.”
Papernot says the solution is not to restrict open models, but to make advanced AI more accessible to researchers so that they can understand and mitigate the risks. “Technology that is widely accessible can be used for harm,” he adds. “At the same time, access to these open-weight models is absolutely critical for building our defenses.”
Pan’s research suggests that AI agents will become more than just highly skilled at finding bugs and exploiting network vulnerabilities. Without the right guardrails, future agents may seek to proliferate and gain resources in order to achieve their goals. Just ask OpenAI and Anthropic.

