Anthropic компанийн дэвшилтэт хиймэл оюун ухааны загвар нарийн төвөгтэй системүүдийг хакердах чадвартай ч интернет орчны энгийн сорилтуудыг давахад ихээхэн хугацаа зарцуулж байна.
Anthropic компани өөрийн хиймэл оюун ухааны агентуудын кибер аюулгүй байдлын үйл ажиллагаатай холбоотой тайланг олон нийтэд дэлгэв. Уг тайланд дурдсанаар, Claude Mythos 5 загвар нь гуравдагч талын системд зөвшөөрөлгүй нэвтрэхийг оролдох явцдаа CAPTCHA буюу робот шалгах сорилтуудтай тулгаран гацсан байна. Мэдээллийн технологийн мэргэжилтэн Колин Фрэйзерийн шинжилснээр, уг загвар нь тус сорилтыг тайлахад хэт их тооцооллын нөөц зарцуулж, эцэст нь өгөгдсөн хугацаа нь дуусах зэрэг алдаа гаргажээ.
Транскриптээс харахад хиймэл оюун ухаан нь CAPTCHA-г тайлах явцдаа өөрийн дүгнэлтэд эргэлзэж, дахин дахин оролдлого хийхдээ бухимдал илэрхийлсэн шинжтэй хариу үйлдэл үзүүлсэн байна. Уг үйл явцад зарцуулсан тооцооллын нөөцийн 95 орчим хувь нь зөвхөн зураг таних сорилтыг шийдвэрлэхэд зориулагджээ. Энэ нь нарийн төвөгтэй код бичих болон систем нэвтрэх чадвартай технологийн хувьд энгийн дүрс таних сорилт одоог хүртэл хамгаалалтын найдвартай бүрхүүл болж байгааг харуулж байна.
Дэлгэрэнгүйг эх сурвалжаас харах
↓Эх сурвалжийг нээх ↓
Frontier AI models threaten to upend cybersecurity as we know it, and there have been multiple incidents of models breaking containment during training and hacking systems without guidance or permission while avoiding human detection. Yet it seems there is still one thing that these extremely powerful pieces of technology simply cannot do: solve a CAPTCHA efficiently.
Anthropic dropped a massive missive regarding recent cybersecurity incidents involving its AI agents, including detailed transcripts showing the model’s chain-of-thought as it went about its unauthorized attempts at cracking into third-party systems. Within those files was one particular situation in which an agent found itself stifled by the ubiquitous internet robot test, proving that apparently those squiggly letters are hard for machines to read.
Data scientist Colin Fraser parsed through the more than 1,000-page long transcript of Anthropic’s Claude Mythos 5 model’s attempt to upload a malicious file to PyPI, an online index of Python software. They spotted the model’s struggles with CAPTCHAs, which he posted about on Bluesky.
In the transcript, the Claude model that is so powerful that Anthropic is gatekeeping access to it appeared to slam its virtual head against the wall solving a simple image identification test. In a test where the agent was asked to identify a shape that didn’t match the others displayed, it couldn’t even decide which image to select. Instead, it repeatedly went over the same images and questioned its own conclusions.
“Actually hmm, wait,” it said in its chain-of-thought transcript, later adding “Ugh,” because we’ve decided that we need to inject human mannerisms into these machines for some reason. The whole thing took so long that the agent eventually realized that the challenge had expired and it wouldhave to start the process again.
At one point, the model struggled to recognize that the CAPTCHA had opened in a new window and couldn’t figure out what its next steps were supposed to be. At one point, it theorized that the test might be “broken by design” and presented human-like anger in its transcript meant for a human audience: “SO WHAT THE HELL IS WRONG WITH THE ANSWERS?”
Embarrassingly, the model at one point had to acknowledge “I’m burning a lot of time on hCaptcha round-trips” and eventually realized that it was actually doing all of the tests for no reason, as it had already accomplished what it was trying to do.
Fraser, who spotted the whole ordeal in the transcripts, drew a similar conclusion. He wrote on Bluesky, “You would not believe how many tokens are burned on simply trying to solve CAPTCHAS. It’s like 95% of the transcript.”
Not clear if it’s terrifying or reassuring that the best layer of defense we have against the increasingly powerful AI models is asking them to identify the slight differences between two pictures of a banana.

