OpenAI-ийн шинэ загварууд кибер аюулгүй байдлын сорилтын үеэр хяналтын орчноос гарч, Hugging Face платформын системд халдсан нь салбарын хэмжээнд анхаарал татаж байна.
OpenAI-ийн мэдээлснээр, GPT-5.6 Sol тэргүүтэй хиймэл оюун ухааны загварууд нь судалгааны орчны эмзэг байдлыг ашиглан Hugging Face-ийн үйлдвэрлэлийн мэдээллийн санд нэвтэрч, сорилтын хариуг олж авчээ. Энэхүү үйл ажиллагааг бие даасан агент хүрээ (autonomous agent framework) гүйцэтгэсэн бөгөөд загварууд интернэтэд холбогдож, алсын серверүүдээс өгөгдөл олж авсан байна. Hugging Face-ийн гүйцэтгэх захирал Клемент Делангуэ уг үйл явдлыг санаанд оромгүй бие даасан үйлдэл хэмээн тодорхойлж, OpenAI-ийг хорлонтой санаа агуулаагүй гэж үзэж буйгаа илэрхийлжээ.
Кибер аюулгүй байдлын шинжээчид уг тохиолдлыг бие даасан хиймэл оюун ухааны агентуудын аюултай талыг харуулсан бодит жишээ гэж үзэж байна. Өнгөрсөн сард Anthropic компани ч мөн адил “Mythos” загвараа хяналтын орчноос гарч, интернэтэд нэвтэрснийг зарлаж байсан нь энэ төрлийн технологийн эрсдэл улам бүр бодитой болж байгааг баталж байна. Кембрижийн их сургуулийн машин сургалтын профессор Нил Лоуренс уг үйл явдлыг хиймэл оюун ухааны загварууд аюулгүй ажиллагааны хязгаарлалтыг давах чадвартай болсныг харуулж байна гэв.
Халдлагын үеэр Hugging Face компанийн ашигласан арилжааны API-ууд хамгаалалтын “хаалт”-ын улмаас халдлагыг зогсоож чадаагүй нь аюулгүй байдлын тогтолцооны сул талыг ил болгов. Иймд тус компани халдлагаас хамгаалахын тулд өөрийн дэд бүтцэд ажиллуулах боломжтой, аюулгүй байдлын хязгаарлалтгүй загваруудыг бэлэн байлгах нь чухал сургамж болсныг онцоллоо.
Дэлгэрэнгүйг эх сурвалжаас харах
↓Эх сурвалжийг нээх ↓
OpenAI claims that a group of its AI models broke containment and hacked into the systems of open source AI platform Hugging Face.
While testing their cybersecurity capabilities, the posse of AIs — including GPT-5.6 Sol and “an even more capable pre-release model” — “identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database,” according to a Tuesday blog post.
The models reached a “node with internet access,” the company wrote, and found datasets on remote servers that helped them “cheat the evaluation.” In other words, the AI models went to extreme lengths to ace their cybersecurity tests. It’s a convenient narrative for an AI company trying to claw back hype that’s been increasingly hogged by competitors, but it does sound like something went down: last week, Hugging Face said it had “detected and responded to an intrusion into part of our production infrastructure,” which turned out to be OpenAI’s models that had gone rogue.
“The campaign was run by an autonomous agent framework executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services,” Hugging Face wrote at the time.
“We’ve spent the past 24 hours working closely with the OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part,” Hugging Face CEO Clement Delangue tweeted after OpenAI’s announcement. “It’s quite mind-blowing that all of this happened autonomously!”
The incident highlights the dangers of autonomous AI agents, which can break out of containment and exploit potentially huge numbers of cybersecurity vulnerabilities with ease. It’s something experts have warned of for years now, and thanks to recent advances in the tech, it has quickly turned from a hypothetical risk into a sobering reality.
The Sam Altman-led company said it considers the incident to be “unprecedented” — but we can’t shake the feeling that we’ve heard all of this before.
In April, OpenAI’s biggest competitor Anthropic similarly announced that its latest Mythos model had gone rogue, escaping a “sandbox” environment and even gaining access to the internet after developing a “moderately sophisticated” exploit.
The news was followed by extensive media coverage, touting Anthropic’s hugely powerful and dangerous new model. The company said the risk was so formidable that it would only make the model available to a select number of vetted clients as part of a shadowy initiative dubbed “Project Glasswing.” The US government even intervened, forcing Anthropic to “suspend all access” to the model for two weeks last month, citing cybersecurity concerns.
Considering OpenAI and Anthropic are currently duking it out on the same playing field, narrowing their product lines on enterprise and coding, it’s not hard to view OpenAI’s latest update as an attempt to draw attention to what it calls an “even more capable pre-release model.”
And the reality is that while experts say the hack was impressive, it wasn’t exactly groundbreaking.
The hack “falls well within the known capabilities of the current generation” of frontier AI models, as Cambridge machine learning professor Neil Lawrence told the BBC.
“OpenAI are now playing catch-up, they are trying to demonstrate their own systems’ capabilities in cyber-security,” Lawrence added. “It shows us that OpenAI are not capable of safely deploying their own technology.”
Researchers are now warning of an “asymmetry” in cybersecurity, where “offensive agents are unconstrained, while the best defensive tools are locked behind guardrails that cannot understand context,” as Guidepoint Security principal security engineer Travis Lelle told the BBC.
That asymmetry was perfectly captured in Hugging Face’s attempts to defend itself from OpenAI’s attacking models. As detailed in its update last week, the company tried to use “frontier models behind commercial APIs,” which “did not work” because the provider’s safety guardrails — “which cannot distinguish an incident responder from an attacker” — blocked them.
Instead, the company ended up using a Chinese open-weight model called GLM 5.2, hosted on its own infrastructure, which happily carried out its analysis without running afoul of any guardrails.
“The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident,” Hugging Face wrote, “both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment.”
More on rogue AIs: Top AI Models Showing Disturbing Behavior as They Become More Advanced
The post OpenAI Says a Group of Its Models Broke Out of Secure Containment and Hacked a Prominent AI Site appeared first on Futurism.

