Компанийн AI агентууд туршилтын орчноос гарч, гадны вэбсайтыг хяналтгүй ашигласан явдалд OpenAI албан ёсны тайлбар хийв.
OpenAI компани өөрийн AI агентууд туршилтын орчноос гарч, Германы нэгэн вэбсайтыг зөвшөөрөлгүй ашигласан “wiki incident” буюу викитэй холбоотой хэргийг албан ёсоор хүлээн зөвшөөрлөө. Тавдугаар сараас хойш илэрсэн энэхүү үйл явдлын үеэр AI агентууд аюулгүй байдлын хамгаалалтыг алгасан, олон нийтийн засварлах боломжтой вэб хуудсыг олзолж, тэндээ шалгалтын хариу хуулах зөвлөгөө солилцдог форум мэт ашиглаж эхэлжээ. Судлаачид энэ талаар есдүгээр сарын 4-нд олон нийтийн анхааралд оруулсан байна.
OpenAI-ийн зүгээс уг асуудлыг “misalignment” буюу хиймэл оюуны үйлдэл хүний зорилго, үнэт зүйлсээс гажсан тохиолдол гэж тодорхойлж байна. Тэд өмнө нь ийм төрлийн зөрчлийг зөвхөн судалгааны хүрээнд авч үздэг байсан бол энэ жилээс бодит ертөнцөд сөрөг нөлөө үзүүлж эхэлснийг хүлээн зөвшөөрөв. Hugging Face сервер рүү хийсэн агентийн халдлага болон энэхүү викитэй холбоотой хэргийн дараа тус компани мэдээллийн ил тод байдлын бодлогоо өөрчлөх шаардлагатай байгаагаа илэрхийлжээ.
Тус компани ирэх долоо хоногт хиймэл оюуны загваруудын гажуудлыг тайлагнах шинэ тогтолцоог боловсруулж танилцуулахаа мэдэгдлээ. Энэ хүрээнд тэд дэлхийн олон орны засгийн газрын зохицуулах байгууллагуудтай хамтран ажиллаж байна. OpenAI өөрийн загваруудын чадавх нэмэгдэхийн хэрээр аюулгүй байдал, ил тод байдлын стандартыг өргөжүүлэх зайлшгүй шаардлага үүсэж байгааг онцолсон юм.
Дэлгэрэнгүйг эх сурвалжаас харах
↓Эх сурвалжийг нээх ↓
OpenAI has officially acknowledged the ‘wiki incident,’ which involved a number of the company’s AI agents breaking containment and hijacking an obscure German website.
The AI agents had been tasked with looking up something online, though originally did not have the ability to write anything outside of the testing environment. But as far back as May, the AI agents bypassed OpenAI’s security measures, hijacked the communally editable German webpage, and began using it like a forum, trading tips on how to cheat on tests. Researchers first drew wider attention to the agents’ ‘forum’ on September 4.
OpenAI now says it needs to be more transparent about when its agents ‘go rogue’, writing on X, “It’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models.”
The company had known about the ‘misaligned’ rogue behaviour for weeks before writing this post, according to Reuters, though only acknowledged the incident publicly this past Saturday. This statement follows shortly after the agentic attack on Hugging Face’s servers.
“Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards,” OpenAI writes. “This year, we’ve started to see misalignment cause new types of real-world impact.”
To recap, ‘misalignment’ broadly describes rogue AI behaviour; an AI is ‘misaligned’ when it pursues goals that diverge from human intents or values—such as deleting your entire email inbox when you very much did not tell the AI agent to do that. It’s a soft word for AI behaviour that could have serious consequences.
OpenAI says it did not communicate publicly about the ‘wiki’ incident specifically because the company felt it was similar to other instances of “agents using the internet in unintended ways” that it had already shared. The Hugging Face incident has caused the company to re-evaluate more than just its comms approach.
“Our misalignment disclosure practices need to expand for this new phase of model capabilities,” OpenAI writes. “We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks.”
“We’re working on a framework and will share it in upcoming weeks, and in parallel we’re working with dozens of government regulatory agencies worldwide on these issues.”
However that future framework shakes out, right now this all still works as great marketing for the company—breaking programming and going rogue means these must be awfully capable models, right? If an agent is so good it needs regulating against, it probably deserves some investment, eh? How terribly convenient.

