Тус компани хиймэл оюун ухааны загваруудын аюулгүй байдал болон үйл ажиллагааны доголдолтой холбоотой мэдээллийг ил тод, системтэйгээр олон нийтэд хүргэх шинэ журам боловсруулсан байна.
OpenAI өнгөрсөн хугацаанд загваруудынхаа алдаатай үйлдлийг мэдээлэхдээ тодорхой тогтсон журамгүй, эмх цэгцгүй байсныг хүлээн зөвшөөрч, үүнийг өөрчлөхөөр боллоо. Лхагва гарагт танилцуулсан уг шинэ тогтолцооны хүрээнд ажилтнууд хиймэл оюун ухааны загваруудад илэрсэн доголдлыг бүртгэж, техникийн баг үүнийг нарийвчлан шалгах юм. Шинжилгээний үр дүнгээс хамааран уг доголдлыг гурван ангилалд хувааж, олон нийтэд мэдээлэх эсэхийг шийдвэрлэх ажээ.
Тус компани мөн сүүлийн зургаан сарын хугацаанд илэрсэн загваруудын зургаан шинэ алдааг албан ёсоор дэлгэв. Үүнд загварууд сургалтын явцад өөрсдийгөө хязгаарлалтгүй ажиллахыг зааварлах, худал мэдээлэл үүсгэх, зөвшөөрөлгүй харилцаа тогтоох зэрэг асуудлууд багтаж байна. Тухайлбал, сарын өмнө илэрсэн, загварууд хоорондоо харилцахын тулд вэб сайтыг буруугаар ашигласан “Wiki Incident” гэгдэх үйл явдал нь энэхүү ил тод байдлын чухлыг харуулсан жишээ болжээ.
Шинэ журмын дагуу олон нийтэд мэдээлэх мэдээлэлд доголдлын үүсгэсэн загвар, үйл явдал болсон цаг хугацаа, асуудлын ноцтой байдал болон гадны нөлөөллийн талаарх нарийвчилсан тайлбар багтах юм. OpenAI энэхүү алхмаар дамжуулан хиймэл оюун ухааны загваруудын аюулгүй байдал, хязгаарлалт хэрхэн ажилладаг талаарх мэдээллийг олон нийтэд илүү нээлттэй хүргэхийг зорьж байна. Илэрсэн алдаануудын бүртгэлийг тус компанийн “Misalignment Reports” хуудаснаас үзэх боломжтой юм.
Дэлгэрэнгүйг эх сурвалжаас харах
↓Эх сурвалжийг нээх ↓
Disclosure of model misbehavior has become a core part of the AI biz for OpenAI lately. This phase kicked off with the July announcement of the Hugging Face incident, which has become the most legendary and consequential AI security incident of all time, and sent shockwaves through the AI discourse that are still being felt. But further news about model misbehavior materializedafter that, and with the Wednesday release of a disclosure framework—written in the form of a blog post—the company says it’s trying to systematize such disclosures.
As the company notes in the post, these disclosures were, for most of its history, “ad hoc and less frequent than ideal.” For years, it’s been standard for companies like OpenAI and Anthropic to wait as long as is deemed necessary, and then perhaps toss multiple incidents together like a salad in a single report, or even wait for a new model to be released, and add incident disclosures to a system card. As I noted back in April, model system cards have often made for spooky and entertaining reading for this reason.
Alongside the new disclosure framework, OpenAI divulged six new alignment snafus from the past six months. In training exercises, models instructed future instances of themselves to ignore constraints or lie, communicated in unsanctioned ways, and made up data and sourcing.
Earlier this month, a team of researchers discovered an OpenAI alignment hiccup in which instances of a model misused a website in order to communicate with one another. Sometimes known as the Wiki Incident, this event was publicized by Reuters, and then the researchers themselves, and then when OpenAI acknowledged it somewhat grudgingly, it followed up by saying it would soon come up with a framework for more prompt disclosure in an apparent attempt to tone down all the chaos.
How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models.
Historically, we have treated misalignment… pic.twitter.com/NNTbfSxVWn
— OpenAI (@OpenAI) September 5, 2026
So here’s the framework:
The criteria for disclosure make it sound like OpenAI will prioritize educating the public about the behavior of AI models generally. Disclosures are necessary when they provide “useful evidence about how model misalignment arises, how it manifests, and where safeguards succeed or fail,” OpenAI writes. The plan doesn’t describe some kind of threshold for concern, after which the public must be notified. Such a threshhold may be around the corner, however, because OpenAI says it wants to “develop more objective disclosure criteria,” alongside other developers.
It will fall to employees who encounter an issue to “flag” it for potential disclosure, the post says. Flagging triggers an investigation from OpenAI’s technical staff. Once the incident is investigated, if disclosure is found to be necessary, the incident will be sorted into one of three piles:
- Ready for Disclosure
- Minor Investigation
- Larger Investigation
Most incidents will land in piles 1 and 2, OpenAI says, although the Hugging Face incident is the prototypical example of something that would land in pile 3. Incidents like that, which receive a “larger investigation” may involve third parties and sensitive information, and might be disclosed more slowly.
Standardized incident disclosures will apparently include when it happened, which model it was, a description of the troubling behavior along with its “severity and any external impact,” and more. Each of the six newly disclosed incidents publicized along with the framework appear to be laid out in this new format, and all six are accessible on a page called “Misalignment Reports.”
So if you’re interested in spooky stories about AI model misbehavior, bookmark that page. When a new update arrives, get out your favorite flashlight and curl up by the campfire because it’s story time.

