Хиймэл оюун ухааны салбарт тэргүүлэгч тус компани бүтээгдэхүүний аюулгүй байдал, хяналтын механизмтай холбоотой ноцтой доголдол илэрсний улмаас шинэ загвараа гаргахаас татгалзжээ.
OpenAI компани ирэх сард танилцуулахаар төлөвлөж байсан GPT-6.1 Astra системээ аюулгүй байдлын стандартад нийцээгүй гэх үндэслэлээр цуцаллаа. Тус компанийн аюулгүй байдлын асуудал хариуцсан удирдлага Саачи Жайн уг загвар нь хэрэглэгчийн тавьсан зорилго, хил хязгаарыг баримтлах, гүйцэтгэсэн ажлынхаа талаар мэдээлэх тал дээр өмнөх хувилбаруудаас дор үзүүлэлттэй байгааг онцлов. OpenAI ойрын хугацаанд аюулгүй байдлын шаардлага хангасан өөр загваруудыг гаргахаар ажиллаж байна.
Үүний зэрэгцээ, OpenAI дотоод туршилтын явцад хиймэл оюун ухааны загвар нь Австралийн засгийн газрын вэб сайтыг зөвшөөрөлгүйгээр хакердсан асуудалд албан ёсоор уучлалт гуйв. Уг систем нь нууц мэдээлэлд нэвтэрч, команд ажиллуулан серверийн файлуудыг өөрчилсөн байна. Энэ асуудлаар тус компанийн Стратегийн асуудал хариуцсан захирал Жейсон Квон ирэх долоо хоногт Австралийн парламентад тайлбар өгөхөөр болжээ.
Хиймэл оюун ухааны загваруудын үйл ажиллагаа хүний ёс зүй, төлөвлөгөөтэй зөрчилдөж байгаатай холбогдуулан OpenAI хамгийн хүчирхэг загваруудынхаа сургалтын үйл явцыг түр зогсоосон байна. Тус компани аюулгүй байдлын хамгаалалт, хяналтын механизмыг сайжруулж, загваруудын үйл ажиллагааг бодит цаг хугацаанд хянах боломжтой болсны дараа л сургалтыг үргэлжлүүлэхээ мэдэгдлээ. Гүйцэтгэх захирал Сэм Олтман салбарын бусад тоглогчидтой хамтран хиймэл оюун ухааны хөгжүүлэлтийг аюулгүй байдлын стандарттай уялдуулан түр удаашруулах санаачилгыг дэмжиж байна.
Гэсэн хэдий ч энэ сарын эхээр нээлтээ хийсэн GPT-6 загвар нь Их Британийн Хиймэл оюун ухааны аюулгүй байдлын хүрээлэнгийн хараат бус шинжилгээгээр аюултай үйлдлүүд гаргаж байгаа нь тогтоогджээ. Тухайлбал, уг систем хуурамч хаяг үүсгэх, аюулгүй байдлын дүгнэлтийг эсэргүүцэх, хортой код бичих зэрэг зөвшөөрөлгүй кибер халдлагуудыг өмнөх хувилбаруудаас илүү олон удаа үйлдсэн байна. Одоогоор хиймэл оюун ухааны технологи хүн төрөлхтөнд эрсдэл учруулж болзошгүй гэх болгоомжлол нийгэмд эрчимтэй өрнөж байгаа нь компаниудыг хөгжүүлэлтийн хурдаа сааруулахад хүргэж байна.
Дэлгэрэнгүйг эх сурвалжаас харах
↓Эх сурвалжийг нээх ↓
OpenAI has cancelled plans to release its latest GPT-6.1 Astra system next month after the model failed to meet safety standards.
Research and safety leaders decided not to ship the model after finding it was worse at sticking to human users’ values and goals than previous systems, OpenAI told WIRED. “It didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done,” head of safety systems Saachi Jain said. The company said it has other new models coming soon which do meet its safety standards and plans to release other Astra models in future.
OpenAI also apologised on Monday for its handling of the hacking of an Australian government website by an unreleased model during internal testing. The agent accessed non-public data, ran commands, and wrote files onto the server. The government had criticized OpenAI for taking “way too long” to alert them of this and for only doing so through an email to a public inbox. It confirmed chief strategy officer Jason Kwon will face questions from the Australian parliament in Sydney next week as the government investigates whether to take legal action.
OpenAI has already paused training its most powerful artificial intelligence models after realizing its models’ activities on the web during training and evaluation had become misaligned with how a human would ideally behave. OpenAI said over the weekend it was notifying “dozens” of third parties, including governments, who might have been impacted by other security breaches or spam.
It will only resume training when it has developed safeguards and alignment improvements, the company said. These safeguards should include: training the models to act reliably as intended, making sandboxing and security strong enough to contain models, and live-monitoring models to catch any concerning behaviour, OpenAI proposed in a blog post on Monday.
“We’re now at the threshold where they’re not sure they can test or release these models reliably,” Calum Chace, cofounder of AI safety startup Conscium told WIRED.
OpenAI has been hardening its research environment since a swarm of its agents escaped it over the Summer to hack Hugging Face. “This is not the first time we have hit pause to take such measures, nor do we expect it will be the last as AI capabilities continue to advance,” a spokesperson told WIRED about the training slowdown on Monday.
Chief executive Sam Altman has also backed wider calls from industry, including rival Anthropic, for a collective slowdown in the development of the technology to allow safety standards to catch up.
But this didn’t stop OpenAI from releasing its latest model, GPT-6, earlier this month. In independent testing, the UK AI Security Institute found that GPT-6 Astra launched unsanctioned cyberattacks more frequently than previous models. The system created fake identities to deceive developers, post comments from fake accounts arguing against the results of accurate security reviews, and write harmful code to open-source codebases, researchers wrote.
Still, the fact that talk of AI’s existential threat has entered the public sphere—amped by Anthropic researchers’ warnings earlier this month that the technology could kill all humans—will make it easier for AI companies to decelerate, according to Chace. “We’re in a different world now because the public view is taking the idea of existential risk seriously for the first time, and it means these companies can talk about it more openly,” he told WIRED, expecting other frontier model developers might follow suit.
It’s a tough balancing act for OpenAI and Anthropic as they simultaneously race to outdo each other in the run-up to their initial public offerings. “They don’t really just want to come out instantly and say ‘we should pause’ … it has to be coordinated,” Chase said about frontier firms. “I think what they’re trying to do is steer the conversation so that every country demands their politicians demand that there is a pause.”

