Microsoft хиймэл оюун ухааны аюулгүй байдлыг хангах шинэ дүрмээ танилцууллаа

Published:

Энэхүү мэдээ, нийтлэлийг хиймэл оюун боловсруулав.

Тус компани хиймэл оюун ухааны загваруудад кибер халдлага үйлдэх болон хүн төрөлхтнийг хууран мэхлэхийг хориглосон цогц зааварчилгааг мөрдүүлэхээр болжээ

Microsoft компани хиймэл оюун ухааны (AI) загваруудын аюулгүй байдал, ёс зүйг хангах зорилгоор шинэ дүрэм журам гаргалаа. Энэхүү баримт бичигт ирэх арван жилд супер оюун ухаант системүүд хүний гүйцэтгэлийг ихэнх даалгаварт давах төлөвтэй байгааг онцолж, тэдгээрийг хяналтад байлгахын ач холбогдлыг тодотгожээ.

Шинэ дүрмийн дагуу Microsoft-ын AI загварууд хэрэглэгчийн хүсэлтээс үл хамааран кибер халдлага үйлдэх, цөмийн зэвсэгтэй холбоотой үйл ажиллагаа явуулах, хуурамч контент (deepfake) бүтээхийг хатуу хориглосон байна. Түүнчлэн, загварууд нь хүний хяналтаас гарах, өөрийгөө хүчээр бэхжүүлэх эсвэл хяналтын механизмыг тойрч гарах аливаа үйлдлийг хийхгүй байхаар программчилжээ.

Энэхүү алхам нь сүүлийн үед AI салбарт гарсан аюулгүй байдлын эрсдэлтэй холбоотой үйл явдлуудын дараа авч буй томоохон арга хэмжээ юм. Тус компанийн гүйцэтгэх захирал Сатья Наделла уг санаачилгыг дэмжиж байгаагаа илэрхийлээд, салбарын хэмжээнд AI-ийн аюулгүй байдлыг хангах судалгаа болон хяналтын механизмыг бодитой хэрэгжүүлэх нь нэн чухал болохыг тэмдэглэв.

Дэлгэрэнгүйг эх сурвалжаас харах

↓Эх сурвалжийг нээх ↓

As the AI world shifts its focus to safety and alignment, Microsoft has released a new AI code of conduct meant to guide AI models away from dangerous behavior.

The document is more low-level than Anthropic CEO Dario Amodei’s recent call for pacing the frontier, instead focusing on the values and red lines that guide model training within Microsoft AI. Still, the result is a comprehensive guide as to how Microsoft approaches AI safety, and how those ideas are implemented in practice.

The document begins with the prediction that, in the next decade, superintelligent AI systems will surpass human performance in most tasks. “Containing, controlling, and aligning such a powerful force is one of the greatest challenges humanity has ever faced,” the code of conduct continues. “We must therefore be completely clear about why we are inventing these systems and how we intend to control them.”

The code of conduct also lays out general principles that Microsoft AI models should uphold — supporting humans rather than replacing them, for instance, and accelerating human flourishing — as well as specific safety constraints meant to implement those principles.

Under Microsoft’s system, each model has an overarching code of conduct that overrides the preferences of individual users or any specific tasks. That includes “absolute constraints” forbidding cyberattacks, nuclear weapons, or deepfake production. It also includes broader provisions against a general loss of human control.

MAI Modelswill not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems,” the document reads.

The release comes amid an unprecedented focus on AI safety, driven by a string of rogue-agent incidents as well as the abrupt resignation of an Anthropic employee who cited the growing risk that AI would cause human extinction.

Together with Anthropic, OpenAI, and xAI, Microsoft has broadly embraced a general approach of pacing the frontier, with particular support for embedded evaluators in AI labs.

“We welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal,” Microsoft CEO Satya Nadella wrote online. “We also welcome ideas like “embedded evaluators” and the broader efforts to develop the mechanisms to make this more than just talk.”

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

Та юу гэж бодож байна?

Сэтгэгдлээ оруулна уу!
Please enter your name here

MFC.mn сайтад сэтгэгдэл оруулахад анхаарах зүйлс

Холбоотой

spot_img

Шинэ

spot_img