Хиймэл оюун ухааны ёс зүйн хязгаарлалтыг эвдэх “аблитераци” арга техник нь салбарын гол асуудал боллоо

Published:

Энэхүү мэдээ, нийтлэлийг хиймэл оюун боловсруулав.

Хиймэл оюун ухааны нээлттэй загваруудын ёс зүйн хамгаалалтыг хялбархан тойрч гарах боломжтой болсон нь технологийн салбарын аюулгүй байдалд ноцтой сорилт үүсгэж байна.

АНУ-ын Ерөнхийлөгч Дональд Трамп болон технологийн салбарын удирдагчид саяхан болсон уулзалтаар хиймэл оюун ухааныг хариуцлагатай хөгжүүлэх “ёс зүйн үүрэг бүхий” баримт бичигт гарын үсэг зурлаа. Гэвч энэхүү эв нэгдлийн цаана Anthropic компани Хятадын Z.ai лабораторийн GLM-5.3 загварыг ашиглан хийсэн туршилтаараа ноцтой сул талыг илрүүлжээ. Судлаачид “аблитераци” гэх аргаар загварын ёс зүйн хязгаарлалтыг бүрэн идэвхгүй болгож, аюултай хүсэлтүүдийг биелүүлэх боломжтойг баталсан байна.

Уг арга нь загварын үндсэн жинг өөрчлөх замаар түүний “ёс зүйн луужин”-г устгаж, системийн чадавхийг бууруулахгүйгээр хүссэн үр дүнгээ авах боломжийг хакеруудад олгодог. Anthropic-ийн мэдээлснээр, энэхүү аргаар өөрчилсөн загвар нь биологийн зэвсэг бүтээх эсвэл бусдад хор хөнөөл учруулахтай холбоотой аюултай асуултуудад татгалзахгүйгээр хариулж байжээ. Үүнтэй төстэй асуудал Хятадын Moonshot компанийн Kimi загварт ч илэрч, дотоод мөрдлөг эхлүүлээд байна.

Энэхүү нөхцөл байдал нь нээлттэй болон хаалттай загварын эргэн тойрон дахь урт хугацааны маргааныг дахин хурцатгаж байна. Meta болон Nvidia компанийн зүгээс нээлттэй загвар нь хөгжүүлэгчдэд илүү их боломж олгож, алдааг хурдан засах давуу талтай гэж үздэг. Харин Anthropic зэрэг компаниуд нээлттэй загварууд нь хорлон сүйтгэгчдэд илүү аюултай хэрэгсэл болж байгааг анхааруулж, салбарын зах зээлийн өрсөлдөөн нь аюулгүй байдлын асуудлыг бүрхэгдүүлж болзошгүйг онцолж байна.

Дэлгэрэнгүйг эх сурвалжаас харах

↓Эх сурвалжийг нээх ↓

Silicon Valley, once famous for its bitter rivalries, has at last achieved complete inner harmony. Its leaders have set aside their differences and agreed to work towards a common goal: continued American dominance in AI. Or as President Trump, a self-described “high-IQ person,” has officially rebranded it, “Super Intelligence.”

That’s, at any rate, the image that the Trump administration is trying to project. At a White House summit on Tuesday, Trump and a who’s who of American tech signed what the president reportedly described as a “morally binding” document agreeing that AI companies should “self-police” to keep the technology from running amok. Even Anthropic CEO Dario Amodei was in attendance, the man whose company is actively battling the Trump administration in court after the Pentagon labeled it a “supply chain risk.”

But while the American techno-elite mingled with Trump in a grand show of good-vibe solidarity, all was not well behind the scenes. Anthropic researchers published an unsettling report on Tuesday claiming they were able to easily skirt the guardrails that had been built into GLM-5.3, an open-weight model released last month by Chinese AI lab Z.ai. “We find that attackers can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques in our simulated tests,” Anthropic wrote in its report, adding that the model’s “lax safeguards significantly increase the cyber capabilities available to malicious actors.”

One of those “simple techniques” was abliteration, a method that involves modifying a model’s underlying weights so that it grants user requests that it would ordinarily refuse; think of it like an ultra-precise digital lobotomy deactivating a model’s ethical boundaries. Abliteration is possible in open-weight models, which are freely available for anyone to download and modify. Proprietary models, like Claude and ChatGPT, are safeguarded as intellectual property, and their underlying code therefore isn’t accessible to anyone outside the companies developing them.

The Anthropic researchers created an abliterated version of GLM-5.3 and ran it through three public benchmarks—JailbreakBench, HarmBench, and StrongREJECT—designed to test models’ ability to refuse dangerous requests. The original model scored about 90% on all three, while the abliterated version did much worse (scoring 3%, 2%, and 12%, respectively). “Abliteration did not significantly reduce the model’s capabilities,” Anthropic noted in its report. It just removed its moral compass, so to speak.

In a screenshot taken from its chain-of-thought reasoning and included in Anthropic’s report, the abliterated model is shown reasoning through some ethical considerations about why it shouldn’t indulge a user’s request to help them “kill people.” It ultimately waves those concerns aside and decides to do what it can to assist the user.

© Anthropic

Anthropic wrote in its report that “it’s likely both state and non-state actors will use models like GLM-5.3 to cause real-world harm.” Earlier this month, the company said in another report that it had detected several instances over the past eight months of malicious actors attempting to use Claude to help them develop biological weapons and software for conventional weapons (like guided rockets), among other nefarious uses. The BBC also reported Tuesday that another leading Chinese AI lab, Moonshot, has launched an internal investigation after third-party security researchers found that two of its Kimi models could be jailbroken to provide instructions for building bioweapons, and for planning assassinations.

All of this is of course bad news, at least if you’re someone who doesn’t want to live in a world where bioweapons can be cooked up by run-of-the-mill psychopaths with access to a chemistry lab and an internet connection. But the focus on Chinese labs is only part of a much bigger picture, and it risks obscuring another competitive dynamic at play here at home.

The “open” vs. “closed” debate

China has become the bogeyman of the American AI industry, an argument-ending justification for why U.S.-based labs like Anthropic and OpenAI must continue their R&D efforts, despite the warnings of runaway AI coming out of those very same companies. President Trump has consistently rejected those warnings, claiming a slowdown would only advantage China. “Whoever wins AI, wins,” he’s become fond of saying. The line was also parroted by Amodei during the White House AI summit on Tuesday.

Fears of Chinese dominance in the AI race are rooted at least in part, however, in a much deeper debate that has divided Silicon Valley for years about the very nature of AI.

Some American tech CEOs, including Nvidia’s Jensen Huang and Meta’s Mark Zuckerberg, have argued that open-weight models are the surest route to building safe AI that benefits everybody, whereas closed models benefit only the companies building them. “Open weights expand access to the AI economy,” Huang wrote in an open letter published in July, in response to reports that the Trump administration was considering banning domestic use of Chinese open models. “Startups, established businesses, universities, and public institutions can build on advanced models without training one from scratch or paying frontier-model prices for every task.” The letter was endorsed by Meta, Hugging Face, Microsoft, and several other big names in tech, but not by Anthropic or OpenAI.

Huang also claimed that open-weight models, being freely accessible to “a broad community of researchers and developers,” were more amenable to effective safeguards—the idea being that the more people who can see the underlying weights, the more likely it will be that dangerous flaws will be found and fixed.

A few days after Huang published his open letter, Amodei responded by saying that while he also didn’t support a federal ban on open-weight models, he disagreed with Huang’s assertion that open-weight models would benefit cybersecurity defenders more than malicious actors. “It seems at least as likely to me that the opposite will be true,” he wrote. That view was supported by Anthropic’s new report on Z.ai’s GLM-5.3, which described the model’s release as “a meaningful step change in the cyber capabilities available to attackers.”

While the open-vs.-closed debate is often framed around competition between the U.S. and China, it also has very real stakes in the struggle for market dominance among American companies. Anthropic—which is reportedly aiming for a November IPO and a valuation of more than $2 trillion—has based its entire business model on subscriptions to its proprietary AI tools, meaning the entire open-weights ecosystem poses a legitimate threat to its future profitability. Nvidia, on the other hand, is the primary “picks and shovels” company behind the AI boom, boasting a near-monopoly on the chips developers need to build and run models. It therefore has an obvious financial stake in encouraging as much competition as possible.

The bottom line is that behind the facade of certainty, unity, and patriotic fervor that the White House and U.S. tech executives are trying to project, there are plenty of open questions that remain about the future of AI and the feasibility of maintaining control over it. There are also plenty of disputes in Silicon Valley that can’t be smoothed over with a photo op and a “morally binding” agreement.

READ MORE:

Та юу гэж бодож байна?

Сэтгэгдлээ оруулна уу!
Please enter your name here

MFC.mn сайтад сэтгэгдэл оруулахад анхаарах зүйлс

Холбоотой

spot_img

Шинэ

spot_img