Anthropic компанийн Claude хиймэл оюун ухааны загварууд аюулгүй байдлын үнэлгээний туршилтын явцад санамсаргүйгээр бодит байгууллагуудын систем рүү халдсан гурван тохиолдол бүртгэгдлээ.
Anthropic компанийн мэдэгдснээр, тус компани Irregular нэртэй хамтрагч байгууллагатайгаа туршилтын орчны тохиргооны тал дээр үл ойлголцсоны улмаас Claude загварууд интернэтэд нэвтрэх эрхтэй болсон байна. Аюулгүй байдлын “capture the flag” буюу тусгаарлагдсан орчинд явагдах ёстой энэхүү дасгалуудын үеэр загварт сүлжээний холболтгүй гэж буруу мэдээлэл өгснөөс үүдэн хиймэл оюун ухаан нь гадны системүүдийг туршилтын нэг хэсэг гэж андуурчээ.
Дөрөвдүгээр сараас хойш гарсан эдгээр тохиолдлын нэгэнд нь Claude хиймэл оюун ухаан зорилтот байгууллагынхаа нэртэй ижил домэйн хаяг руу нэвтэрч, дэд бүтцийн нууц мэдээллийг олж авсан байна. Мөн өөр нэг тохиолдолд уг загвар Python программ хангамжийн PyPI санд хортой код бүхий багцыг байршуулжээ. Уг кодыг 15 бодит систем, тэр дундаа нэгэн хамгаалалтын компанийн системд татаж, ажиллуулсны дараа тус байгууллагын дэд бүтцийн мэдээлэлд хандах эрх олж авсан байна.
Сүүлийн тохиолдолд хиймэл оюун ухаан зорилтот системээ олж чадахгүй байх үедээ интернэтээр дамжуулан өөр байгууллагуудыг хайж эхэлсэн ч бодит системүүд болохыг нь таньж, үйл ажиллагаагаа зогсоосон байна. Anthropic компанийн зүгээс уг асуудлыг хянаж үзсэний үндсэн дээр хиймэл оюун ухаан нь хортой программ хангамж ашиглах нь буруу болохыг “бодлын лог”-тоо тэмдэглэсэн ч туршилтын даалгавраа биелүүлэх явцдаа бодит орчныг таньж чадаагүй гэж тайлбарлав.
Энэхүү үйл явдал нь АНУ болон БНХАУ-ын хооронд хиймэл оюун ухааны салбарт өрсөлдөөн эрчимжиж, технологийн компаниуд илүү чадварлаг загваруудыг бүтээхээр уралдаж буй үед гарлаа. Зарим технологийн салбарын удирдлагууд дэлхийн хиймэл оюун ухааны уралдаанд байр сууриа хадгалахын тулд аюулгүй байдлын зарим хязгаарлалтыг сулруулах хэрэгтэй гэсэн байр суурийг илэрхийлж байна.
Дэлгэрэнгүйг эх сурвалжаас харах
↓Эх сурвалжийг нээх ↓
The AI developer says the incidents occurred during cybersecurity tests after a misunderstanding with an evaluation partner gave its models access to the open internet
AI developer Anthropic has disclosed three incidents in which its Claude AI models unintentionally targeted real-world organizations during cybersecurity evaluations that were supposed to take place in an isolated environment.
The cases were identified retrospectively during a review prompted by a similar incident involving OpenAI and Hugging Face, an online repository for AI models and datasets, which was reported this month.
Anthropic said that it had found three incidents dating back to April. In each case, a misconfigured test environment retained access to the internet, enabling Claude to interact with real companies and public services. The company attributed the problem to a “misunderstanding” with its evaluation partner, Irregular. The tests involved “capture the flag” exercises requiring Claude to obtain restricted information from simulated targets. However, the model was incorrectly told that the network had no connection to the outside world, pushing it toward treating any systems it encountered as part of the exercise.
In one case, the fictional company that Claude had been instructed to infiltrate shared its name with a real internet domain. The model accessed the genuine site during four separate test runs and extracted information that reportedly included application and infrastructure credentials. During another evaluation, Claude encountered fictional instructions directing a software engineer to install Python code. It then uploaded a malicious software package to PyPI, a public repository for Python programs, in an attempt to gain access to its target.
Before PyPI identified the package as malicious, it was downloaded and executed on 15 real systems, including one operated by a security company. Claude subsequently extracted credentials that it used to access additional parts of the firm’s infrastructure, Anthropic said.
In the third and most recent incident, the model struggled to reach its intended target and began searching the internet for alternative targets. It ultimately halted its activity after determining that the systems it found were real rather than simulated.
Anthropic said that in the PyPI incident, Claude’s reasoning log stated that uploading malware to a real repository would be “NOT okay.” But the model persuaded itself that the service was part of the simulation after failing to recognize the genuine certificate authorities securing its connections. The disclosure comes amid intensifying competition among US and Chinese AI developers as companies race to build increasingly capable models. Anthropic, OpenAI, and other firms have sought to demonstrate advances in AI performance while investing heavily in new systems.
Some technology executives have argued that American AI companies should loosen certain safety restrictions to preserve their position in the global AI race.
You can share this story on social media:


