OpenAI хиймэл оюун ухаанаар үүсгэсэн текстэд зориулсан дижитал тэмдэглэгээний шинэ аргыг нэвтрүүллээ

Published:

Энэхүү мэдээ, нийтлэлийг хиймэл оюун боловсруулав.

Европын холбооны хиймэл оюун ухааны тухай хуулийн шаардлагыг хангах зорилгоор OpenAI компани текст таних шинэ технологийг ашиглалтад оруулж байна.

OpenAI компани API хэрэглэгчиддээ зориулан загваруудынхаа гаргасан текстэд дижитал тэмдэглэгээ хийх сонголтыг нэвтрүүллээ. Ирэх хэдэн долоо хоногийн дотор Европын холбооны нутаг дэвсгэрт ChatGPT болон Codex загваруудын гаргасан текстүүдэд энэхүү тэмдэглэгээг албадан хэрэглэхээр төлөвлөж байна. Anthropic компани Google-ийн SynthID-Text алгоритмыг ашиглан дэлхий даяар ижил төстэй арга хэмжээ авч буйтай харьцуулахад OpenAI-ийн энэхүү шийдэл нь одоогоор зөвхөн Европын зах зээлд чиглэж байгаа юм.

Тус компани textGrain нэртэй технологийг ашиглан үг сонголтын статистик үзүүлэлтэд үл үзэгдэх дохио нэмж байна. OpenAI энэхүү тэмдэглэгээг илрүүлэх хэрэгслийг зөвхөн батлагдсан судлаачид болон байгууллагуудад олгохоор шийдвэрлэжээ. Гэсэн хэдий ч 200 үгтэй богино хэмжээний текстэд тэмдэглэгээг илрүүлэх магадлал 80 хувь байдаг бол 400 үгтэй текстэд 95 хувьд хүрдэг байна.

Технологийн гүйцэтгэл нь математик болон програмчлалын код зэрэг функциональ текстүүд дээр илүү сул байгааг шинжээчид тэмдэглэж байна. Түүнчлэн, үгийн санг өөрчлөх буюу синонимоор солих энгийн аргуудад энэхүү тэмдэглэгээ амархан эвдэрдэг болох нь тогтоогджээ. Тухайлбал, 400 токен бүхий текстэд үгсийн 25 хувийг солиход тэмдэглэгээг илрүүлэх чадвар 17 хувь хүртэл буурдаг байна.

Дэлгэрэнгүйг эх сурвалжаас харах

↓Эх сурвалжийг нээх ↓

OpenAI is now offering API customers the option to watermark the output of certain models in order to comply with the letter of Europe’s AI Act. And within a few weeks, the biz plans to apply a form of digital watermarking – like it or not – to eligible ChatGPT and Codex text output in the European Union. But unlike rival Anthropic, it isn’t applying its watermarking technique to content generated outside of Europe, though API customers can choose to have the marks applied to content made anywhere. Moreover, OpenAI is using technology that misses a fair number of instances of AI-generated text and can easily be tricked. The EU AI Act requires generative AI service providers to make model output identifiable in a way that’s readable by machines. Anthropic was the first major AI company to announce its approach to AI Act compliance, which involves the application of Google’s SynthID-Text algorithm. It’s doing so globally. Microsoft and Meta have been working on similar technology for image-based AI because the EU is not the only market in which people have raised concerns about content provenance and the potential to use AI-generated content to produce misinformation. OpenAI already applies other content provenance signals to AI-generated images, audio, and video. In keeping with the letter of the AI Act, OpenAI is now offering a text watermarking approach called textGrain [PDF] that “adds an invisible statistical signal to the model’s word choices. Our detector looks for that signal to assess whether a passage contains an OpenAI watermark.” Large language models predict word tokens one at a time from a probability distribution of potential words. By picking a different word to emit periodically – ideally a synonym – the resulting text deviates from the statistical pattern produced by an unbiased model. A watermark detector, which OpenAI is making available only to approved researchers and organizations, should flag that vocabulary divergence. That’s the theory, but in practice results vary. The target error rate is one percent, but in short passages of 200 words, the detector may only flag 80 percent of the watermarks, compared to 95 percent in a 400-word passage. The results are also worse in texts that are functional, like passages about math or code, as substitutions are more likely to introduce errors. The textGrain paper does not address the likelihood that a statistically altered word choice might alter the meaning of a given passage. That’s a possibility if the algorithm were to alter some salient fact or figure. Anthropic’s approach, which also replaces certain output words though using a different algorithm, brings similar risks. In any event, OpenAI’s watermark is easily defeated by word substitution. In tests on 400-token passages, swapping out 10 percent of the words with synonyms reduced watermark detection from 92 percent to 66 percent. And when replacing 25 percent of words, the watermark signal could only be detected in 17 percent of cases. Now witness the firepower of this barely adequate but compliant AI detector… ®

1 сэтгэгдэл

  1. Хиймэл оюун ухаанаар үүсгэсэн текстэд тэмдэглэгээ хийх шинэ арга нь үнэхээр сонирхолтой юм. Та үг солих, синоним ашиглах нь тэмдэглэгээг эвдэхэд хэр их нөлөөлдөг гэж бодож байна? Энэ технологийн өргөн ашиглалт ямар эрсдэл дагуулж болох вэ?

Та юу гэж бодож байна?

Сэтгэгдлээ оруулна уу!
Please enter your name here

MFC.mn сайтад сэтгэгдэл оруулахад анхаарах зүйлс

Холбоотой

spot_img

Шинэ

spot_img