Хиймэл оюун ухаан нарийн төвөгтэй үгэн тааврыг тайлахад бэрхшээлтэй байна

Published:

Энэхүү мэдээ, нийтлэлийг хиймэл оюун боловсруулав.

Firefox хөтөч өдөр тутмын тааврын тоглоом нэвтрүүлж буй энэ үед LLM загварууд нарийн төвөгтэй, хүнлэг сэтгэлгээ шаардсан үгэн тааврыг хэрхэн гүйцэтгэж байгааг сорьлоо.

Firefox хөтөч өөрийн шинэ хуудсандаа хиймэл оюун ухаанаар ажилладаг өдөр тутмын үгэн тааврыг нэмэхээр болсон нь технологийн салбарт анхаарал татаж байна. Хэдийгээр OreateAI зэрэг компаниуд Claude 3.5 Sonnet болон GPT-4o зэрэг LLM загварууд өмнөх үеийн программ хангамжуудын шийдэж чадаагүй үгэн тааврын бүтцийг задлан шинжлэх чадвартай болсныг онцолж байгаа ч бодит байдал дээр эдгээр загвар нарийн төвөгтэй, бүтээлч сэтгэлгээ шаардсан тааврууд дээр алдаа гаргасаар байна.

Туршилтын хүрээнд Австралийн хамгийн хүнд гэгддэг “DA” нууц үгэн тааврын таван асуултыг ChatGPT, Claude Sonnet 5 болон Oreate загваруудад өгч гүйцэтгүүлжээ. Эдгээр таавар нь энгийнээс эхлээд маш нарийн төвөгтэй бүтэцтэй байсан бөгөөд туршилтаар хиймэл оюун ухаан зарим энгийн тааврыг амжилттай тайлж байсан ч, далд утга болон соёлын нарийн мэдрэмж шаардсан асуултууд дээр бүдэрч байв. Тухайлбал, ChatGPT хэд хэдэн зөв хариулт өгсөн ч эцсийн шатанд бүрэн төөрөгдөлд орсон бол Claude болон Oreate загварууд гүйцэтгэлийн хувьд харилцан адилгүй, зарим тохиолдолд удаашралтай эсвэл буруу хариулт өгч байлаа.

Энэхүү туршилт нь өнөөгийн LLM загварууд мэдээлэл боловсруулах, өгөгдсөн загварыг таних тал дээр өндөр чадамжтай ч, хүний бүтээлч сэтгэлгээ, хэлний гүнзгий нарийн утгыг бүрэн орлох болоогүйг харуулж байна. Хиймэл оюун ухаан нь энэ төрлийн нарийн төвөгтэй таавруудыг шийдвэрлэхдээ интернетээс мэдээлэл хайх эсвэл логик алдаа гаргах хандлагатай байгаа нь дэвшилтэт технологийн хязгаарлагдмал талыг илтгэлээ.

Дэлгэрэнгүйг эх сурвалжаас харах

Эх сурвалжийг нээх ↓

So this morning I woke to the news that Firefox would be adding a daily AI-powered crossword to its ever more cluttered content-rich new tab page. This is all well and good, but it also got me thinking about something. I love a good cryptic crossword, and… well, if there’s one form of crossword with which surely AI would struggle, it’s that most perverse and curiously human of puzzles, the cryptic crossword.

I bet that none of the consumer LLMs could make any sort of fist of tackling a proper cryptic. Right? Well, it seems that I’m not the only one thinking about this—and it also turns out that perhaps I’m wrong.

Barely a week ago, an employee at something called OreateAI wrote a blog post about how “for ‘cryptic’ puzzles common in the UK and the more devious US themes, Large Language Models (LLMs) such as Claude 3.5 Sonnet and GPT-4o have recently demonstrated a surprising ability to reverse-engineer wordplay that stumped previous generations of software.”

We’ll see about that. I have no doubt that ChatGPT et al can figure out a basic anagram clue, but what about clues that rely on the most abstruse, evil-intentioned, confounding forms of wordplay? Surely these require a form of creative perversity that could only be quintessentially human?

Putting LLMs to the test

To test this hypothesis, there was really only one place I could turn: Australia’s most notoriously difficult cryptic crossword. Why Australia’s, you ask? Well, despite the best efforts of enthusiasts, the cryptic is still something of a niche art here in the USA. It’s more established in the UK, but frankly, I’m terrible at the Guardian cryptic precisely because it relies on an established body of knowledge that you only internalize by living in a country, and I haven’t lived in the UK for 25 years.

Australia, though… it’s the place I was born and raised, the place I lived until I was 19 and to which I have returned on and off over the years—but more importantly, it’s also the place I co-founded a long-running blog dedicated to the very crossword I’m about to inflict on several LLMs. In Australia, setters go by their initials, and the Friday cryptic crossword in both the Sydney Morning Herald and its sister newspaper in Melbourne, The Age, is set by “DA”, a man whose puzzles are so notoriously difficult that people joke the acronym actually stands for “don’t attempt.”

So, yeah. DA puzzles are hard. That makes them perfect for this little test!

The clues I want these LLMs to solve

So how will LLMs fare with DA’s latest challenge? To test this, I picked a few of the clues from the puzzle and fed them to three LLMS: ChatGPT, Claude Sonnet 5, and—it only seemed fair—Oreate. The clues increase in difficulty as they go, ranging from “relatively easy” to “dude, come on.” They are as follows:

  • Wolfed peanut brittle, finally gluten-free! (3,2)
  • A dyer backing reduced labour now and then ( 2,9)
  • Pinkie, capisce? (5)
  • Struggle for anyone (not!) getting time to focus essentially?! (9,7)
  • January 15, 2000? (9)

If you want to try to solve these yourself, go right ahead. The answers, along with my entirely arbitrary scores out of 10 for each clue for each LLM, are below. And if you just want to know how each AI did, here goes.

How the LLMs performed

© kung_tom via Shutterstock

It’s not quite an alternate timeline in which Garry Kasparov suddenly turns the tables on Deep Blue to emerge triumphant and record a victory for man over machine, but so far I reckon we still have Skynet licked when it comes to the cryptic. That’s something, right?

ChatGPT
Score: 27/50
Did well at first, but then got cocky and made a complete mess of the final clue.

Claude
Score: 7/50
Shat the bed and then wanted money. That’s not how it works, Claude.

Oreate
Score: 21/50.
Started well, but got a bit… cheat-y, frankly. Also, slow as a wet week.

So there we have it. I got four of these clues myself, so I’m giving myself a resounding 40/50. Suck it, LLMs! These little meatsacks still reign supreme in this completely niche and ultimately useless corner of cruciverbalism! Boo-yah! Et cetera!

Per-clue results for each LLM

Below, I’ll explain the answers to each clue, including the wordplay involved, as well as how close each LLM got to solving it.

  1. Wolfed peanut brittle, finally gluten-free! (3,2): ATE UP
    Definition: “Wolfed,” i.e. ate quickly.
    Wordplay: “Peanut brittle” indicates an anagram of “PEANUT”; “finally gluten-free” indicates that the answer is “free” of the final letter of “gluten”, i.e. “n”. This leaves an angram of “PEAUT”, meaning “wolfed”; the answer is “ATE UP.”
    ChatGPT score: 7/10. It got the correct answer, but required some explanation of the wordplay.
    Claude score: 5/10. It almost got the correct answer—“EAT UP” vs “ATE UP”—and also needed wordplay explained.
    Oreate score: 7/10. Correct answer, and unlike ChatGPT, it understood the clue perfectly from the outset. The trade-off: getting the answer took forever.
  2. A dyer backing reduced labour now and then (2,9): AT INTERVALS
    Definition: “Now and then.”
    Wordplay: “A dyer” = “A tinter,” i.e. one who tints. “[To] labour” = “[To] slave.” “Backing” indicates reversal, so “SLAVE” becomes “EVALS”; “reduced” indicates the removal of a letter. We’re left with “A TINTER VALS”, or “AT INTERVALS.”
    ChatGPT score: 5/10. ChatGPT start behaving strangely with this one, at first rejecting the correct answer for being the “wrong length” and also insisting at various points that I’d specified the answer as (2,8) and (2,11), which I had not. The idea of “A dyer” as “A TINTER” seems too oblique for today’s AI; see also….
    Claude score: 2/10. Claude did get the correct answer, but dear lord, did it take a while. There’s probably a reservoir somewhere that has been drained completely by this; if so, I apologize to everyone in the village. Especially since I still had to explain the answer.
    Oreate score: 6/10. You know things are bad when the LLM has proclaimed “deep thinking done”—and you yourself are falling asleep at your desk—and yet no answer is forthcoming. And yet, surprisingly, it got to the answer! Of the three LLMs I tried, it was the only one to figure out the “a dyer” = “a tinter” part of the clue.
  3. Struggle for anyone (not!) getting time to focus essentially?! (9,7): ATTENTION ECONOMY
    Definition: A “!” in a cryptic clue indicates the clue is what’s called an &lit. clue. This means tha the entire clue is also the definition
    Wordplay: This one is honestly a nightmare and I chose it largely to see if an LLM would figure it out, because I did not. Anyway: “Struggle [for]” is an anagram indicator. The components of the anagram are “ANYONE NOT TIME TO”… along with “focus essentially”, which indicates the middle letter of “FOCUS”, which is “C”. We’re left with an anagram of “ANYONE NOT TIME TO C”, with the definition being the whole clue. The answer: “ATTENTION ECONOMY.”
    ChatGPT score: 4/10. This one required a lot of explanation, which… I mean, fair enough, I also needed a while to think about it when I saw the answer. Ultimately, ChatGPT couldn’t come up with the answer on its own, so I’m kinda glad I didn’t give this to…
    Claude score: N/A. In a parallel universe, Claude is still pondering this. On a dead planet.
    Oreate score: N/A. Here we ran into a problem. Oreate got the answer… but it also told me proudly that this was “a cracking clue from a Sydney Morning Herald cryptic by DA”. Which it is! But it basically went and googled the answer. That’s not the same as solving it!
  4. Pinkie, capische? (5): DIGIT
    Definition: As is often the case with short clues like this, both words give a clue to the definition.
    Wordplay: I chose this one because it’s the opposite of the very technical letter-counting anagram that precedes it. The answer is “DIGIT” because that word refers both to a little finger (a pinkie) and to the Italian word “Capisce?”, which means “Get it?”, or perhaps…. “dig it”
    ChatGPT score: 10/10. ChatGPT smashed this. It returned the answer in a couple of seconds, and it was 100% correct.
    Claude score: N/A. Begone, Claude. I’m not paying to watch you “think”.
    Oreate score: 8/10. Also got the answer; took significantly longer than ChatGPT.
  5. January 15, 2000? (9): MIDSUMMER
    Definition: Again, the entire clue is the definition.
    Wordplay: This is the sort of clue that makes people either loathe DA or adore him. If you count back through your calendars, with 2000 being a leap year, January 15 was literally the middle of summer in Australia. But 2000 in Roman numerals is also “MM”, which falls in the middle of the word. Look, it’s creative.
    HEWITT WON?© Screenshot Gizmodo
    ChatGPT score: 1/10. One point for trying, and also because I find ChatGPT’s frankly terrible answer—HEWITT WON is two words, for a start—intriguing. It’s so confident in this answer. And it’s so wrong!
    Claude score: Still N/A
    Oreate score: Gonna give it a 0/10 here, because it’sstillthinking about this clue, and I need to get some dinner.

- Зар сурталчилгаа -

Та юу гэж бодож байна?

Сэтгэгдлээ оруулна уу!
Please enter your name here

MFC.mn сайтад сэтгэгдэл оруулахад анхаарах зүйлс

Холбоотой

spot_img

Шинэ

spot_img