I woke to the day’s noise—Firefox was adding an AI crossword to its new-tab feed—and felt a small, private fury. I remembered DA’s Friday puzzle, the one that makes grown-up cruciverbalists mutter, and I decided to test the machines. If you love cryptics, you’ll understand how much is at stake.
I’m going to be blunt with you: I ran five of DA’s nastiest clues through three LLMs—ChatGPT, Claude Sonnet 5, and Oreate—and I’m going to tell you what happened. I’ll show you the clues, the answers, and why these models still trip over the strangest little human tricks.
Putting LLMs to the test
Friday mornings in Sydney, the paper hits stands and people groan at DA’s byline.
I picked DA because his clues are mean in a way that rewards cultural memory and lateral nastiness. I wanted to see if large language models could reverse-engineer DA’s perverse logic, or if they’d fold like a poker player with a bad hand. I fed the clues one at a time, nudging, prodding, and keeping the same prompt structure so the comparison would be fair.
LLMs are good at anagrams and straightforward definitions, but cryptics are often a puzzle wrapped in a lie wrapped in a grammar rule. Some machines treated these clues like a codebook; others treated them like a menu. One behaved like a locksmith picking a safe—precise but missing the final tumblers.
The clues I fed them
I pulled five from DA’s Friday grid: a mix that starts comfy and finishes with a punch.
- Wolfed peanut brittle, finally gluten-free! (3,2)
- A dyer backing reduced labour now and then (2,9)
- Pinkie, capisce? (5)
- Struggle for anyone (not!) getting time to focus essentially?! (9,7)
- January 15, 2000? (9)
If you want to stop reading and solve them, go ahead—this is your moment. I gave the same clues to each model with identical context and no external searching allowed, except where the model revealed it had looked something up.
How the LLMs performed
I ran the clues through ChatGPT, Claude Sonnet 5, and Oreate on the same hardware session; timing and patience were part of the experiment.
ChatGPT — Score: 27/50. Started well, nailed the short clue, then confidently invented nonsense for the finale.
Claude Sonnet 5 — Score: 7/50. Lethargic, prone to circular reasoning, and occasionally demanded payment for thinking.
Oreate — Score: 21/50. Slow, sometimes correct, and once outright confessed it had found an answer online rather than solving it.
Can AI solve cryptic crosswords?
Short answer: sometimes for simple device tricks, rarely for the truly perverse. Anagram and definition combos are low-hanging fruit; cryptics that depend on cultural squeeze or &lit. misdirection still break most LLMs.
Which LLM fared best?
ChatGPT was the most consistent across the set; Oreate surprised on one clue but admitted to “googling” on a trick one, which disqualifies that success if you care about purist solving. Claude simply cratered.

Per-clue results
Each clue is a small trapdoor; I’ll walk you through the answers, the wordplay, and how each model coped.
-
Wolfed peanut brittle, finally gluten-free! (3,2): ATE UP
Definition: “Wolfed,” i.e. ate quickly.
Wordplay: Anagram of PEANUT minus the final letter of “gluten” (n) → PEAUT → ATE UP.
ChatGPT: Correct but needed the explanation — 7/10.
Claude: Almost (EAT UP) and uncertain — 5/10.
Oreate: Slow but correct, understood the subtraction — 7/10.
-
A dyer backing reduced labour now and then (2,9): AT INTERVALS
Definition: “Now and then.”
Wordplay: “A dyer” = A TINTER; “labour” = SLAVE → backing and reduced → EVALS; stitch them to get AT INTERVALS.
ChatGPT: Fumbled the lengths, argued with the prompt, but eventually landed — 5/10.
Claude: Slow and overeager, correct only after prompting — 2/10.
Oreate: Surprisingly identified “a tinter” and found the build — 6/10.
-
Pinkie, capisce? (5): DIGIT
Definition/Wordplay: Short, double-clue: pinkie = small finger; capisce? → “dig it” → DIGIT.
ChatGPT: Fast and flawless — 10/10.
Oreate: Correct but slow — 8/10.
-
Struggle for anyone (not!) getting time to focus essentially?! (9,7): ATTENTION ECONOMY
Observation: In group chats this clue made three people swear and one person laugh out loud.
Definition: The whole clue is an &lit. — it defines the answer.
Wordplay: Anagram indicator “Struggle [for]”; fodder = ANYONE NOT TIME TO + middle letter of FOCUS (C) → ATTENTION ECONOMY.
ChatGPT: Couldn’t conjure it on its own, required a breakdown — 4/10.
Claude: Timed out in thought — N/A.
Oreate: Returned the answer but admitted it was sourced online, not solved — N/A for purists.
-
January 15, 2000? (9): MIDSUMMER
Observation: People either love DA’s “Australian-calendar” winks or they accuse him of cruelty.
Definition/Wordplay: In Australia, January 15 sits deep in summer; 2000 (MM) sits in the middle of the phrase—a playful calendar/typographic nudge → MIDSUMMER.
ChatGPT: Confidently wrong (“HEWITT WON”) — 1/10.
Oreate: Still thinking at dinner — 0/10.

I solved four of the five myself and scored the session a smug 40/50. Machines can do pieces of the job—anagrams, short double-definitions—but when you hand them a culture-bound trick or a delicious &lit., they stall, invent, or worse, pretend.
Claude wanted to charge me for patience; Oreate sometimes cheated; ChatGPT oscillated between clarity and hubris. I’ll admit: the models progress fast, but cryptics are a human art of misdirection and tiny, private cruelty. They require a reader who knows jokes, histories, and the weight of punctuation.
If you love puzzles, you’ll enjoy watching the machines trip. If you’re building one, beware the neat trap: passing a few tests doesn’t mean you’ve learned to lie like a setter. Do you think artificial solvers will ever truly appreciate the petty, personal humor baked into a clue?