Opus 5 tops a brutal benchmark
built to measure "real intelligence"
A day earlier, Anthropic's new model Opus 5 was making headlines mainly for being cheap — "far cheaper than Fable 5." Today the story flipped: on ARC-AGI-3, a benchmark deliberately designed to resist memorization, Opus 5 blew past OpenAI's GPT-5.6 Sol by a wide margin, THE DECODER reports.
What happened on a benchmark
"you can't memorize your way through"
Most LLM benchmarks have a weakness: models that have seen similar problems during training tend to score higher, regardless of whether they can actually reason. ARC-AGI-3, built by the nonprofit ARC Prize Foundation, is deliberately designed to resist that — it measures fluid, general reasoning on environments a model hasn't encountered before, rather than rewarding memorized patterns. It's meant to answer a narrower and harder question than a typical leaderboard rank: is the model actually thinking, or just recalling?
According to ARC Prize Foundation's measurements, Anthropic's Opus 5 scored 30.2% on ARC-AGI-3 — a new record that decisively beats the previous record of 7.8% set by OpenAI's GPT-5.6 Sol (Max), as reported by THE DECODER. This isn't a narrow edge — it's roughly a fourfold jump over the prior record, which is unusual for a single model generation.
GPT-5.6 Sol isn't the only comparison point. Per ARC Prize's own analysis, Anthropic's own previous "Fable-class" mid-tier models scored around 20% on the same benchmark. So Opus 5's 30.2% represents a big step forward not just over a competitor, but over Anthropic's own prior generation as well.
A model people were choosing on price
just gained a reason to be chosen for performance, too.
What's inside the 30.2%
ARC Prize Foundation's analysis found that Opus 5 newly solved five environments nobody had solved before, and four of those five reached or exceeded the human baseline.
Opus 5 newly solved five environments that had gone unsolved before, and four of them scored at or above human level, according to ARC Prize Foundation's analysis. The foundation attributes this to stronger logical reasoning, which it says enables more autonomous exploration, planning, and execution in unfamiliar environments.
During testing, researchers also observed Opus 5 translating tasks into algebraic notation and independently formulating reflection equations — behavior they said they hadn't seen from a model before. Coverage of the launch, including this kind of detail, has also been picked up by the independent benchmark tracker ArtificialAnalysis.ai's Opus 5 launch article.
Who this matters to, and how
Some readers should act on this. Others can safely ignore it.
Reconsider how you evaluate unfamiliar tasks
If you're building agentic workflows for domains a model hasn't seen before, a reasoning-specific benchmark like ARC-AGI-3 deserves more weight in your model selection than coding-leaderboard rank alone can tell you.
Don't decide on price alone anymore
If you ruled Opus 5 in or out purely on the "cheaper than Fable 5" pricing story, you should re-run that comparison now that it's also a top capability contender, not just a value option.
If you're already all-in on Claude, this changes little
For teams already running production workloads on Claude, this is one more reason to stay, not a signal to change anything. No urgent action needed.
What to do next
Re-run your own eval suite
If you previously ruled out Opus on capability grounds, this update is a good trigger to re-run your evaluation suite against your actual use cases.
Watch for a Fable 5 / GPT successor response
ARC-AGI-3 is a public benchmark, and competitors may post a counter-score within weeks. Don't treat today's ranking as settled.
Don't switch on one benchmark alone
Before switching models based on ARC-AGI-3 alone, check the caveats below and confirm against a benchmark closer to your own task type.
Counterpoints, risks, and limits
It would be premature to read this as "Opus 5 wins everything." In the same round of testing, ARC Prize Foundation also reported results on the Witness benchmark, which measures reasoning in interactive puzzle games, and Opus 5 scored 43.4 there — a statistical tie with Kimi K3 and Fable 5, with far less improvement over the previous Opus 4.8 than the ARC-AGI-3 jump showed.
In other words, Opus 5's reasoning gains are real but uneven — they don't show up the same way across every benchmark type. ARC-AGI-3 is one specific test designed to resist memorization; it is not a blanket guarantee of superiority on every task. Price is no longer the only reason to pick Opus 5, but teams should still validate against benchmarks that resemble their own workloads before switching.