Opus 5 Does Fable 5's Job for Half the Price

Opinion

Fable 5 came back on the 1st of July with a meter attached. Max plans get it at 50% of weekly limits from the 20th. Pro moved to usage credits with a one-off hundred dollars to soften the landing. The best model any of us had used was available again, and you had to decide whether a task was worth spending it on.

Opus 5 landed on Friday at $5 per million input tokens and $25 per million output. Fable 5 is $10 and $50.

Introducing Claude Opus 5

Same price as Opus 4.8, two months after Opus 4.8. I have no idea how the unit economics on that work.

Anthropic’s numbers put Opus 5 ahead of Fable 5 on agentic terminal coding, knowledge work, agentic search, computer use, business workflows and novel problem solving. The places it loses, it loses by a tenth of a percent.

Opus 5 compared with Fable 5, Opus 4.8 and GPT-5.6 Sol across benchmarks

Frontier-Bench v0.1 has Opus 5 at 43.3%, Fable 5 at 33.7% and Opus 4.8 at 21.1%. Doubling your own previous generation in two months while beating a model that costs twice as much is not a result I expected to see this year.

The cost curve is more useful than the score.

Frontier-Bench v0.1 score plotted against cost per attempt, with Opus 5’s curve above and to the left of Fable 5’s

Fable has to climb to roughly $28 an attempt to reach its top score. Opus 5 passes that mark at about $8.50 and keeps going. Its whole line sits above and to the left of Fable’s, so there’s no budget on this benchmark where spending Fable money buys a better answer. Six weeks ago Fable was untouchable.

CursorBench is tighter. At max effort Opus 5 comes within 0.5% of Fable’s peak at half the cost per task, and half a percent is run-to-run noise.

CursorBench agentic coding by effort level, showing Opus 5 reaching Fable 5’s peak score at roughly half the cost

Design is what I actually care about and the thing benchmarks measure worst. Anthropic says this is the best animations, games and 3D work they’ve had out of an Opus model. The two demos they shipped alongside the announcement, a wind tunnel and an interactive cell, are the sort of thing I’d have handed to Fable a month ago.

What I want to see hold up is the self-checking. On their frontend benchmark the model opened its own pages in a browser at desktop and phone widths, found a product hidden below the mobile fold and a checkout button sitting off screen, and fixed both before handing the work back. Every model I’ve used will happily tell you a layout is finished while it’s falling apart at 390 pixels wide. A model that goes and looks is worth more to me than four points on a coding eval.

Computer use, same shape. 70.6% on OSWorld 2.0 against Fable’s 66.1%, and Opus 5 beats Fable’s best result at about a third of the cost.

OSWorld 2.0 agentic computer use by effort level

Where it loses: FrontierCode by a tenth of a percent, which is a tie. The legal agent benchmark, 13.3% to 11.7%. Professional health questions, where Mythos 5 is well ahead at 66.0% against 59.8%. And DeepSWE, which GPT-5.6 Sol takes at 72.7% with Opus 5 on 68.8% and Fable on 69.7%. The coding crown isn’t the clean sweep the coverage makes it sound like.

The effort dial does more work here than on any model before it. Low, medium, high, xhigh, max, and on CursorBench that ladder runs from about $2.50 a task to about $8. The bottom of it is the surprise. Opus 5 at low effort scores 62.8%, which is better than Opus 4.8 managed at max, for less than half the money. Fast mode is there too, 2.5 times the output speed at twice the base price, which puts you back at Fable rates for speed rather than intelligence.

Two things to check before you swap the model string. Thinking is on by default now, so any route that never set the thinking parameter and sized its token limit tightly around the answer will start truncating. And you can only disable thinking at high effort or below, so thinking off plus xhigh returns a 400 instead of doing something sensible.

Opus 5 is the default on Max and the strongest thing on Pro, so most people reading this already have it without touching a setting.

The open question is what Fable is for now. On Anthropic’s own charts it’s matched or beaten on nearly everything I’d use it for, at twice the price, and what it still holds is legal work, health questions and a tenth of a percent on FrontierCode. That’s a thin strip of ground for a flagship.

Sources: