Opus 5.5 Just Mogged GPT-6 Astra
I’ve been using Opus 5.5 in Claude Code a lot since Anthropic released it on Tuesday, and my weekly usage has barely gone down. When I was using GPT-6 Astra in Codex it burnt through my usage so fast, and that was on low reasoning.
Anthropic just mogged OpenAI. Opus 5.5 is cheaper and faster than Astra, it beats it on most of the coding benchmarks, and it isn’t eating my plan.
I wasn’t the only one getting hammered by Astra either. On 6 September a Plus user opened an issue on the Codex GitHub repo saying they started a fresh session at 2:51pm and hit their limit at 3:11pm. Twenty minutes. Someone on the Pro 20x plan posted the same day that they were already down 20% of their weekly usage that morning, and people on the $20 plan were joking about getting 11 minutes of Astra.
Tibo Sottiaux, who runs Codex, spent that week handing out resets to keep everyone going. There was a banked reset for every day people waited for access, a full one on the 5th when Astra rolled out, and another for everyone on the 7th. Nice gesture, but I don’t want a plan that only works when someone at OpenAI remembers to hit the reset button.
Weirdly, Astra doesn’t use many tokens. Artificial Analysis measured it using the fewest of any agent in their Coding Agent Index. Going by the complaints, some of it is Codex resending the whole conversation every turn, and some of it is Astra building stuff you never asked for. People have posted about it answering a small feature request with five levels of verification, smoke tests and SHA256 hashes. That’s your week’s usage going on tests for a button.
Anthropic went the other way. They say limits on Pro, Max and Team go about 25% further on Opus 5.5 than on Opus 5 because the tokens are cheaper, and they’ve bumped the five-hour limits by 20% on top of that. Everyone got a spare reset to use whenever they want as well.
Opus 5.5 also says a lot less than Opus 5 did. Someone on dev.to ran 468 graded calls through both at their default settings. Opus 5.5 got the same answers with a median of 193 output tokens, where Opus 5 used 448. In tool loops it was 427 against 666.
Right, the API. Opus 5.5 is $4 per million input tokens and $20 per million output. Astra is $10 and $50, so Opus is 40% of the price. Cached input is 20 cents against Astra’s $1, which is a big deal for coding agents because most of what they do is re-read the same context over and over. Astra also charges $20 and $75 for the whole request once you go over 272K input tokens. Opus doesn’t have anything like that.
I’d forgive it for being a bit worse at that price, but it isn’t. Anthropic’s table has it ahead of Astra on Terminal-Bench 4.0 by 8.5 points (66.4% to 57.9%), on Humanity’s Last Exam with tools by 10.5 and on GDPval-AA by 304 Elo. Astra wins AutomationBench by 1.4 points, which is nothing, and Terminal-Bench-Science by about six, which is a proper win. FrontierCode is 54.4% to 53.3%, so call that one even.
Anthropic put per-task costs on their charts too. On Terminal-Bench 4.0 at medium effort, Opus gets 57.6% for $2.94 a task and Astra gets 53.9% for $6.15. On high it’s 64.2% for $3.88 against 57.9% for $7.21. On GDPval, Opus on medium beats Astra on max for about a fifth of the price.
Those are Anthropic’s numbers, so I checked Artificial Analysis as well. Opus 5.5 on max scores 58 on their Intelligence Index and Astra on max scores 53. Opus on high scores 54 for $1.82 a task, so it beats Astra’s best result for a bit over half the money ($3.26). On medium they’re basically level at 51 and 50, and Opus is 20 cents a task cheaper.
Speed isn’t close. Opus does 78 to 92 tokens a second depending on effort and Astra does 45 to 52. Astra starts answering sooner on medium (about 6 seconds against 22), so a quick question feels a bit snappier in Codex, but Opus has caught up by about 1,700 tokens of output and any coding session goes well past that.
Max is where Opus gets expensive. It writes a lot more than Astra when you let it think. Artificial Analysis needed 38 million output tokens to run their index on Opus at medium and 19 million on Astra at medium. On max it was 260 million against 60 million.
You can see it in the Coding Agent Index. Claude Code with Opus 5.5 on max scores 66, the best they’ve recorded, and Codex with Astra on max scores 62. The Opus run costs $13.04 a task, though, and the Astra one costs $7.09, because Opus on max writes around 333K output tokens per task. One of Every’s testers blew through their allowance with it and lost their last day of testing.
So don’t leave it on max. It ships on medium in Claude Code, and on medium it’s cheaper than Astra and scores higher. Put it on high for the bigger jobs you leave running, where it’s still cheaper than Astra on high. I’d only go to max when you’ve seen high fail at something.
OpenAI released GPT-6 Sol minutes after Anthropic’s announcement, at $2 and $10, which is half the price of Opus 5.5. GPT-6 Luna came with it at 10 cents and 50 cents. Sol’s a good deal and it beats Opus 5 and Fable 5.1 on OpenAI’s benchmarks. It also sits below Astra, and Astra’s price didn’t change. Anthropic put out a better top model and OpenAI responded by making its cheaper one cheaper.
Astra launched at the same price as Fable 5.1, which I said at the time told you who OpenAI were going after. Now there’s an Opus that beats Fable on everything in Anthropic’s launch table for 40% of the price, and Astra is still charging Fable prices. That’s embarrassing.
(It’s also Anthropic’s first release since Dario’s essay about pacing the frontier. So much for pacing.)
Every did a vibe check with a few people who’d moved to Codex this year, and Opus 5.5 is dragging them back to Claude. One of them said they can’t afford another $200 subscription, so something has to go. Dan Shipper still splits work 80/20 in Codex’s favour but is spending far more tokens on Claude now.
Last week I had SWE-2 building the first pass and Fable or Astra cleaning up after it. Astra’s out. Opus 5.5 on high does the cleanup now and costs me less than either of them.
Comments