Claude Sonnet 5.5 takes third place on Arena’s Agent Arena leaderboard
Anthropic's mid-tier model posted a 12.52% net improvement score and now trails only two of its own pricier siblings
Anthropic’s newest mid-tier model has landed in third place on Arena.ai’s Agent Arena leaderboard. Claude Sonnet 5.5 posted a net improvement score of 12.5%, which puts it ahead of OpenAI’s top entry.
The model sits behind just two competitors, and both are also made by Anthropic.
How the leaderboard shakes out
Claude Sonnet 5.5 launched on September 28, 2026. Within days, it had climbed to the number three spot on Agent Arena.
Its precise net improvement score is 12.52%, with a margin of plus or minus 3.09%.
The two models above it are Claude Fable 5.1 (Max) at 14.31% and Claude Opus 5.5 (High) at 13.82%. Directly below sits OpenAI’s GPT-6 Astra (Max) at 12.27%.
Agent Arena grades models on millions of real-world agentic tasks, where an AI has to use tools, complete multi-step jobs and satisfy actual users. The scoring looks at tool reliability, task completion and user feedback.
Cost is the other headline number. Arena lists Sonnet 5.5 at $2.74 per task, while early October snapshots put its median cost at around $2.78 per task. Sonnet 5.5 also uses more tokens than most of its peers.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
The benchmarks beyond Arena
On the Artificial Analysis Intelligence Index, it scored 56, good for second place behind Opus 5.5 (Max).
On Terminal-Bench 4.0, Sonnet 5.5 scored 70.6%. For comparison, its predecessor, Sonnet 5, managed roughly 10 to 14% on the same test.
It also edged out Opus 5.5, which scored 66.4% on Terminal-Bench 4.0.
Speed improved too. Sonnet 5.5 runs more than 30% faster than Sonnet 5.
The pricing did not move. It still costs $2 per million input tokens and $10 per million output tokens, the same as before.
Sonnet 5.5 offers a context window of 1 million tokens and a maximum output of 128,000 tokens.
What this means for the AI model race
The first thing to watch is that error bar. Sonnet 5.5’s score carries a margin of plus or minus 3.09%, and the gap between the top four models is smaller than that. GPT-6 Astra (Max) trails Sonnet 5.5 by just 0.25 percentage points.
Higher token usage can eat into the cost advantage, depending on the workload. A model that is cheap per token but verbose may not always be cheap per job.