Claude Sonnet 5.5 takes third place on Arena’s Agent Arena leaderboard

Claude Sonnet 5.5 takes third place on Arena’s Agent Arena leaderboard

Anthropic's mid-tier model posted a 12.52% net improvement score and now trails only two of its own pricier siblings

Anthropic’s newest mid-tier model has landed in third place on Arena.ai’s Agent Arena leaderboard. Claude Sonnet 5.5 posted a net improvement score of 12.5%, which puts it ahead of OpenAI’s top entry.

The model sits behind just two competitors, and both are also made by Anthropic.

How the leaderboard shakes out

Claude Sonnet 5.5 launched on September 28, 2026. Within days, it had climbed to the number three spot on Agent Arena.

Its precise net improvement score is 12.52%, with a margin of plus or minus 3.09%.

The two models above it are Claude Fable 5.1 (Max) at 14.31% and Claude Opus 5.5 (High) at 13.82%. Directly below sits OpenAI’s GPT-6 Astra (Max) at 12.27%.

Advertisement

Agent Arena grades models on millions of real-world agentic tasks, where an AI has to use tools, complete multi-step jobs and satisfy actual users. The scoring looks at tool reliability, task completion and user feedback.

Cost is the other headline number. Arena lists Sonnet 5.5 at $2.74 per task, while early October snapshots put its median cost at around $2.78 per task. Sonnet 5.5 also uses more tokens than most of its peers.

The benchmarks beyond Arena

On the Artificial Analysis Intelligence Index, it scored 56, good for second place behind Opus 5.5 (Max).

On Terminal-Bench 4.0, Sonnet 5.5 scored 70.6%. For comparison, its predecessor, Sonnet 5, managed roughly 10 to 14% on the same test.

It also edged out Opus 5.5, which scored 66.4% on Terminal-Bench 4.0.

Speed improved too. Sonnet 5.5 runs more than 30% faster than Sonnet 5.

The pricing did not move. It still costs $2 per million input tokens and $10 per million output tokens, the same as before.

Sonnet 5.5 offers a context window of 1 million tokens and a maximum output of 128,000 tokens.

What this means for the AI model race

The first thing to watch is that error bar. Sonnet 5.5’s score carries a margin of plus or minus 3.09%, and the gap between the top four models is smaller than that. GPT-6 Astra (Max) trails Sonnet 5.5 by just 0.25 percentage points.

Higher token usage can eat into the cost advantage, depending on the workload. A model that is cheap per token but verbose may not always be cheap per job.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.
Claude Sonnet 5.5 takes third place on Arena’s Agent Arena leaderboard
Claude Sonnet 5.5 takes third place on Arena’s Agent Arena leaderboard

Anthropic's mid-tier model posted a 12.52% net improvement score and now trails only two of its own pricier siblings

Anthropic’s newest mid-tier model has landed in third place on Arena.ai’s Agent Arena leaderboard. Claude Sonnet 5.5 posted a net improvement score of 12.5%, which puts it ahead of OpenAI’s top entry.

The model sits behind just two competitors, and both are also made by Anthropic.

How the leaderboard shakes out

Claude Sonnet 5.5 launched on September 28, 2026. Within days, it had climbed to the number three spot on Agent Arena.

Its precise net improvement score is 12.52%, with a margin of plus or minus 3.09%.

The two models above it are Claude Fable 5.1 (Max) at 14.31% and Claude Opus 5.5 (High) at 13.82%. Directly below sits OpenAI’s GPT-6 Astra (Max) at 12.27%.

Advertisement

Agent Arena grades models on millions of real-world agentic tasks, where an AI has to use tools, complete multi-step jobs and satisfy actual users. The scoring looks at tool reliability, task completion and user feedback.

Cost is the other headline number. Arena lists Sonnet 5.5 at $2.74 per task, while early October snapshots put its median cost at around $2.78 per task. Sonnet 5.5 also uses more tokens than most of its peers.

The benchmarks beyond Arena

On the Artificial Analysis Intelligence Index, it scored 56, good for second place behind Opus 5.5 (Max).

On Terminal-Bench 4.0, Sonnet 5.5 scored 70.6%. For comparison, its predecessor, Sonnet 5, managed roughly 10 to 14% on the same test.

It also edged out Opus 5.5, which scored 66.4% on Terminal-Bench 4.0.

Speed improved too. Sonnet 5.5 runs more than 30% faster than Sonnet 5.

The pricing did not move. It still costs $2 per million input tokens and $10 per million output tokens, the same as before.

Sonnet 5.5 offers a context window of 1 million tokens and a maximum output of 128,000 tokens.

What this means for the AI model race

The first thing to watch is that error bar. Sonnet 5.5’s score carries a margin of plus or minus 3.09%, and the gap between the top four models is smaller than that. GPT-6 Astra (Max) trails Sonnet 5.5 by just 0.25 percentage points.

Higher token usage can eat into the cost advantage, depending on the workload. A model that is cheap per token but verbose may not always be cheap per job.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.