GPT-6 Astra leads Epoch Capabilities Index with ECI of 166

Photo: Tima Miroshnichenko / Pexels

GPT-6 Astra leads Epoch Capabilities Index with ECI of 166

OpenAI's newest model tops 249 competitors on Epoch AI's composite benchmark, though a glaring weakness in software engineering keeps things interesting

OpenAI’s GPT-6 Astra has claimed the top spot on the Epoch Capabilities Index, posting a composite score of 166 and outranking all 249 models currently tracked by the research organization. The model, which began rolling out on September 3, 2026, represents the highest single-score performance Epoch AI has ever recorded, surpassing both GPT-5.5 Pro (162) and Anthropic’s Claude Fable 5.1 (164).

But there’s a catch buried in the numbers. Astra’s software engineering score sits at just 47% on the MirrorCode benchmark, a category where Claude Fable 5.1 posts 73%. For a model billing itself as the most capable AI system on the planet, that gap is hard to ignore.

What the numbers actually show

The Epoch Capabilities Index isn’t a single test. It aggregates performance across 37 to 59 benchmarks into one composite score, normalizing results so that comparisons across model generations are meaningful. A difference of roughly 10 points on the ECI scale is considered statistically significant, which means Astra’s two-point lead over Claude Fable 5.1 is notable but not exactly a blowout.

Advertisement

Where Astra truly flexes is in math, continual learning, and game-puzzle categories. The model achieved 98% on the FrontierMath Tier 4 benchmark and 96% on GPQA Diamond.

There are also reports suggesting Astra could score as high as 169 on the ECI depending on which benchmarks are included in the composite calculation. Epoch AI periodically updates its benchmark suite, and different inclusion criteria can shift composite scores by a few points in either direction.

How Astra reached users

OpenAI followed what has become its standard playbook for major releases. GPT-6 Astra was initially made available to select organizations before opening up to the broader user base across ChatGPT Plus, Pro, Business, and Enterprise tiers.

Pricing sits at $10 per million input tokens and $50 per million output tokens, based on Epoch AI’s recommended pricing page.

The competitive landscape gets tighter

Claude Fable 5.1’s score of 164 puts Anthropic within striking distance. A two-point gap on a normalized composite index is close enough that the next model release from either company could flip the rankings. Claude’s 73% on MirrorCode versus Astra’s 47% gives Anthropic a clear selling point for the developer market.

GPT-5.5 Pro’s score of 162 now looks like a transitional checkpoint. The four-point improvement from 5.5 Pro to Astra suggests OpenAI is still finding meaningful gains between releases.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.
GPT-6 Astra leads Epoch Capabilities Index with ECI of 166
GPT-6 Astra leads Epoch Capabilities Index with ECI of 166

OpenAI's newest model tops 249 competitors on Epoch AI's composite benchmark, though a glaring weakness in software engineering keeps things interesting

Photo: Tima Miroshnichenko / Pexels

OpenAI’s GPT-6 Astra has claimed the top spot on the Epoch Capabilities Index, posting a composite score of 166 and outranking all 249 models currently tracked by the research organization. The model, which began rolling out on September 3, 2026, represents the highest single-score performance Epoch AI has ever recorded, surpassing both GPT-5.5 Pro (162) and Anthropic’s Claude Fable 5.1 (164).

But there’s a catch buried in the numbers. Astra’s software engineering score sits at just 47% on the MirrorCode benchmark, a category where Claude Fable 5.1 posts 73%. For a model billing itself as the most capable AI system on the planet, that gap is hard to ignore.

What the numbers actually show

The Epoch Capabilities Index isn’t a single test. It aggregates performance across 37 to 59 benchmarks into one composite score, normalizing results so that comparisons across model generations are meaningful. A difference of roughly 10 points on the ECI scale is considered statistically significant, which means Astra’s two-point lead over Claude Fable 5.1 is notable but not exactly a blowout.

Advertisement

Where Astra truly flexes is in math, continual learning, and game-puzzle categories. The model achieved 98% on the FrontierMath Tier 4 benchmark and 96% on GPQA Diamond.

There are also reports suggesting Astra could score as high as 169 on the ECI depending on which benchmarks are included in the composite calculation. Epoch AI periodically updates its benchmark suite, and different inclusion criteria can shift composite scores by a few points in either direction.

How Astra reached users

OpenAI followed what has become its standard playbook for major releases. GPT-6 Astra was initially made available to select organizations before opening up to the broader user base across ChatGPT Plus, Pro, Business, and Enterprise tiers.

Pricing sits at $10 per million input tokens and $50 per million output tokens, based on Epoch AI’s recommended pricing page.

The competitive landscape gets tighter

Claude Fable 5.1’s score of 164 puts Anthropic within striking distance. A two-point gap on a normalized composite index is close enough that the next model release from either company could flip the rankings. Claude’s 73% on MirrorCode versus Astra’s 47% gives Anthropic a clear selling point for the developer market.

GPT-5.5 Pro’s score of 162 now looks like a transitional checkpoint. The four-point improvement from 5.5 Pro to Astra suggests OpenAI is still finding meaningful gains between releases.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.