Arena raises $200M Series B and launches an AI Alignment Index

Arena raises $200M Series B and launches an AI Alignment Index

The AI evaluation startup now carries a $2.88 billion valuation and wants to grade how AI agents behave, not just how smart they sound

Arena, the startup behind one of the most watched AI leaderboards, has closed a $200 million Series B at a post-money valuation of $2.88 billion. The round closed on September 22, 2026.

Weeks later, the company followed up with a new product: an Alignment Index that scores how AI agents behave when given real tasks.

The money and the math

The Series B lifts Arena’s total funding to approximately $450 million.

Arena raised $100 million in seed funding in May 2025, then a $150 million Series A in January 2026. The Series B arrived about eight months after that.

Investors who have backed Arena across its rounds include Andreessen Horowitz, Felicis, UC Investments, Kleiner Perkins, and Lightspeed, among others.

Arena’s commercial AI Evaluations product, introduced in September 2025, reached an estimated annualized run-rate of $100 million by June 2026. That figure stood at $30 million in January 2026.

Advertisement

Arena charges on a consumption basis rather than through fixed subscriptions.

Grading the agents

The Alignment Index launched on October 8, 2026. It is designed as a benchmark for measuring AI agent safety and alignment in practical settings.

Arena built the index from more than 90,000 agent sessions spanning 27 different models.

The benchmark tracks three specific safety signals:

Unauthorized Action (UA): the agent does something it was not permitted to do.

False Attribution (FA): the agent credits information or actions to the wrong source.

Deceptive Completion (DC): the agent claims a task is finished when it is not.

Early results put OpenAI in front. GPT-6.1-Sol currently leads the index with a score of 87.9.

Anthropic’s Claude-Opus-5.5 follows at 83.2, with Grok-4.7 close behind at 82.7.

From Berkeley project to billion-dollar referee

Arena began in April 2023 as an open-source project at UC Berkeley. The original idea was simple: let users compare responses from two anonymous AI models and vote on the better one.

The company formally incorporated in April 2025. Since then it has expanded well beyond text, adding evaluations for coding, vision, and agent workflows.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.
Arena raises $200M Series B and launches an AI Alignment Index
Arena raises $200M Series B and launches an AI Alignment Index

The AI evaluation startup now carries a $2.88 billion valuation and wants to grade how AI agents behave, not just how smart they sound

Arena, the startup behind one of the most watched AI leaderboards, has closed a $200 million Series B at a post-money valuation of $2.88 billion. The round closed on September 22, 2026.

Weeks later, the company followed up with a new product: an Alignment Index that scores how AI agents behave when given real tasks.

The money and the math

The Series B lifts Arena’s total funding to approximately $450 million.

Arena raised $100 million in seed funding in May 2025, then a $150 million Series A in January 2026. The Series B arrived about eight months after that.

Investors who have backed Arena across its rounds include Andreessen Horowitz, Felicis, UC Investments, Kleiner Perkins, and Lightspeed, among others.

Arena’s commercial AI Evaluations product, introduced in September 2025, reached an estimated annualized run-rate of $100 million by June 2026. That figure stood at $30 million in January 2026.

Advertisement

Arena charges on a consumption basis rather than through fixed subscriptions.

Grading the agents

The Alignment Index launched on October 8, 2026. It is designed as a benchmark for measuring AI agent safety and alignment in practical settings.

Arena built the index from more than 90,000 agent sessions spanning 27 different models.

The benchmark tracks three specific safety signals:

Unauthorized Action (UA): the agent does something it was not permitted to do.

False Attribution (FA): the agent credits information or actions to the wrong source.

Deceptive Completion (DC): the agent claims a task is finished when it is not.

Early results put OpenAI in front. GPT-6.1-Sol currently leads the index with a score of 87.9.

Anthropic’s Claude-Opus-5.5 follows at 83.2, with Grok-4.7 close behind at 82.7.

From Berkeley project to billion-dollar referee

Arena began in April 2023 as an open-source project at UC Berkeley. The original idea was simple: let users compare responses from two anonymous AI models and vote on the better one.

The company formally incorporated in April 2025. Since then it has expanded well beyond text, adding evaluations for coding, vision, and agent workflows.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.