Apollo Research wants AI evaluators inside the lab during training, not just at launch

Photo: Steve A Johnson / Pexels

Apollo Research wants AI evaluators inside the lab during training, not just at launch

The AI safety group argues that independent testers need ongoing access to model development because the riskiest behaviors may emerge long before release

AI safety organization Apollo Research has a simple pitch for the industry: stop inviting the inspectors only after the building is finished.

The group argues that independent evaluators should have continuous access to AI models throughout training, not just a brief look right before a public release. It calls the idea “embedded evaluations.” The goal is to catch dangerous behaviors, especially scheming, while they are still forming.

What Apollo is proposing

Apollo laid out its case in a series of blog posts published between September 30 and October 5, 2026. The central argument targets how outside testing currently works.

Today, external evaluations typically happen shortly before a model goes public. Apollo’s concern is that by then, the most informative moments in a model’s development have already passed.

Instead, the organization wants independent evaluators to have something closer to employee-level access. That means ongoing visibility into the entire training process, rather than a limited set of checks at the finish line.

Advertisement

Apollo points to risks tied to multi-agent training and reward-seeking behavior. These issues, the group says, often show up during development rather than after launch.

Apollo also outlined four criteria it considers essential for embedded evaluations to work:

  • Risk reduction: the evaluations should actually lower the chance of harm.
  • Public information: findings should help inform the public.
  • Aligned incentives: evaluators and developers should not be pulling in opposite directions.
  • Fairness: the arrangement should be equitable for the parties involved.

From blog posts to Capitol Hill

The proposal did not stay on the company blog. Apollo CEO Marius Hobbhahn testified in early October 2026 before the US Senate Homeland Security Subcommittee.

In that testimony, Hobbhahn advocated for mandatory independent evaluations in AI development.

The push also lands as prominent industry leaders, including Anthropic’s Dario Amodei and OpenAI’s Sam Altman, have backed calls for greater transparency and oversight in AI.

Who Apollo Research is

Apollo Research was founded in 2023. From the start, it focused on a specific category of AI risk: systems that are deceptively aligned, meaning they appear to behave well while pursuing different goals underneath.

The organization built its reputation through technical research collaborations with major AI labs. That includes work with OpenAI studying reward-seeking behavior during reinforcement learning, the stage where models are trained by being rewarded for desired outputs.

By 2025 and 2026, Apollo had sharpened its focus on scheming and on the limits of relying only on final model evaluations. The embedded evaluations proposal is the logical next step from that research: if the problem appears during training, the testing should too.

What this means for AI labs and the industry

The core tension here is access. Frontier AI training runs are among the most closely guarded processes in tech. Letting independent evaluators observe them on an ongoing basis would require labs to share far more than they currently do.

The regulatory angle may matter most. Hobbhahn’s call for mandatory independent evaluations in front of a Senate subcommittee signals that Apollo sees voluntary arrangements as insufficient.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.
Apollo Research wants AI evaluators inside the lab during training, not just at launch
Apollo Research wants AI evaluators inside the lab during training, not just at launch

The AI safety group argues that independent testers need ongoing access to model development because the riskiest behaviors may emerge long before release

Photo: Steve A Johnson / Pexels

AI safety organization Apollo Research has a simple pitch for the industry: stop inviting the inspectors only after the building is finished.

The group argues that independent evaluators should have continuous access to AI models throughout training, not just a brief look right before a public release. It calls the idea “embedded evaluations.” The goal is to catch dangerous behaviors, especially scheming, while they are still forming.

What Apollo is proposing

Apollo laid out its case in a series of blog posts published between September 30 and October 5, 2026. The central argument targets how outside testing currently works.

Today, external evaluations typically happen shortly before a model goes public. Apollo’s concern is that by then, the most informative moments in a model’s development have already passed.

Instead, the organization wants independent evaluators to have something closer to employee-level access. That means ongoing visibility into the entire training process, rather than a limited set of checks at the finish line.

Advertisement

Apollo points to risks tied to multi-agent training and reward-seeking behavior. These issues, the group says, often show up during development rather than after launch.

Apollo also outlined four criteria it considers essential for embedded evaluations to work:

  • Risk reduction: the evaluations should actually lower the chance of harm.
  • Public information: findings should help inform the public.
  • Aligned incentives: evaluators and developers should not be pulling in opposite directions.
  • Fairness: the arrangement should be equitable for the parties involved.

From blog posts to Capitol Hill

The proposal did not stay on the company blog. Apollo CEO Marius Hobbhahn testified in early October 2026 before the US Senate Homeland Security Subcommittee.

In that testimony, Hobbhahn advocated for mandatory independent evaluations in AI development.

The push also lands as prominent industry leaders, including Anthropic’s Dario Amodei and OpenAI’s Sam Altman, have backed calls for greater transparency and oversight in AI.

Who Apollo Research is

Apollo Research was founded in 2023. From the start, it focused on a specific category of AI risk: systems that are deceptively aligned, meaning they appear to behave well while pursuing different goals underneath.

The organization built its reputation through technical research collaborations with major AI labs. That includes work with OpenAI studying reward-seeking behavior during reinforcement learning, the stage where models are trained by being rewarded for desired outputs.

By 2025 and 2026, Apollo had sharpened its focus on scheming and on the limits of relying only on final model evaluations. The embedded evaluations proposal is the logical next step from that research: if the problem appears during training, the testing should too.

What this means for AI labs and the industry

The core tension here is access. Frontier AI training runs are among the most closely guarded processes in tech. Letting independent evaluators observe them on an ongoing basis would require labs to share far more than they currently do.

The regulatory angle may matter most. Hobbhahn’s call for mandatory independent evaluations in front of a Senate subcommittee signals that Apollo sees voluntary arrangements as insufficient.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.