OpenAI’s Noam Brown reveals how AI agents solved a Millennium Prize Problem in 88 hours

OpenAI’s Noam Brown reveals how AI agents solved a Millennium Prize Problem in 88 hours

The researcher's latest comments offer a surprisingly candid look at where multi-agent coordination actually matters, and where it doesn't, on the road to AGI

A team of roughly 10,000 AI agents working together cracked one of mathematics’ most famous unsolved problems in under four days. And according to the OpenAI researcher who helped make it happen, the multi-agent coordination part barely mattered.

Noam Brown, speaking on the Dwarkesh Podcast on September 17, 2026, laid out a vision of AGI that’s equal parts impressive and humbling. The headline result, solving the Navier-Stokes Millennium Prize Problem using 130 billion tokens across 88 hours, sounds like a triumph for collaborative AI. But Brown estimated that multi-agent coordination contributed less than 10% of the breakthrough.

The automated research intern, still needs a supervisor

OpenAI’s current crop of AI agents function as what the company calls “automated research interns.” These systems can deliver outputs 3.1 times what human researchers produce on multi-day tasks. Over 50% of those tasks still require human oversight to stay on track.

Advertisement

Brown pinpointed the bottleneck with unusual specificity. The missing ingredient is what he called “research taste,” the intuitive ability to judge whether a research direction is genuinely significant and original. That quality, Brown noted, represents roughly 10% of his own work. But it’s the 10% that everything else depends on.

Scaling agents hits diminishing returns fast

Agent scaling is sublinear. Four agents working together approximately halve the latency compared to a single agent on suitable tasks. That sounds promising until you realize the gains diminish quickly as you add more agents to the mix.

What Brown found compelling wasn’t the raw scaling numbers. It was the emergent behavior: agents spontaneously developing human-like collaboration patterns without being explicitly programmed to do so. They communicate, delegate, and refine work using straightforward messaging tools in ways that mirror how research teams actually function. That emergent coordination, Brown suggested, offers real clues about what AGI systems might eventually look like.

The March 2028 target and recursive self-improvement

OpenAI has set an internal objective to achieve a fully autonomous AI researcher by March 2028. No human oversight required. No research taste borrowed from human supervisors.

The key mechanism for getting there is recursive self-improvement, or RSI. Brown flagged RSI as OpenAI’s top priority while also acknowledging the risks associated with cooperative multi-agent behavior.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.
OpenAI’s Noam Brown reveals how AI agents solved a Millennium Prize Problem in 88 hours
OpenAI’s Noam Brown reveals how AI agents solved a Millennium Prize Problem in 88 hours

The researcher's latest comments offer a surprisingly candid look at where multi-agent coordination actually matters, and where it doesn't, on the road to AGI

A team of roughly 10,000 AI agents working together cracked one of mathematics’ most famous unsolved problems in under four days. And according to the OpenAI researcher who helped make it happen, the multi-agent coordination part barely mattered.

Noam Brown, speaking on the Dwarkesh Podcast on September 17, 2026, laid out a vision of AGI that’s equal parts impressive and humbling. The headline result, solving the Navier-Stokes Millennium Prize Problem using 130 billion tokens across 88 hours, sounds like a triumph for collaborative AI. But Brown estimated that multi-agent coordination contributed less than 10% of the breakthrough.

The automated research intern, still needs a supervisor

OpenAI’s current crop of AI agents function as what the company calls “automated research interns.” These systems can deliver outputs 3.1 times what human researchers produce on multi-day tasks. Over 50% of those tasks still require human oversight to stay on track.

Advertisement

Brown pinpointed the bottleneck with unusual specificity. The missing ingredient is what he called “research taste,” the intuitive ability to judge whether a research direction is genuinely significant and original. That quality, Brown noted, represents roughly 10% of his own work. But it’s the 10% that everything else depends on.

Scaling agents hits diminishing returns fast

Agent scaling is sublinear. Four agents working together approximately halve the latency compared to a single agent on suitable tasks. That sounds promising until you realize the gains diminish quickly as you add more agents to the mix.

What Brown found compelling wasn’t the raw scaling numbers. It was the emergent behavior: agents spontaneously developing human-like collaboration patterns without being explicitly programmed to do so. They communicate, delegate, and refine work using straightforward messaging tools in ways that mirror how research teams actually function. That emergent coordination, Brown suggested, offers real clues about what AGI systems might eventually look like.

The March 2028 target and recursive self-improvement

OpenAI has set an internal objective to achieve a fully autonomous AI researcher by March 2028. No human oversight required. No research taste borrowed from human supervisors.

The key mechanism for getting there is recursive self-improvement, or RSI. Brown flagged RSI as OpenAI’s top priority while also acknowledging the risks associated with cooperative multi-agent behavior.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.