AI cracks second FrontierMath benchmark problem on absolute Galois groups, signaling a shift for computational research

Via storyofmathematics.com

AI cracks second FrontierMath benchmark problem on absolute Galois groups, signaling a shift for computational research

Claude Fable 5 and GPT-5.5 Pro both generated verified solutions to a number theory problem that would take human experts months to solve, raising questions about AI's growing role in fields from pure math to cryptography.

Two of the world’s most advanced AI models just solved the same unsolved math problem within two weeks of each other. The problem involves the absolute Galois group of the 2-adic numbers.

Epoch AI’s FrontierMath benchmark recorded its second verified solution on June 24, 2026. Anthropic’s Claude Fable 5 got there first on June 9, followed by OpenAI’s GPT-5.5 Pro fifteen days later. Both solutions were accepted by the benchmark’s verifier, though the broader mathematical community notes that full proof verification is still ongoing as of July 6.

Advertisement

What exactly did these models solve

The absolute Galois group of ℚ₂ (the 2-adic numbers) sits at the intersection of number theory and abstract algebra. Galois groups describe the symmetries of solutions to polynomial equations. The 2-adic numbers are an alternative number system that mathematicians use to study properties related to the prime number 2. While solutions are known for other prime cases, the case of ℚ₂ had eluded resolution until these AI developments.

The problem is classified as a “solid result” in number theory. Human experts would typically need one to three months to produce a solution. David Roe from MIT, who contributed the problem to the FrontierMath benchmark, rated the probability of eventual solvability at 95-99%.

FrontierMath exists specifically because earlier benchmarks were too easy. It comprises an Open Problems track featuring contributions from expert mathematicians addressing genuine unsolved research questions.

The July 6 update from the FrontierMath community emphasizes that these AI-generated solutions still require additional verification to confirm they constitute complete mathematical proofs. Passing an automated verifier is not the same as surviving peer review.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

AI cracks second FrontierMath benchmark problem on absolute Galois groups, signaling a shift for computational research

AI cracks second FrontierMath benchmark problem on absolute Galois groups, signaling a shift for computational research

Claude Fable 5 and GPT-5.5 Pro both generated verified solutions to a number theory problem that would take human experts months to solve, raising questions about AI's growing role in fields from pure math to cryptography.

Via storyofmathematics.com

Two of the world’s most advanced AI models just solved the same unsolved math problem within two weeks of each other. The problem involves the absolute Galois group of the 2-adic numbers.

Epoch AI’s FrontierMath benchmark recorded its second verified solution on June 24, 2026. Anthropic’s Claude Fable 5 got there first on June 9, followed by OpenAI’s GPT-5.5 Pro fifteen days later. Both solutions were accepted by the benchmark’s verifier, though the broader mathematical community notes that full proof verification is still ongoing as of July 6.

Advertisement

What exactly did these models solve

The absolute Galois group of ℚ₂ (the 2-adic numbers) sits at the intersection of number theory and abstract algebra. Galois groups describe the symmetries of solutions to polynomial equations. The 2-adic numbers are an alternative number system that mathematicians use to study properties related to the prime number 2. While solutions are known for other prime cases, the case of ℚ₂ had eluded resolution until these AI developments.

The problem is classified as a “solid result” in number theory. Human experts would typically need one to three months to produce a solution. David Roe from MIT, who contributed the problem to the FrontierMath benchmark, rated the probability of eventual solvability at 95-99%.

FrontierMath exists specifically because earlier benchmarks were too easy. It comprises an Open Problems track featuring contributions from expert mathematicians addressing genuine unsolved research questions.

The July 6 update from the FrontierMath community emphasizes that these AI-generated solutions still require additional verification to confirm they constitute complete mathematical proofs. Passing an automated verifier is not the same as surviving peer review.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.