Epoch’s Automation Reports find frontier AI still can’t do Epoch’s job

Epoch’s Automation Reports find frontier AI still can’t do Epoch’s job

Top models from Anthropic and OpenAI tie at approximately 65% on tasks drawn from Epoch AI's own research, but judgment and style remain weak spots

Epoch AI spends much of its time measuring how capable AI models are becoming. With its latest release, the organization turned the question inward: could these models actually do Epoch’s own work?

According to the new Epoch Automation Reports, not yet. Frontier models get close on clearly defined tasks, but they still fall short of fully automating the work, especially when assignments turn open-ended or call for judgment.

A job trial, not a pop quiz

Epoch AI introduced the Automation Reports on October 8, 2026. Instead of relying on abstract puzzles, the benchmark pulls its tasks from the organization’s internal research.

The evaluation covers 11 distinct tasks spread across five categories. Human graders score each model’s output against the same internal quality rubrics Epoch uses for its own work.

Advertisement

The scoreboard

Two models share the top spot. Anthropic’s Claude Fable 5.1 and OpenAI’s GPT-6 Astra each posted an average score of approximately 65% across the tasks.

Grok 4.6 followed at 59%. Qwen 3.8 Max came in at 53%, edging out Kimi K3 at 52%.

Gemini 3.8 Flash rounded out the listed results at 42%. That leaves a spread of more than 20 points between the leaders and the bottom of the reported field.

The models performed well on defined subtasks such as coding and computational analysis. Where they faltered was in creative and judgment-based work, with recurring problems in style adherence and content targeting.

Every model in the evaluation lagged on outputs that depend on implicit conventions — the house style, the expected framing, the sense of what belongs in a report and what does not.

Where this fits in Epoch’s work

The Automation Reports are part of Epoch AI’s broader benchmarking efforts. They complement the Epoch Capabilities Index, or ECI, with the aim of offering a more holistic view of model performance than existing frameworks provide.

Around October 7, 2026, Epoch also rolled out InnovationEval, which tests how well AI can automate research itself.

What this means

For companies deciding where to deploy AI, the findings draw a useful line between structured and open-ended work. Tasks with a clear right answer, like code or computation, look increasingly within reach for the top models.

There are caveats worth keeping in mind. The benchmark reflects 11 tasks drawn from a single organization’s research, so the results describe how well models handle Epoch’s work specifically, not every knowledge job. Human grading against internal rubrics also brings a degree of subjectivity that automated scoring avoids, making direct comparisons with other benchmarks harder.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.
Epoch’s Automation Reports find frontier AI still can’t do Epoch’s job
Epoch’s Automation Reports find frontier AI still can’t do Epoch’s job

Top models from Anthropic and OpenAI tie at approximately 65% on tasks drawn from Epoch AI's own research, but judgment and style remain weak spots

Epoch AI spends much of its time measuring how capable AI models are becoming. With its latest release, the organization turned the question inward: could these models actually do Epoch’s own work?

According to the new Epoch Automation Reports, not yet. Frontier models get close on clearly defined tasks, but they still fall short of fully automating the work, especially when assignments turn open-ended or call for judgment.

A job trial, not a pop quiz

Epoch AI introduced the Automation Reports on October 8, 2026. Instead of relying on abstract puzzles, the benchmark pulls its tasks from the organization’s internal research.

The evaluation covers 11 distinct tasks spread across five categories. Human graders score each model’s output against the same internal quality rubrics Epoch uses for its own work.

Advertisement

The scoreboard

Two models share the top spot. Anthropic’s Claude Fable 5.1 and OpenAI’s GPT-6 Astra each posted an average score of approximately 65% across the tasks.

Grok 4.6 followed at 59%. Qwen 3.8 Max came in at 53%, edging out Kimi K3 at 52%.

Gemini 3.8 Flash rounded out the listed results at 42%. That leaves a spread of more than 20 points between the leaders and the bottom of the reported field.

The models performed well on defined subtasks such as coding and computational analysis. Where they faltered was in creative and judgment-based work, with recurring problems in style adherence and content targeting.

Every model in the evaluation lagged on outputs that depend on implicit conventions — the house style, the expected framing, the sense of what belongs in a report and what does not.

Where this fits in Epoch’s work

The Automation Reports are part of Epoch AI’s broader benchmarking efforts. They complement the Epoch Capabilities Index, or ECI, with the aim of offering a more holistic view of model performance than existing frameworks provide.

Around October 7, 2026, Epoch also rolled out InnovationEval, which tests how well AI can automate research itself.

What this means

For companies deciding where to deploy AI, the findings draw a useful line between structured and open-ended work. Tasks with a clear right answer, like code or computation, look increasingly within reach for the top models.

There are caveats worth keeping in mind. The benchmark reflects 11 tasks drawn from a single organization’s research, so the results describe how well models handle Epoch’s work specifically, not every knowledge job. Human grading against internal rubrics also brings a degree of subjectivity that automated scoring avoids, making direct comparisons with other benchmarks harder.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.