Photo: Tima Miroshnichenko / Pexels
InnoEval and new benchmarks show AI models struggle with original research
A wave of 2025-2026 studies finds frontier AI models are far better at remixing known science than inventing new methods
Frontier AI models can summarize a thousand papers before your coffee cools. Ask them to invent a genuinely new research technique from scratch, and the results get a lot less impressive.
That is the shared conclusion of a cluster of studies from 2025 and 2026, led by a new evaluation framework called InnoEval. Together, they suggest AI models underperform when asked to develop new research methods without prior information to lean on.
What the research actually found
InnoEval first appeared in an arXiv paper in February 2026 and was later presented at ICML 2026. Its goal is to measure innovation in a structured way, using what its authors describe as knowledge-grounded, multi-perspective evaluations.
That approach paid off on its own terms. InnoEval beat traditional LLM-as-judge methods by up to 16.18% in F1 scores on specific tasks.
Better measurement, however, produced some unflattering readings. The more carefully researchers looked, the less originality they found.
Take the “Reconstruction” benchmark. It tested whether AI models could recover the core idea of a paper using only the reference list that existed before publication. Across 643 papers, as of August 2026, models recovered the central idea only 3ā15% of the time.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
A study published in Science in October 2026 compared AI-generated hypotheses with human ones directly. The AI hypotheses scored an average novelty of 3.406. Human hypotheses averaged 3.968.
The diversity problem
A separate large-scale study, also from October 2026, analyzed more than 121,000 preprints. It found that LLMs often generate ideas that are narrowly focused and show little diversity.
Two more benchmarks, InnoGym and InnovatorBench, were introduced across 2025 and 2026 to probe the same question from a workflow perspective. Rather than grading a single answer, they test complete LLM pipelines on a range of real-world tasks. The results reported low success rates in producing genuinely new methods.
Why this keeps coming up
Until recently, a common way to judge AI-generated ideas was to ask another AI model whether they were novel. Frameworks like InnoEval exist because that approach was not reliable enough. The 16.18% F1 improvement suggests grounding evaluations in real knowledge catches things that a lone LLM judge misses.
What this means for AI labs, researchers and investors
Current models appear well suited as research assistants that surface literature and test hypotheses. They appear much less suited to serving as the source of the hypotheses themselves, at least when working without strong prior context. The 121,000-preprint study hints at the risk that a tool which reliably proposes the same narrow cluster of ideas could quietly reduce the diversity of research directions a lab explores.
None of this means AI cannot contribute to original research. What the studies do establish is a clearer baseline, with numbers like a 3ā15% recovery rate and a novelty score of 3.406 against 3.968 that future systems will have to beat.