Nvidia researchers improve AI agent reliability with a judging model
A new method called Mid-Harness has a verifier model pick the best command before a terminal agent runs anything
Nvidia researchers have found a fairly human fix for unreliable AI agents: think before you type. Their new method, called Mid-Harness, has an agent generate several possible actions, then lets a separate judge model choose which one actually runs.
On one benchmark, that extra moment of deliberation lifted the first-try success rate from 50.00% to 68.03%. For software that operates inside a command line, where one bad command can sink an entire task, that matters.
How Mid-Harness works
Mid-Harness has a generator model sample multiple candidate actions at each step, and a second model, the verifier, scores them. Only the top choice is forwarded for execution.
The target is what researchers call terminal agents. These are AI systems that work in command-line interfaces or call external tools, typing real commands into real environments.
Those environments are stochastic, meaning outcomes can vary unpredictably from one run to the next. An agent might usually know the right command and still fumble it at the worst possible moment.
The paper frames this as a gap between generation and reliable execution. A model can be capable of producing a useful command without consistently producing it when it counts.
Mid-Harness aims to close that gap without retraining the generator. It simply adds a filter between thinking and doing.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
The numbers on TerminalBench-Lite
The team tested the method on TerminalBench-Lite, a benchmark for terminal agent tasks. Their headline result paired a TMAX-9B generator with GPT-5.6 Sol acting as the verifier.
With eight sampled actions per step, Pass@1 climbed from 50.00% to 68.03%. Pass@1 measures how often the agent succeeds on its first and only attempt, which is the scenario that matters most in production.
According to the research findings, the most effective results emerged when TMAX-9B served as both generator and verifier, grading its own homework. It suggests a model may be better at recognizing a good command than reliably producing one on the first try.
The researchers also compared two strategies for spending extra compute. Action scaling, the Mid-Harness approach, samples many options at each individual step. Trajectory scaling instead runs entire task attempts multiple times and picks the best overall run.
According to the findings, action scaling delivered better outcomes at a lower token cost than trajectory scaling alone. The two approaches also worked well together. The research reports consistent improvements across diverse models and benchmarks, not just the single headline pairing.
Background: Nvidia’s push on agent reliability
The paper was published on December 16, 2026, by authors including Minki Kang and Ehsan Hosseini-Asl of Nvidia. Earlier in 2026, Nvidia introduced ACES, a framework published around August 2026 for measuring the real runtime impact of agent skills. In September 2026, it followed with the Open Agent Safety Platform.
What this means
If action-level verification beats trajectory-level retries on token efficiency, as the findings indicate, it offers a cheaper path to better reliability. Checking each step costs less than redoing the whole job.
If a 9-billion-parameter model can effectively judge its own candidate actions, teams may not need to pay for a second, larger model to act as the referee, simplifying deployment and keeping costs down.
There are caveats. A 68.03% Pass@1 rate is a real improvement, but it still means roughly one task in three fails on the first attempt. Benchmarks are also controlled settings, and how well Mid-Harness holds up across messier real-world environments, longer tasks and different toolchains remains to be tested.