Meta official logo (public domain, Wikimedia Commons) — CryptoBriefing brand treatment
Meta study shows two AI coding agents catch more bugs than one with a bigger budget
Pairing coding agents to review each other's patches beat simply giving a single agent more resources, adding to Meta's growing pile of automated code review data
Give one AI coding agent more budget and it gets a little better at finding bugs. Give it a partner to check its work, and it gets a lot better.
That is the central finding of new research tied to Meta: having two coding agents review each other’s patches improves bug detection more than increasing a single agent’s budget.
What Meta’s numbers actually show
The peer-review finding lands on top of a substantial body of Meta data on automated code review. The headline system is RADAR, Meta’s internal review tool.
RADAR has reviewed over 535,000 diffs. Of those, more than 331,000 were landed, meaning they were merged into the codebase. RADAR also trimmed median review wall time by 35%.
Meta says RADAR uses risk calibration, which means it adjusts how cautious it is based on how dangerous a given change looks. With that calibration in place, Meta reports a lower revert rate than manual review. Production incidents fell to one-fiftieth of what manual reviews produced.
Then there is Meta’s Engineering Agent, which was tasked with repairing test failures over a three-month trial. Of the fixes it generated, 80% received human review. Of those reviewed fixes, approximately 25.5% were accepted into production.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
Testing that hunts for bugs on its own
Meta’s Just-in-Time testing framework combines large language models, program analysis and mutation testing. Mutation testing works like a fire drill for your test suite: you plant small, deliberate faults in the code and check whether the tests notice.
Across more than 22,000 generated tests, the approach reportedly produced a fourfold increase in bug detection. For significant failures, the improvement reached up to twentyfold.
The broader case for agents checking agents
A protocol called Adversarial Review, tested on the LiveCodeBench coding benchmark, achieved an 87% pass rate. That beat single-agent setups and even some configurations using five agents.
Another system, Wink, focuses on catching coding agents when they go off the rails. It recovered from approximately 90% of misbehaviors in coding agent trajectories across more than 10,000 real-world instances.
What this means for engineering teams
The Engineering Agent’s fixes still passed through human review, and only approximately 25.5% of those reviewed made it to production. Engineers are shifting from writing every fix to judging which machine-written fixes deserve to ship.