Meta study shows two AI coding agents catch more bugs than one with a bigger budget

Meta official logo (public domain, Wikimedia Commons) — CryptoBriefing brand treatment

Meta study shows two AI coding agents catch more bugs than one with a bigger budget

Pairing coding agents to review each other's patches beat simply giving a single agent more resources, adding to Meta's growing pile of automated code review data

Give one AI coding agent more budget and it gets a little better at finding bugs. Give it a partner to check its work, and it gets a lot better.

That is the central finding of new research tied to Meta: having two coding agents review each other’s patches improves bug detection more than increasing a single agent’s budget.

What Meta’s numbers actually show

The peer-review finding lands on top of a substantial body of Meta data on automated code review. The headline system is RADAR, Meta’s internal review tool.

RADAR has reviewed over 535,000 diffs. Of those, more than 331,000 were landed, meaning they were merged into the codebase. RADAR also trimmed median review wall time by 35%.

Advertisement

Meta says RADAR uses risk calibration, which means it adjusts how cautious it is based on how dangerous a given change looks. With that calibration in place, Meta reports a lower revert rate than manual review. Production incidents fell to one-fiftieth of what manual reviews produced.

Then there is Meta’s Engineering Agent, which was tasked with repairing test failures over a three-month trial. Of the fixes it generated, 80% received human review. Of those reviewed fixes, approximately 25.5% were accepted into production.

Testing that hunts for bugs on its own

Meta’s Just-in-Time testing framework combines large language models, program analysis and mutation testing. Mutation testing works like a fire drill for your test suite: you plant small, deliberate faults in the code and check whether the tests notice.

Across more than 22,000 generated tests, the approach reportedly produced a fourfold increase in bug detection. For significant failures, the improvement reached up to twentyfold.

The broader case for agents checking agents

A protocol called Adversarial Review, tested on the LiveCodeBench coding benchmark, achieved an 87% pass rate. That beat single-agent setups and even some configurations using five agents.

Another system, Wink, focuses on catching coding agents when they go off the rails. It recovered from approximately 90% of misbehaviors in coding agent trajectories across more than 10,000 real-world instances.

What this means for engineering teams

The Engineering Agent’s fixes still passed through human review, and only approximately 25.5% of those reviewed made it to production. Engineers are shifting from writing every fix to judging which machine-written fixes deserve to ship.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.
Meta study shows two AI coding agents catch more bugs than one with a bigger budget
Meta study shows two AI coding agents catch more bugs than one with a bigger budget

Pairing coding agents to review each other's patches beat simply giving a single agent more resources, adding to Meta's growing pile of automated code review data

Meta official logo (public domain, Wikimedia Commons) — CryptoBriefing brand treatment

Give one AI coding agent more budget and it gets a little better at finding bugs. Give it a partner to check its work, and it gets a lot better.

That is the central finding of new research tied to Meta: having two coding agents review each other’s patches improves bug detection more than increasing a single agent’s budget.

What Meta’s numbers actually show

The peer-review finding lands on top of a substantial body of Meta data on automated code review. The headline system is RADAR, Meta’s internal review tool.

RADAR has reviewed over 535,000 diffs. Of those, more than 331,000 were landed, meaning they were merged into the codebase. RADAR also trimmed median review wall time by 35%.

Advertisement

Meta says RADAR uses risk calibration, which means it adjusts how cautious it is based on how dangerous a given change looks. With that calibration in place, Meta reports a lower revert rate than manual review. Production incidents fell to one-fiftieth of what manual reviews produced.

Then there is Meta’s Engineering Agent, which was tasked with repairing test failures over a three-month trial. Of the fixes it generated, 80% received human review. Of those reviewed fixes, approximately 25.5% were accepted into production.

Testing that hunts for bugs on its own

Meta’s Just-in-Time testing framework combines large language models, program analysis and mutation testing. Mutation testing works like a fire drill for your test suite: you plant small, deliberate faults in the code and check whether the tests notice.

Across more than 22,000 generated tests, the approach reportedly produced a fourfold increase in bug detection. For significant failures, the improvement reached up to twentyfold.

The broader case for agents checking agents

A protocol called Adversarial Review, tested on the LiveCodeBench coding benchmark, achieved an 87% pass rate. That beat single-agent setups and even some configurations using five agents.

Another system, Wink, focuses on catching coding agents when they go off the rails. It recovered from approximately 90% of misbehaviors in coding agent trajectories across more than 10,000 real-world instances.

What this means for engineering teams

The Engineering Agent’s fixes still passed through human review, and only approximately 25.5% of those reviewed made it to production. Engineers are shifting from writing every fix to judging which machine-written fixes deserve to ship.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.