Via news.microsoft.com
Microsoft’s MDASH outperforms Claude Mythos, GPT-5.6 in cybersecurity test
Microsoft's multi-agent AI system scored 88.45% on the CyberGym benchmark, beating out single-model competitors from Anthropic and OpenAI
Microsoft just dropped what might be the most compelling argument yet for why the future of AI security isn’t about building one really smart model. It’s about building a hundred of them and teaching them to work together.
The company unveiled MDASH, short for Microsoft Security multi-model agentic scanning harness, a cybersecurity system that orchestrates over 100 specialized AI agents to hunt down and fix vulnerabilities across complex codebases. On the CyberGym benchmark, MDASH scored 88.45%, comfortably beating Anthropic’s Claude Mythos Preview at 83.1% and OpenAI’s GPT-5.5 at 81.8%.
A swarm beats a single brain
Rather than relying on a single large language model to do everything, Microsoft’s Autonomous Code Security team built an orchestrated system where each agent handles a specific slice of the vulnerability detection and remediation process.
The approach already proved its worth during the May 2026 Patch Tuesday, where MDASH identified 16 new Windows vulnerabilities. Four of those were critical remote code execution flaws.
Microsoft’s team developed the system drawing on experience from DARPA’s AI Cyber Challenge. The system is integrated with Microsoft Defender and GitHub Code Security.
By early June 2026, MDASH’s CyberGym scores climbed to 96.55%. When the team integrated lighter models like MAI-Cyber-1-Flash in July 2026, the system still held at 96%.
Why crypto and Web3 should pay attention
Smart contract exploits and bridge hacks have drained billions from the crypto ecosystem over the past few years. If multi-agent AI systems like MDASH can scan complex codebases across multiple programming languages with near-perfect accuracy, the technology could eventually be adapted or replicated for auditing Solidity, Rust, and Move smart contracts.
What this means for investors
MDASH is currently available only in limited private preview for select clients.
Microsoft is betting that the future of AI-powered security belongs to orchestrated multi-agent systems rather than monolithic models. If that bet proves correct, it puts Anthropic and OpenAI at a structural disadvantage in the security vertical, since both companies have primarily competed on the strength of individual models.
MDASH’s ability to automate both detection and remediation could compress the timeline for enterprise vulnerability management from weeks to hours.