Microsoft’s MDASH outperforms Claude Mythos, GPT-5.6 in cybersecurity test

Via news.microsoft.com

Microsoft’s MDASH outperforms Claude Mythos, GPT-5.6 in cybersecurity test

Microsoft's multi-agent AI system scored 88.45% on the CyberGym benchmark, beating out single-model competitors from Anthropic and OpenAI

Microsoft just dropped what might be the most compelling argument yet for why the future of AI security isn’t about building one really smart model. It’s about building a hundred of them and teaching them to work together.

The company unveiled MDASH, short for Microsoft Security multi-model agentic scanning harness, a cybersecurity system that orchestrates over 100 specialized AI agents to hunt down and fix vulnerabilities across complex codebases. On the CyberGym benchmark, MDASH scored 88.45%, comfortably beating Anthropic’s Claude Mythos Preview at 83.1% and OpenAI’s GPT-5.5 at 81.8%.

A swarm beats a single brain

Rather than relying on a single large language model to do everything, Microsoft’s Autonomous Code Security team built an orchestrated system where each agent handles a specific slice of the vulnerability detection and remediation process.

Advertisement

The approach already proved its worth during the May 2026 Patch Tuesday, where MDASH identified 16 new Windows vulnerabilities. Four of those were critical remote code execution flaws.

Microsoft’s team developed the system drawing on experience from DARPA’s AI Cyber Challenge. The system is integrated with Microsoft Defender and GitHub Code Security.

By early June 2026, MDASH’s CyberGym scores climbed to 96.55%. When the team integrated lighter models like MAI-Cyber-1-Flash in July 2026, the system still held at 96%.

Why crypto and Web3 should pay attention

Smart contract exploits and bridge hacks have drained billions from the crypto ecosystem over the past few years. If multi-agent AI systems like MDASH can scan complex codebases across multiple programming languages with near-perfect accuracy, the technology could eventually be adapted or replicated for auditing Solidity, Rust, and Move smart contracts.

What this means for investors

MDASH is currently available only in limited private preview for select clients.

Microsoft is betting that the future of AI-powered security belongs to orchestrated multi-agent systems rather than monolithic models. If that bet proves correct, it puts Anthropic and OpenAI at a structural disadvantage in the security vertical, since both companies have primarily competed on the strength of individual models.

MDASH’s ability to automate both detection and remediation could compress the timeline for enterprise vulnerability management from weeks to hours.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Microsoft’s MDASH outperforms Claude Mythos, GPT-5.6 in cybersecurity test

Microsoft’s MDASH outperforms Claude Mythos, GPT-5.6 in cybersecurity test

Microsoft's multi-agent AI system scored 88.45% on the CyberGym benchmark, beating out single-model competitors from Anthropic and OpenAI

Via news.microsoft.com

Microsoft just dropped what might be the most compelling argument yet for why the future of AI security isn’t about building one really smart model. It’s about building a hundred of them and teaching them to work together.

The company unveiled MDASH, short for Microsoft Security multi-model agentic scanning harness, a cybersecurity system that orchestrates over 100 specialized AI agents to hunt down and fix vulnerabilities across complex codebases. On the CyberGym benchmark, MDASH scored 88.45%, comfortably beating Anthropic’s Claude Mythos Preview at 83.1% and OpenAI’s GPT-5.5 at 81.8%.

A swarm beats a single brain

Rather than relying on a single large language model to do everything, Microsoft’s Autonomous Code Security team built an orchestrated system where each agent handles a specific slice of the vulnerability detection and remediation process.

Advertisement

The approach already proved its worth during the May 2026 Patch Tuesday, where MDASH identified 16 new Windows vulnerabilities. Four of those were critical remote code execution flaws.

Microsoft’s team developed the system drawing on experience from DARPA’s AI Cyber Challenge. The system is integrated with Microsoft Defender and GitHub Code Security.

By early June 2026, MDASH’s CyberGym scores climbed to 96.55%. When the team integrated lighter models like MAI-Cyber-1-Flash in July 2026, the system still held at 96%.

Why crypto and Web3 should pay attention

Smart contract exploits and bridge hacks have drained billions from the crypto ecosystem over the past few years. If multi-agent AI systems like MDASH can scan complex codebases across multiple programming languages with near-perfect accuracy, the technology could eventually be adapted or replicated for auditing Solidity, Rust, and Move smart contracts.

What this means for investors

MDASH is currently available only in limited private preview for select clients.

Microsoft is betting that the future of AI-powered security belongs to orchestrated multi-agent systems rather than monolithic models. If that bet proves correct, it puts Anthropic and OpenAI at a structural disadvantage in the security vertical, since both companies have primarily competed on the strength of individual models.

MDASH’s ability to automate both detection and remediation could compress the timeline for enterprise vulnerability management from weeks to hours.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.