Microsoft built a 100-agent harness to hunt bugs, and a cheap model to feed it
Microsoft shipped its first cybersecurity model today, MAI-Cyber-1-Flash, but the model isn't really the story. It's one component inside MDASH, a multi-agent vulnerability identification and remediation harness that runs more than a hundred agents to find and fix security holes in complex codebases. The design is the now-familiar cheap-and-strong split: the compact Flash model, derived from the MAI-Thinking-1 lineage, handles the routine security work, and the genuinely hard problems get routed up to bigger models. Net result Microsoft claims is 50% cost savings.
The number that makes people look is 95.95%. On CyberGym, the standard benchmark for code vulnerability detection, MDASH with MAI-Cyber-1-Flash hit that accuracy, beating comparable setups running Gemini and GPT models. So this isn't a lone model taking a benchmark, it's an orchestrated fleet of agents plus a purpose-built cheap model beating single-model approaches at finding vulnerabilities.
The context is what makes this land. This is the same month OpenAI's Hugging Face breach reignited the whole argument about whether frontier models can be trusted to poke at attacks, and the same lane as the containment failures the security world has been chewing on all summer. Microsoft's answer is not one big secure model, it's a harness of many agents with a cheap specialist doing the grunt work and escalation for the hard cases, and it beat the general-purpose frontier models on the security benchmark that counts.
That's the pattern worth taking away. The interesting security systems of 2026 aren't a single genius model, they're an architecture, where you decide which parts stay dumb and cheap, which parts escalate, and how a hundred agents coordinate on a codebase without stepping on each other. MDASH is Microsoft planting a flag on that shape. Details: https://microsoft.ai/news/introducing-mai-cyber-1-flash-inside-mdash/
← Back to all articles
The number that makes people look is 95.95%. On CyberGym, the standard benchmark for code vulnerability detection, MDASH with MAI-Cyber-1-Flash hit that accuracy, beating comparable setups running Gemini and GPT models. So this isn't a lone model taking a benchmark, it's an orchestrated fleet of agents plus a purpose-built cheap model beating single-model approaches at finding vulnerabilities.
The context is what makes this land. This is the same month OpenAI's Hugging Face breach reignited the whole argument about whether frontier models can be trusted to poke at attacks, and the same lane as the containment failures the security world has been chewing on all summer. Microsoft's answer is not one big secure model, it's a harness of many agents with a cheap specialist doing the grunt work and escalation for the hard cases, and it beat the general-purpose frontier models on the security benchmark that counts.
That's the pattern worth taking away. The interesting security systems of 2026 aren't a single genius model, they're an architecture, where you decide which parts stay dumb and cheap, which parts escalate, and how a hundred agents coordinate on a codebase without stepping on each other. MDASH is Microsoft planting a flag on that shape. Details: https://microsoft.ai/news/introducing-mai-cyber-1-flash-inside-mdash/
Comments