T
hree minutes. That was the length of the demo Microsoft ran Monday at a small event in San Francisco: a simulated enterprise codebase, a single prompt, and then a cluster of vulnerabilities identified, triaged by severity, matched to a detection rule, and patched with working code. Dave Weston, the engineer who built the system, stood next to the screen and said what used to take hours of work from appsec hunters and remediation engineers now happens in minutes.
The tool doing the work is MAI-Cyber-1-Flash, Microsoft's first purpose-built cybersecurity model, launched alongside a broader agentic platform called Project Perception. Both are Microsoft's answer to a race that Anthropic and OpenAI have effectively been running alone until now Anthropic with its Mythos models, OpenAI with a security-focused release of its own.
Microsoft's opening move wasn't to claim it built the smartest model in the room. It claimed something more specific: that its combination is nearly as good, and roughly half the price.
Run inside MDASH Microsoft's existing multi-agent harness for finding and fixing software vulnerabilities MAI-Cyber-1-Flash paired with OpenAI's GPT-5.4 scored 95.95% on CyberGym, a benchmark that tests how well AI systems reason through large codebases to find real, exploitable bugs. Microsoft's Mustafa Suleyman, the former DeepMind co-founder who now runs Microsoft AI, rattled off the field it claims to have beaten at the event: Gemini, GPT-5.5 Cyber, GPT-5.6 Sol, and Anthropic's Mythos 5. He put the margin over Mythos at 12 percentage points.
The cost framing matters as much as the score. In this setup, MAI-Cyber-1-Flash handles roughly 90% of the workload, with GPT-5.4 called in only for the hardest 10% of cases a routing decision Microsoft says cuts the bill by half compared with its previous best MDASH configuration, which leaned on GPT-5.4, GPT-5.4 mini, and GPT-5.3 Codex together. "We have world-leading performance at 50% of the cost," Suleyman said, according to CNBC.
It's worth being precise about what that 95.95% actually measures, since it's easy to read past. The headline score belongs to the MDASH system running MAI-Cyber-1-Flash alongside GPT-5.4 not to the new model working alone, and not on general security tasks.
CyberGym's Level 1 tier, the one Microsoft cites, is built around reproducing already-known vulnerabilities rather than discovering novel ones. Access, for now, is limited to approved MDASH customers through a private preview on Azure AI Foundry.
It isn't available as a standalone model or a general-purpose API for anyone to test independently.
MAI-Cyber-1-Flash is really the opening scenario for something bigger. Project Perception, which enters public preview on August 3, is designed to coordinate teams of security agents Microsoft describes them in red, blue, and green groupings across detection, posture management, and remediation, with vulnerability management as the first workflow it's been pointed at. Hayete Gallot, who returned to Microsoft in February to run the security business, framed the pitch around a threat that's no longer hypothetical: attackers already use AI to move faster, so defenders need AI tools built for the same speed.
That's not an abstract talking point in Microsoft's own numbers.
A related deployment of the MDASH system reportedly caught $7.7 million worth of vulnerabilities over three months before Monday's announcement the kind of figure a sales team likes to have in its pocket. And Suleyman was blunt that this is a first step rather than a finished product: Microsoft's data, its harness architecture, and its accumulated security expertise, he said, give it a moat that will keep improving, and the next model is going to be considerably better than this one.
Coverage since Monday has split roughly into two camps: outlets running the benchmark numbers more or less as Microsoft presented them, and outlets pointing out how much of the announcement rests on Microsoft's own testing, on a narrow slice of one benchmark, and on a product still gated behind a private preview.
The Register's take was characteristically dry, framing the whole rollout as Microsoft's solution to AI-driven security threats being more AI, wrapped in more acronyms. There's also a smaller, slightly awkward detail buried in the architecture: for all the talk of an in-house model, roughly a tenth of the actual work in this "Microsoft" system still runs on a model built by OpenAI.
None of that undercuts the basic strategic point Microsoft is making, which has less to do with cybersecurity specifically and more to do with how the company wants enterprises to think about buying AI going forward not chasing the single most capable model available, but routing intelligently between a cheap one and an expensive one depending on what a task actually needs. Whether CISOs buy that argument as readily as Microsoft's sales deck assumes they will is the part nobody can benchmark yet.












