Microsoft reports that MAI-Cyber-1-Flash, running inside MDASH, achieves 96% on the CyberGym benchmark — a custom Microsoft benchmark designed to test vulnerability detection, validation, and patching — which is 12 percentage points above Anthropic's Claude Mythos 5 (which scored roughly 84%) . According to Microsoft, the configuration also delivers roughly 50% cost savings compared to relying on third-party frontier models alone, because the in-house model handles about 90% of MDASH workflow tasks while reserving expensive models like OpenAI's GPT-5.4 only for the most complex edge cases
.
Independent verification note: These performance and cost figures are vendor-stated claims from Microsoft. According to reporting from Digital Applied, no third party has independently verified the 96% CyberGym score, and the CyberGym leaderboard entries are self-reported by participating labs . The New York Times also noted that Microsoft did not share the model with independent testers for evaluation before release
.
Microsoft published a Model Card for MAI-Cyber-1-Flash documenting its safety calibration, fine-tuning data, and intended use within the MDASH harness . The system uses a debate mechanism between multiple models plus a dedicated proof pipeline to eliminate false positives before any finding reaches a human engineer
.
MDASH (Multi-Model Agentic Scanning Harness) is the underlying system that MAI-Cyber-1-Flash operates inside. It orchestrates multiple LLMs — including Microsoft's model, OpenAI's GPT-5.4, and others — in a coordinated scanning, validation, and remediation workflow . The model-routing logic assigns routine tasks to the cheaper MAI-Cyber-1-Flash and escalates only the hardest reasoning tasks to larger models
.
MDASH has already demonstrated its value in production, before the formal announcement:
The vulnerabilities were found using a multi-model debate system where AI agents argued about whether a finding was exploitable before it was escalated to human engineers .
Project Perception is the commercial platform built on top of MDASH and MAI-Cyber-1-Flash. Microsoft positions it as a "new cyber stack" that combines signals, security context, specialized models, and agents into a continuously learning defense system designed for "using AI to defend against AI" .
The system routes tasks across models based on efficiency — it can deploy Microsoft's MAI-Cyber-1-Flash, OpenAI's GPT-5.4, or Anthropic models depending on which is best-suited for the specific workflow step .
While initially focused on software vulnerability management, Microsoft described Project Perception as expanding into broader security workflows — continuous monitoring, threat investigation, and automated remediation — forming a "continuously learning system of defense" .
Project Perception enters public preview on August 3, 2026 for Microsoft business customers already testing the MDASH agent harness . It will initially be available inside Microsoft Defender, with plans to roll out across all Microsoft Security products
.
Monday's launch comes amid an intensifying AI cybersecurity arms race. Anthropic had previously set the pace with its Claude Mythos security model, and Google and OpenAI were also investing heavily in AI-powered vulnerability discovery . Microsoft explicitly framed the announcement around the need to counter AI-powered attackers with AI-powered defense, arguing that traditional security tools cannot keep pace
.
The model-routing architecture — using a small, cheap specialist model alongside expensive frontier models — is a competitive differentiator on cost. Enterprises facing budget pressure on AI security tools see the 50% cost reduction as a major selling point .
With Project Perception entering public preview on August 3 and MDASH already proving itself in production with 16 real-world CVEs found in a single patch cycle, Microsoft is betting that live vulnerability discovery track records will matter more than benchmark scores alone in enterprise purchasing decisions . The question now is whether independent validation will back up Microsoft's claims — and how quickly enterprises will trust an AI system to continuously probe their own infrastructure.