Dark-themed email security dashboard showing threats, filtering results, containment, and remediation metrics.
Microsoft says Defender for Office 365 missed 221 high-severity email threats per 1,000 protected users from May through July 2026, 55.4% fewer than the next-closest secure email gateway in its latest comparison. The important operational finding is less flattering to any single-product narrative: Microsoft’s own data says third-party integrated email-security tools still add their clearest measurable value in sorting promotional and bulk mail, while their average incremental catches for malicious messages and spam remain small.

The figures were published September 17 by Microsoft Security as the fifth consecutive quarterly update to its Defender email-security benchmarking. They are useful evidence for Microsoft 365 administrators, but they are not an independent bake-off. Microsoft collects the telemetry, defines the threat categories, chooses the compared deployment populations, and publishes the results. The company’s public benchmark page now reflects the May–July period, but the underlying data set, customer mix, sample sizes, confidence intervals, and full vendor-by-vendor gateway totals are not disclosed.

That does not make the numbers worthless. It means IT teams should treat them as a directional operational signal—particularly about their own tenant’s mail flow and post-delivery remediation—not a procurement scorecard that settles the Defender-versus-gateway debate.

The 221 figure measures a narrow definition of failure​

Microsoft’s headline comparison covers secure email gateways, or SEGs: products that sit in front of Microsoft 365 in the mail path and attempt to stop threats before Exchange Online and Defender process them. Microsoft says it normalizes “missed” high-severity threats by protected user count rather than comparing total catches, an approach intended to reduce the effect of one vendor’s customers receiving more malicious traffic than another’s.

For the current report, Microsoft’s public performance page defines a missed SEG threat as one that was not detected before delivery. That wording deserves close attention. When Microsoft introduced the benchmark in July 2025, it described a SEG miss more broadly: a threat not detected before delivery or not removed shortly after delivery. It also said it applied a stricter rule to Defender itself, counting a message as missed even when Defender remediated it after delivery.

Microsoft has not explained in this quarter’s post whether the scoring model changed, whether the shorter current definition is simply a simplified description, or how the five quarterly periods should be compared if post-delivery handling is accounted for differently. Administrators should not assume a trend line across the reports is strictly comparable until Microsoft publishes the full methodology and confirms the definitions used in every period.

The current release does concede a more uncomfortable point: missed high-severity threats have risen across multiple reporting periods, including for Defender. Microsoft attributes the direction of travel in part to attackers using AI to research targets, tailor language and improve impersonation lures. That is consistent with the practical problem facing mail-security teams: a filter can improve relative to competitors while still confronting a higher absolute volume of convincing attacks.

A 92% post-delivery share is useful, but it is not prevention​

Microsoft says Defender handled 92% of malicious messages caught after delivery in benchmarked environments, with the remaining 8% attributed on average to integrated cloud email security, or ICES, products. These tools generally connect through APIs after mail reaches Microsoft 365, moving suspicious messages from Inbox or other folders to quarantine, junk, deleted items, or a vendor-managed location.

Post-delivery remediation is a genuine protection layer. Threat intelligence changes, campaigns are linked together after the first messages arrive, and a message initially considered benign can later be convicted. Defender’s Zero-hour Auto Purge capability and related remediation workflows exist for exactly this reason. For a security team, fast retroactive removal can turn one delivered phishing email into a contained event rather than a broad user-reporting incident.

But the number must not be read as “92% of malicious email was blocked.” It is Microsoft’s share of the messages that were identified and removed after delivery within the benchmark’s measurement model. A high remediation share can demonstrate strong detection and automated cleanup; it also confirms that some risky messages reached user-accessible mailboxes before new intelligence triggered action.

That timing matters most for credential phishing, business email compromise, QR-code lures, and malicious instructions aimed at AI assistants that read mailbox content. A user who opens a message, scans a QR code, approves an OAuth consent request, or acts on a fraudulent invoice before a purge job runs cannot be protected retroactively by the mailbox cleanup alone.

Proofpoint, a competitor with its own commercial interest in this argument, made that objection in a June blog post responding to vendor efficacy measurement. It argues that a gateway’s early blocks may be invisible to an email-provider benchmark when those messages never enter Microsoft’s environment, and that post-delivery removal should be treated as a backstop rather than proof of equivalent pre-delivery protection. Microsoft has not published data in this release that resolves that architectural disagreement. The two vendors are measuring different deployment positions and different threat populations, so their claims should not be compared as if they came from a shared test corpus.

The ICES result is an argument for targeted layering​

Microsoft’s strongest case for ICES layering is not malware or phishing detection. It is inbox noise. Across the vendors included in the May–July data, Microsoft says ICES tools added an average 19.6% improvement in promotional-email filtering, compared with 0.52% for spam and 0.30% for malicious content.

Those small malicious and spam uplifts are still potentially meaningful in a large tenant. An incremental 0.30% catch rate can represent real messages in an organization with tens of thousands of mailboxes and high inbound volume. But the average is not a universal answer: Microsoft’s public page shows materially different results for individual ICES vendors, including variation in overlap with Defender detections and in non-malicious messages moved by those tools.

The right conclusion for a Microsoft 365 administrator is layer where a demonstrated gap exists. A finance department repeatedly targeted by invoice fraud, an executive population subject to impersonation, or an organization with an unusually high phishing burden may justify an inline gateway or post-delivery tool even if the portfolio-wide average uplift looks modest. Conversely, buying a second security product primarily to reduce newsletters, offers, and low-priority bulk mail should be evaluated as a productivity and user-experience decision, not presented internally as a major anti-phishing control.

Microsoft’s report also puts a practical cost on extra filtering: its ICES comparison records “non-malicious” detections, which can include false positives or messages moved because of customer preferences. Security teams adding another filter should measure recovered malicious mail and legitimate business mail delayed, quarantined, or moved out of the Inbox. A tool that prevents several phishing messages but regularly traps vendor invoices, legal notices, or support tickets creates a different operational risk.

What to verify in a Defender for Office 365 tenant​

The report should prompt tenant-specific measurement rather than a rushed replacement or removal of a gateway. Microsoft’s Defender portal includes email and collaboration reporting, Threat Explorer, campaign views, and investigation tools that can be used to identify what reached users, what was remediated later, and which recipients are repeatedly targeted.

Administrators should review the last 30 to 90 days with a narrow set of questions in mind:

  • Compare pre-delivery blocks with post-delivery remediation, and measure the time a malicious message remained accessible before it was removed.
  • Separate phishing, malware, business email compromise, spam, and bulk mail instead of combining them into a single “blocked email” metric.
  • Review user-reported messages that were initially delivered, especially messages received by executives, finance staff, help-desk personnel, and users with payment authority.
  • Audit mail-flow rules, allow lists, spoof-intelligence overrides, and third-party routing configurations that may bypass or weaken Defender inspection.
  • Test whether the existing security stack catches representative attacks before delivery without generating unacceptable false positives in critical business workflows.

Microsoft also says a redesigned machine-learning and AI model stack reduced false negatives by roughly two-thirds and false positives by nearly one-fifth during a four-week internal observation period. Those are Microsoft research figures rather than independently published test results, and the company did not identify the customer population, baseline detection rate, or test dates. They should be validated through each organization’s own alert, quarantine, and user-report data before being used to justify a policy change.

Prompt injection makes inbox protection a broader control point​

The forward-looking part of Microsoft’s announcement is its claim that Defender can detect and isolate malicious prompt-injection instructions in email before delivery, protecting people as well as Copilot, agents, and other AI systems that may read mail. This is a material expansion of the email-security threat model.

Traditional email controls are designed around malicious links, attachments, spoofed identities, malware, and social engineering intended to manipulate a human recipient. A mailbox-connected agent introduces another target: text that attempts to redirect an AI system’s instructions, extract sensitive context, or cause an automated workflow to take an unsafe action. Filtering those instructions before they enter the mailbox is preferable to relying on an agent to recognize adversarial content after it has ingested it.

Microsoft’s benchmark does not quantify how often those prompt-injection detections occur, what patterns trigger them, or whether the capability is available across every Defender for Office 365 licensing tier and deployment model. Those omissions matter for organizations planning to connect agents to shared mailboxes, service inboxes, or executive assistants’ workflows.

For now, the defensible takeaway is straightforward. Microsoft’s new report supports Defender for Office 365 as a substantial part of a modern Microsoft 365 mail-security stack, and it documents the continuing value of post-delivery cleanup. It does not prove that every organization can retire an existing gateway or ICES product. The deciding evidence remains the threat and false-positive data in the tenant that will bear the risk.