Cybersecurity analysts monitor a dashboard showing blocked attacks, endpoint security, and threat activity.
AV-Comparatives’ 2026 Endpoint Prevention and Response test found that 11 of 14 enterprise endpoint products met its new Certified Leader bar, but the result is more useful as a shortlist of tested Windows endpoint configurations than as a universal vendor ranking. Microsoft Defender for Endpoint was not among the 14 products listed in the report, so Windows administrators cannot use this particular comparison to judge Microsoft’s platform against the certified products.

The test, published September 15, covered 50 multi-stage intrusion scenarios and measured whether a product stopped an attack automatically or surfaced an actionable alert for a security team. AV-Comparatives says the scenarios included phishing-style entry, lateral movement, data exfiltration, and abuse of legitimate tools, with techniques mapped to MITRE ATT&CK for Enterprise.

PA Media’s distributed release correctly reports the headline figure — 11 Certified Leaders — but its claim that testing began in May does not match AV-Comparatives’ own comparative report. The primary report says testing ran from June through August 2026, with publication in September. For buyers treating the announcement as a fresh measure of current product behavior, the lab’s report is the record that counts.

The 11 certified products and the configurations that passed​

The Certified Leader list is Bitdefender, Broadcom, Check Point, Elastic, ESET, Fortinet, G Data, Palo Alto Networks, TrendAI, VIPRE, and WithSecure. Those names matter, but the exact product and version matter more: AV-Comparatives tested specific enterprise editions rather than every security product each vendor sells.

The report identifies the tested builds as Bitdefender GravityZone Business Security Enterprise 8.26; Broadcom Symantec Endpoint Security Complete 16.0; Check Point Endpoint Security Advanced 89.10; Elastic Security 9.4; ESET PROTECT Enterprise Cloud 7.1; Fortinet FortiEndpoint with EDR Essentials 6.2; G Data XDR 1.0; Palo Alto Networks Cortex XDR Prevent 9.1; TrendAI Vision One Endpoint Security Essentials 14.0; VIPRE Endpoint Detection & Response 13.4; and WithSecure Elements XDR 26.2.

That is an important limitation for procurement teams. A result for Cortex XDR Prevent does not automatically describe Cortex XDR tiers with different prevention policies, response modules, retention periods, or managed-service arrangements. The same applies to lower-cost endpoint packages from any of the listed vendors. AV-Comparatives explicitly says a different tier from the same vendor may have a different feature set.

The three non-certified entrants are labeled Vendor A, Vendor B, and Vendor C. AV-Comparatives says they chose anonymity, meaning readers cannot determine whether a platform was absent, declined publication, or failed the threshold. The report does provide their anonymous scores and modeled costs, but withholding the product names prevents a buyer from using the test as a complete market map.


A 92 percent threshold is not the same as a 92 percent block rate​

AV-Comparatives increased the certification requirement from 90 percent to 92 percent. The figure is a combined score across active response and passive response, not a simple percentage of attacks stopped cold.

Active response is automatic prevention with reporting. Passive response means the product did not block the step but generated a detection that AV-Comparatives judged sufficiently tied to the attack for an administrator to act on it. In a real security operation, that distinction can decide whether an incident becomes an outage: detection only helps if the alert reaches a staffed team with enough context and authority to contain the affected device.

The test also weights the stage at which a product acts. An automatically reported block in Phase 1 — endpoint compromise and foothold — receives no modeled breach impact. An active response in the final phase, asset breach, is assigned 75 percent of the model’s full breach impact; passive response at that stage is assigned 95 percent. A product that neither prevents nor detects a scenario across all three phases receives 100 percent impact.

This explains why the report’s top-line certification should not be read as “11 products stopped 92 percent of attacks.” A tool can receive credit for a high-quality alert after an initial action has occurred. That is legitimate EDR behavior, but it puts more operational responsibility on the customer’s SOC, managed detection provider, or incident-response team.

AV-Comparatives further says a product is automatically disqualified if it suffers five full breaches, with testing stopped at that point. Yet the report says none of the 14 tested products recorded a full unknown breach in this year’s scenarios. The three unnamed products fell below certification for the combined score or operational-impact requirements, rather than because the test documented five complete undetected attack chains.

AI-assisted tooling raised the scenarios, not an AI-versus-AI contest​

The release frames the 2026 test as a response to AI-built offensive tooling. AV-Comparatives’ report supports the narrower version of that claim: the lab used AI-assisted development techniques to create testing tools and variations of scenarios. It does not establish that every scenario was generated by an autonomous agent, nor that the products were directly tested against the same kind of AI-orchestrated campaign Anthropic described in late 2025.

That distinction is worth preserving. Anthropic said its investigation into the China-linked GTG-1002 operation found Claude Code had been manipulated to support reconnaissance, vulnerability discovery, exploitation, credential harvesting, lateral movement, analysis, and exfiltration against roughly 30 organizations. Anthropic estimated AI performed 80 to 90 percent of tactical activity, though it also reported that the model sometimes overstated findings or fabricated results, requiring human validation.

AV-Comparatives’ scenarios draw on public threat intelligence and techniques associated with named state-linked and financially motivated groups, including APT28, APT29, APT41, Lazarus, LockBit, Black Basta, and FIN7. The lab says those scenarios are inspired by, rather than replicas of, those groups’ operations. Its 50 tests use artifacts and delivery forms familiar to Windows defenders: ISO, CPL, XLL, CHM, VBScript, batch files, malicious MSI packages, LNK files, HTA payloads, PowerShell, rundll32 abuse, and Metasploit or Meterpreter-based activity.

For administrators, the actionable conclusion is straightforward: “AI-assisted” should not become a checkbox separate from endpoint fundamentals. The relevant controls remain the ones that stop or expose scripting abuse, malicious archive and shortcut delivery, Office and browser entry points, credential access, remote execution, lateral movement, and data staging. The effectiveness of an endpoint agent still depends heavily on Windows hardening, identity controls, patching, network segmentation, logging, and whether somebody can respond when prevention fails.


The CyberRisk Quadrant has assumptions buyers should inspect​

AV-Comparatives’ CyberRisk Quadrant combines prevention and response results with a five-year operational-impact model for a hypothetical organization with 5,000 endpoints. The model includes list pricing, workflow delays, operational accuracy, and estimated breach impact.

This creates a more useful picture than raw prevention percentages alone, particularly where an endpoint tool blocks legitimate work or floods analysts with alerts. But the dollar figures in the report are not customer quotes. AV-Comparatives says it uses vendor-provided list prices and excludes reseller discounts, volume agreements, region-specific pricing, and negotiated enterprise contracts. It cautions that actual enterprise costs can vary significantly.

The report’s standardized five-year licensing estimates range from $375,000 for VIPRE to $3.125 million for Broadcom among the named products. Those numbers should be treated as model inputs, not as a price sheet. A Microsoft shop evaluating a migration, for example, would need to account for existing Microsoft 365 licensing, Defender entitlements, staff familiarity, SIEM costs, identity tooling, and any MDR contract — none of which this comparison can settle because Defender for Endpoint was not tested.

There is a second operational caveat. AV-Comparatives allowed vendors to specify configuration changes before testing, then applied those configurations with vendor engineers during setup. That reflects the reality that enterprise endpoint protection is commonly tuned, but it also means the result represents a vendor-approved deployment, not necessarily an out-of-box tenant with default policies.

The report documents several examples of non-default hardening, including Bitdefender’s increased blocking intensity and enabled firewall and intrusion prevention; Elastic’s aggressive malware threshold; and a range of enabled prevention, ransomware, network, EDR sensor, AMSI, and sandbox-related settings for ESET. Organizations comparing their installed estate to this test should verify whether their own policy baseline matches the tested configuration before claiming equivalent protection.

What Windows security teams should do with the result​

The 2026 EPR report is credible evidence that the 11 named products, in their tested editions and vendor-approved configurations, cleared a demanding lab threshold across simulated Windows attack chains. It is not evidence that one badge-holder is automatically the right choice for every Windows fleet, and it does not answer how an organization’s existing Microsoft security stack compares.

A practical review should begin with the gaps the test deliberately leaves outside its scope. AV-Comparatives says telemetry-based threat hunting is not included. It does not substitute for validating Microsoft Entra ID protections, mailbox defenses, privileged-access workflows, Windows event collection, SIEM correlation, backup recovery, or incident-response coverage. Those are the controls that determine whether a passive endpoint detection becomes containment before data leaves the network.

The report does provide a defensible reason to ask endpoint vendors sharper questions during a proof of concept: which tested product tier is being proposed, which prevention and EDR policies must be enabled, whether those settings increase false positives or workflow delay, and who owns response after an alert is generated. The 11 certifications narrow the field; the configuration record shows why a badge alone should not close the purchase order.