HackerOne's 29x Critical Backlog Sits Beside Its Own MTTR Improvement
The two headline numbers first appeared in HackerOne's September 21 press release. According to H1 Platform data, critical vulnerability mean time to remediate (MTTR) improved by more than 50% over the past 12 months. Over the same period, the unresolved critical backlog grew 29x. Security Boulevard carried the same figures and added Sprague's framing. It reported her view that the number of vulnerabilities being found in source code has increased well beyond the ability of the average organization to validate and remediate.
Security Boulevard's report is independent coverage, but it relies on HackerOne's data rather than any separate measurement. The honest way to read these figures is as one large vendor's view of its own customers. They are not an industry-wide census. HackerOne's published material doesn't give absolute backlog counts or a full methodology, so "29x" tells you the direction and rough size of the change, not how many open criticals a typical organization has.
HackerOne's earlier reports point the same way with different numbers. In a May 6 analysis of the 12 months ending March 2026, Chief Product Officer Nidhi Aggarwal and VP of Product Strategy Sandeep Singh wrote that the resolution rate for critical issues fell from over 83% to under 40%. This gap between discovery and resolution caused the backlog of unresolved critical vulnerabilities to grow by about 25x. The 25x and 29x figures come from different publications covering different reporting windows. They shouldn't be averaged or combined. What they share is the trend: two consecutive HackerOne snapshots show critical backlogs growing by more than an order of magnitude.
Why Faster MTTR and a Growing Vulnerability Backlog Can Both Be True
It looks contradictory to fix things twice as fast while falling 29 times further behind. It isn't, because the two metrics measure different things. MTTR only counts vulnerabilities that were actually resolved and measures how long each one took. Backlog counts what is still open. That number depends on how many validated findings arrive compared with how many get closed, not on how fast any single fix happens.
HackerOne's May analysis puts the gap in plain terms. Over the twelve-month period, mean time to remediate (MTTR) across the platform dropped by about 80%, and median MTTR fell by over 70%. Yet the total number of vulnerabilities resolved each month fell by about 46% over the same period, even as overall submissions grew by 76%. The cumulative backlog of validated but unresolved vulnerabilities grew by more than 21x. The authors' explanation is that teams fixed individual issues faster, possibly by being more selective, while total remediation effort did not grow to match discovery.
Here's the arithmetic, which is our explanation rather than a HackerOne formula. At the end of any period, the backlog equals what was open before, plus newly validated findings, minus findings resolved. If validated findings rise sharply and resolutions stay flat or fall, the backlog grows no matter how quickly the resolved items were handled. A team can fix the easy critical in three days and leave the hard one open for three months. MTTR will look great, and the exposure is still there.
Sprague makes the same point in her TechRadar piece. She argues that faster repairs can coexist with a growing backlog when discovery accelerates faster than remediation and engineering capacity doesn't keep up. Any security dashboard that reports MTTR alone is measuring how fast work moves, not how much risk is left.
AI-Assisted Bug Reports Make Validation the Bottleneck
Sprague's second argument is that validation, meaning confirming that a reported flaw is real, in scope and exploitable, is now the choke point. In the TechRadar piece she writes that AI has made convincing security reports cheap to produce. Some point to real weaknesses. Others duplicate known issues, misread the target, or describe theoretical risks that barely matter. A report that took seconds to generate can still cost an experienced analyst hours before it can be safely dismissed. At enterprise scale, she says, well-evidenced findings end up buried among hundreds of plausible-sounding reports that lead nowhere.
HackerOne's own data complicates the "AI slop" framing, and readers should keep that in mind. In an April post, co-founder and CTO Alex Rice reported that March 2026 set a record of 46,947 submissions, up 76% year over year. He said the share validated as exploitable held steady at about 25%. The May analysis reached the same conclusion: signal rates stayed relatively consistent throughout this period. This stability indicates that the increase in volume represents a mix of valid and invalid vulnerabilities rather than an influx of AI-generated noise reports.
Both things can be true, and together they make the case stronger. If a quarter of a much larger stream is valid, there are many more real vulnerabilities to fix, and also many more invalid reports that still take analyst time to reject. Rice added that critical and high-severity findings rose to 32% of validated reports, up from a historical 26–28%. On his account, the extra volume is more severe, not just bigger. He also said HackerOne now uses its own agentic AI system, Hai, for initial classification, deduplication and validation alongside human analysts, and that it couldn't process current volumes without it.
Rice's post also cites curl maintainer Daniel Stenberg as an example. Stenberg complained publicly about low-quality AI-generated reports, then in April 2026 described a flow of high-quality, mostly AI-assisted reports. Rice presents this as evidence that skilled researchers using better tools produce more signal, not more noise.
Sprague's practical fix is stricter evidence standards. Programs should expect researchers to show likely business impact and demonstrate how a vulnerability reproduces, using automation to improve evidence rather than increase volume. She also argues that a researcher's record of valid findings is a useful triage signal that grows more valuable as submissions rise. She attaches an obligation to that: researchers who meet higher standards should get fast, fair triage, real recourse when a valid report is wrongly dismissed, and continued access for newcomers without a reputation yet.
Exploit Chains and Business Context Keep Humans in the CTEM Loop
Sprague ties the argument to Continuous Threat Exposure Management (CTEM). She describes it as a continuous process for understanding the attack surface, finding weaknesses, proving which are exploitable, and sending remediation to the exposures that carry the most business risk. It's an operating model, not a product category. Organizations can adopt it with tools from any vendor.
In her account, the prioritization step is where AI falls short. A model can match findings to known patterns, but it rarely knows which services generate revenue, where regulated data is stored, which dependencies make downtime especially costly, or which compensating controls already exist. She calls the harder problem combination: several moderate-looking findings can form a serious attack path once someone understands how systems interact, which permissions can be abused, and where controls fail across organizational boundaries.
HackerOne's product material takes the same position. It says humans remain essential for complex, context-dependent weaknesses like business logic flaws and multi-step exploit chains. The goal is a hybrid model that scales without losing trust. Security Boulevard reported that Sprague expects organizations to rely more on third-party researchers supported by AI tools to confirm whether a vulnerability is actually reachable in their IT environment as part of what she called a "deep cleaning" process.
The one number in the op-ed that appears nowhere else in HackerOne's published material is Sprague's claim that researchers earned more than $47 million on the H1 Platform in the first half of 2026, up more than 25% year over year. She adds her own caveat: rising total payouts don't mean every researcher is doing well. In her view, AI is absorbing routine, high-frequency findings first. That hurts researchers who relied on them and rewards those who can chain weaknesses, reason about business logic and produce credible proof. That's her analysis of where the market is heading. The payout total alone doesn't show it.
HackerOne's Claude Mythos Integration Is the Commercial Backdrop
The TechRadar essay is a vendor argument, and readers should weigh it that way. Two days before it ran, HackerOne announced the upcoming integration of Anthropic's Claude Mythos across select H1 Platform products: H1 Code Security Audit and H1 Code. The integration will make frontier cyber AI available in existing security workflows, connecting agentic vulnerability discovery with validation, prioritization, and remediation. Per trade coverage of the announcement, H1 Code Security Audit will leverage Claude Mythos to thoroughly scan code repositories, bringing to light complex vulnerabilities that may include extended compositional vulnerabilities across years of code commits. Furthermore, H1 Code will enhance the review of pull requests.
This is a planned release, not a shipping feature. HackerOne says only that the select H1 Platform offerings will soon be available with Mythos, with no general-availability date. Aggarwal said in the release that HackerOne ran and validated its "harness" on its own systems before any customer saw it. That's the company's statement, not a published independent evaluation. The release also says HackerOne does not use confidential researcher submissions or customer vulnerability data to train, fine-tune or otherwise improve generative AI models. That's a policy commitment enterprises may want written into their contracts.
The commercial angle doesn't make the analysis wrong. The MTTR-versus-backlog gap is simple arithmetic that any organization can check against its own ticket data. What the vendor framing does is steer readers toward buying the answer. The underlying practices, which are evidence standards, separate validation tracking, named ownership and retesting, work the same whether a team uses the H1 Platform, another bug bounty provider, or an internal vulnerability management process.
Board Metrics Should Measure Open Exposure, Not Finding Counts
Sprague's final recommendation is about reporting. She argues that boards need measures of risk reduction rather than activity. Finding counts are easy to report and can go up even as an organization gets safer. In her words, confirmed exploitability, remediation speed, recurrence, and the size of the unresolved critical backlog tell the real story.
The May analysis gave a similar prescription: MTTR measures how fast teams fix things, not how much risk they remove, so it should be paired with resolution rate and exposure backlog. It also argued for dedicated remediation capacity and sprints, plus AI on the fix side: AI-assisted fix generation, automated regression testing and agentic workflows. Rice's April post added fixing whole classes of bugs instead of individual findings. When AI finds the same category of flaw across dozens of endpoints, a framework-level fix beats fifty separate patches.
HackerOne has put some of this into its own analytics. Its August 2026 changelog added Days to zero backlog, our first forward-looking metric in Analytics. It turns 90 days of discovery and remediation rates into a timeline to clear the current backlog, with a separate figure for High and Critical. The same changelog says HackerOne corrected MTTR to count valid reports only, and updated the Executive dashboard to match. Customers comparing HackerOne MTTR figures across periods should keep that definition change in mind. Any team can calculate the days-to-zero idea in a spreadsheet: take the open backlog and divide it by the net rate at which it shrinks. If that rate is negative, the backlog never clears.
Sprague also warns about verification at the end of the process. When an AI system proposes a fix, it may carry the same blind spot into judging whether that fix works. Independent retesting is the check that confirms a finding is actually closed.
What this means for you
Security leaders, and sysadmins who manage patching and vulnerability intake, should check whether their current metrics can tell faster fixing apart from less exposure. If your monthly report leads with MTTR or finding counts, add the unresolved validated backlog and its age before you add any more discovery tooling. Teams running bug bounty or vulnerability disclosure programs should expect more submissions and set evidence requirements now, not after the queue fills up. Teams that don't run external programs still face the same pressure from internal AI code scanners and pull-request reviewers.
- Track validated, unresolved critical and high findings by count and age alongside MTTR, because MTTR only measures issues that were actually closed.
- Measure time to validate separately from time to remediate so you can see whether triage or engineering is the real bottleneck.
- Require reproduction steps and a stated business impact for incoming reports, and treat a researcher's history of valid findings as a triage signal.
- Assign a named engineering owner to every confirmed finding so validated risks don't sit between the security and development teams.
- Retest every fix independently, especially fixes proposed by AI tools, before marking a vulnerability closed.
- Treat HackerOne's Claude Mythos integration for H1 Code Security Audit and H1 Code as announced but not yet available, since no general-availability date has been published.
Discovery was always the hard, expensive part of vulnerability management, and the data Sprague cites shows that has changed, at least on the platform HackerOne can see. When real, exploitable findings are this cheap to surface, security programs will be judged by how many critical findings they close and verify each month and how fast the open backlog shrinks. The next milestone is the general availability of HackerOne's Mythos-powered tools. By its own account, those tools will add more findings, which puts even more weight on the validation, prioritization and remediation steps that come after discovery.