Skip to main content
14 speakers. 7 enterprise risk teams. One day in Charlotte, Oct. 22.Save My Seat
BlackKite: Home
Menu
blog

2,300 Mythos Vulnerabilities Disclosed. 421 Patched. That Gap Is in Your Supply Chain.

Published

Sep 25, 2026

Introduction

Since February, Anthropic has surfaced 26,153 candidate vulnerabilities in open-source software using Claude Mythos Preview and other Claude models ( the vulnerability deluge). External security firms reviewed 5,008 of those and confirmed 4,576 as real bugs. Roughly 2,300 have been formally disclosed to maintainers across 392 open-source projects. 421 have been patched by software vendors.

That is the story. Not the discovery number, which everyone quoted in April, but the remediation number, which almost nobody has looked at since, and which is now the most important input to your TPCRM program. And importantly, the numbers mean nothing to non tech audiences. It’s just noise.

That gap is what John Heuer of Proviniti, Johnathan Bald of Black Kite and I spent forty minutes on in the second installment of our Mythos series. Here's the briefing.

Mythos Made Finding Vulnerabilities Easy. Nobody Made Fixing Them Faster.

From Anthropic's May update on Project Glasswing: "Progress on software security used to be limited by how quickly we could find new vulnerabilities. Now it's limited by how quickly we can verify, disclose, and patch the large numbers of vulnerabilities found by AI."

The research is interesting, but this is a supply chain problem. The constraint used to sit with the people looking for flaws. It now sits with the people who have to fix them. Those are not the same people, and in open source they are frequently one (or none) unpaid person with a day job.

When Anthropic announced Project Glasswing on April 7, roughly 50 partners had early access. Those partners have since reported more than 10,000 high- or critical-severity vulnerabilities found in their own code. The members of the launch group run some of the largest engineering teams in the industry, and companies at that scale have the budget, the staff, the expertise, and the release pipeline to act on what a model hands them.

Your mid-tier vendors have none of that, and they are running the same libraries.

What Five Months of Mythos Disclosure Data Says About Third-Party Cyber Risk

What happened

Anthropic's coordinated disclosure dashboard is public and updating. As of late August it showed a 91.4% true positive rate on external review, 462 identifiers assigned across 177 CVEs and 285 GitHub Security Advisories, and 1,815 reports acknowledged by maintainers.

They patched 421.

Why it matters

Acknowledgment is not remediation. A maintainer writing back to say "yes, that's a real bug in my parser" moves the vulnerability from unknown to known and confirmed, which is precisely the state an attacker can use and you cannot.

Anthropic says this themselves on the dashboard, and it is the most useful sentence on the page: the true positive count "should only be taken as one proxy for impact. Another, more reliable one is the number of patches created." The organization doing the finding is telling you to measure the fixing. The 2026 Verizon DBIR already put median time to patch at 43 days, up from 32 the year before, with only 26% of CISA KEV defects fully remediated, down from 38%. Set that against Mandiant's M-Trends 2026 finding that mean time to exploit has dropped to negative seven days, meaning exploitation routinely happens before a patch is even released, and the arithmetic stops being subtle.

Our own 2026 Third-Party Breach Report found the median vendor takes 73 days to disclose a breach, with an average of 117. You will hear about the patch your vendor did not apply roughly four months after it stopped mattering.

What's next

More volume, and not evenly distributed. NVD data shows 45,626 CVEs published through July, and 2026 passed all of 2025 during August. FIRST's 2026 forecast sits at a median near 59,000 and calls 70,000 to 100,000 entirely possible. Every one of those lands on a queue that is already 421-for-2,300.

Most of Those Vulnerabilities Are Pointless. A Few Will Ruin Your Quarter.

Volume is not the threat. Separating the risk from the background noise is the real threat.

Plenty of what AI discovery surfaces may be theoretical. It requires physical access, or a specific build, or a chain of preconditions nobody will ever assemble. If you have lost physical access control, an exotic kernel bug is not your most pressing concern.

Bad actors act like water. They take the easiest path, and the easiest path is almost never the clever new vulnerability in an obscure package. Of the 48,000-plus CVEs published in 2025, the Black Kite Research Group™ manually analyzed 1,240 as high-priority for third-party risk. 329 of those proved highly discoverable through OSINT and picked up FocusTags®. Just 58 carry an EPSS score above 60%, the set the team calls Code Red.

Fifty-eight. That is a number a team can work, and layering your own vendor exposure on top cuts it further still.

That filtering is the real discipline of TPCRM, and it is the part no frontier model does for you. Mythos will tell you a flaw exists. It will not tell you which of your suppliers is running it, whether that supplier sits inside your payment flow, or how many of your other vendors depend on the same library underneath.

I watched this play out at DEF CON this year. Across eight sessions I picked more or less at random, not one speaker raised AI. They were breaking BGP, Windows plug and play, and capturing access tokens over long-range radio. Protocols that we have been running for decades, some of them since before I started in this field, and I’ve been doing this so long, it wasn’t even called cybersecurity back then. AI will make those attacks faster. It will not make them new. We could not patch this stuff when discovery was slow, and discovery is no longer slow.

The Two Vendors You Cannot Fix

Every remediation strategy assumes the vendor is both willing and able. A lot of them are neither.

There are two failure modes here, and they sit at opposite ends of your vendor list.

  • The vendor is too small to act. Mid-market suppliers and open-source maintainers cannot buy frontier model access, and many have no security response process at all. The 421 number is what happens when you send a valid, confirmed, high-severity finding to someone with no one on the other end to receive it. This is the long tail my colleague Ferhat Dikbiyik has spent two years mapping, where more than 36% of discoverable risk now lives, and it is the population you are not assessing, because they never made your critical tier, or anyone elses.
  • The vendor is too big to care. Good luck calling Google or AWS to tell them they are not patching fast enough. Your leverage is a rounding error on their revenue, and no amount of escalation changes that. What you get instead is the right to know your exposure and plan around it.

Jaguar Land Rover is the case study nobody wants to be. A cyber attack in late August 2025 halted UK production for roughly five weeks, and the Bank of England attributed part of a 0.1% monthly contraction in UK GDP to it. The Cyber Monitoring Centre put the cost to the UK economy at £1.9 billion across more than 5,000 organisations. One company, one incident, measurable in national accounts. That is cascading risk with a number attached, and cascading failure rarely starts with the organization that absorbs the damage.

What to Tell Your Executives

Three things, and none of them require explaining what a CVE is.

  • Discovery is solved and it was never our constraint. The industry can now find flaws faster than anyone can fix them. Our exposure is not a function of what gets found. It is a function of whose queue it lands in, and most of those queues belong to companies we do not control.
  • The gap is structural, not a performance problem. Our team is not missing things because they are not good enough. They are missing things because assessment cycles measured in years cannot track a disclosure pipeline measured in days.
  • We need prioritization, not more collection. Fifty-eight out of 48,000 is a workable program. Thirty thousand questionnaire responses is not, and if all 30,000 vendors answered tomorrow, we would have a bigger problem than we started with.

Then bring a number, and make sure it is the right number. Not what the tooling costs, but what the exposure costs. Quantify what a compromise at one of your load-bearing vendors would do to revenue, and bring it as a defensible range rather than a figure to the penny.

The Bottom Line

The deluge did not arrive as a flood. It arrived as a backlog.

Nothing about this is apocalyptic, and I would not sell it to you that way. The models are real, imperfect, and now available to defenders and attackers on the same terms. What changed is not the nature of the risk. It is that the interval between "this flaw exists" and "somebody is using it" is now shorter than the interval between "somebody told my vendor" and "my vendor fixed it."

Closing that gap is not a patching problem. It is a visibility and prioritization problem, and it runs through vendors you do not control and mostly are not watching. That means knowing which of your vendors run the affected software before the disclosure lands, using cyber risk intelligence rather than a letter grade, sending structured evidence through vendor engagement instead of a 500-question survey, and mapping the Nth-party dependencies that put the same library inside forty of your suppliers at once. That last one is concentration risk, and it is the single exposure that gets worse as discovery gets better.

Discovery is working. The patch pipeline is not, and the difference between those two facts is where your next third-party incident is already sitting.

John Heuer had a lot more to say about where AI governance breaks down inside a TPCRM program, and it did not make it into this post. Watch the replay.