Skip to main content
New 2026 Ransomware Report: Why Every Year Becomes the Worst Year on RecordRead the Report
BlackKite: Home
Menu

The AI Scanner Hype Test

Third Party Podcast: AI scanners find flaws faster than anyone can fix them. What that changes for TPCRM.

YouTube video thumbnail

In this article

Check out our podcast, Third-Party. This is the podcast built for the people behind the dashboards. The ones managing 5,000 vendors with a team of three.

WATCH ON YOUTUBE

Found, Not Fixed

More than 10,000 high and critical vulnerabilities surfaced in a matter of weeks. Patching them is now the bottleneck.

That comes from Anthropic's May 2026 Project Glasswing update, which states it plainly: progress used to be limited by how quickly vulnerabilities could be found, and it is now limited by how quickly they can be verified, disclosed and patched. Their own open-source scan illustrates the squeeze. It surfaced 6,202 high and critical flaws, 530 have reached maintainers, and 75 are patched. The finding engine works. The ecosystem is straining to absorb what it produces.

In the latest episode of the Third Party podcast, Jeffrey Wheatman, Bob Maley, and Ferhat Dikbiyik put the new class of AI scanners through a hype test, and the verdict is neither "game changer" nor "marketing nonsense." It is considerably more useful than either.

You Were Never Going to Patch Everything

AI did not break vulnerability management. It removed your ability to pretend.

48,100 CVEs were published in 2025. Projections for this year land somewhere between 70,000 and 90,000. Anyone whose risk mitigation approach is "patch everything" was already living in a world where that was never possible. The AI scanners just made the delusion visible.

Picture rowing upstream with a waterfall behind you. Paddle harder and you still lose, because the current is stronger than you are. The only way out is to change your angle and reach the bank. Prioritization is that angle. It is not a compromise. It is the only strategy the arithmetic permits.

In a third-party context, that angle has three components:

  • Exposure. Which of your vendors actually run the affected software? Without ecosystem visibility, everything downstream is guesswork.
  • Susceptibility. Which of those vendors are realistically exploitable, not just technically affected?
  • Criticality. Which ones, if they went down tomorrow, would interrupt your business?

Continuous cyber risk intelligence exists to answer those three questions at ecosystem scale. No AI model, however capable, solves the problem of knowing which of your hundreds of vendors are exposed right now.

The Bottleneck Was Never Discovery

The industry keeps optimizing the wrong end of the pipeline.

Every argument about AI scanners is an argument about remediation throughput. But organizations do not skip patching production systems because they found the flaw too slowly. They skip it because patching production breaks things, and breaking things shuts the business down. That cost is fixed, it is hostile, and no amount of faster discovery touches it.

Four decades of software engineering research point the same direction: the earlier in the development lifecycle you fix a defect, the cheaper it is. The finding still holds, but he industry aims its most impressive tooling at the single most expensive moment available. Point the same models at code while it's still being written and the flaw never becomes a vendor's production vulnerability, or your third-party exposure.

The honest counterargument is money. Running frontier models across a large codebase is expensive, and discovery is only the first invoice. Verification and remediation are separate bills, and human triage is the line item nobody budgets for. Frontier vendors do pitch pre-production scanning, and the sheer volume they return is itself evidence that shifting left does not close the gap on its own. Even before release, you triage. The tiebreaker is the ship date.

Why Your Remediation Ceiling Might Be Enough

Here is the reframe that makes a small patch count look less like failure.

Of the 48,000-plus CVEs published in 2025, Black Kite research identified 58 that represented a genuine, discoverable, and exploitable threat to enterprise supply chains. Fifty-eight. Read the deluge against that number and the picture inverts. Most of the volume you cannot patch was never going to reach you. The gap is not a capacity problem. It is a targeting problem. Volume is not the win. Precision is. That is the argument running through both the vulnerability deluge analysis and Black Kite's 2026 supply chain vulnerability research.

Which makes one definition the most consequential decision in your program: what counts as bad. Historically, bad has meant technically severe. Bad is business impact. Programs that make that switch get to use their 1% well. Programs that do not will spend it on whatever the scanner sorted to the top.

How to Tell a Real AI Scanner From a Repaint

Every scanner on the expo floor calls itself AI-powered. Some actually are. Most are not.

Buyers reach for explainability as the filter, and as usually stated it is the wrong ask. Nobody can fully explain how these models work, including the labs that build them. They are non-deterministic. Identical inputs produce different outputs. A model can test perfectly, ship clean, and do something unaccountable a week into production.

So stop asking how the model works. Ask how it is governed. That question a vendor can actually answer, and the answers sort the field fast:

  1. Do you monitor for model drift, and how? If the response is "what's drift," the evaluation is over.
  2. What can you do that you could not do before?
  3. What do you do better, and what can you do that your competitors cannot?
  4. Why is AI in this product at all? Could a basic if-else statement accomplish the same thing?

One answer is disqualifying on its own. Any vendor promising to automate everything and reduce your headcount is advertising the absence of a human in the loop, which is a primary AI governance control rather than an optional feature. If that is missing from the marketing, there is no reason to take the meeting.

None of this makes the analyst redundant. Expect 80% to 90% of discovery, remediation, and publication to become autonomous. What survives is the research that generates the prioritization metrics in the first place, plus a human deciding when to escalate. The role does not disappear. It moves up a level.

Don't Miss an Episode!

Subscribe to Third Party on YouTube, the podcast for people who don't need to ask ChatGPT what TPCRM means. New episodes every other week.

Next time on Third Party:

Next time we get into the gap between what vendors tell you and what they actually do. Trust but verify is the cliché. We're talking about what verification looks like now that the data has finally caught up with the marketing.

Subscribe below.

Real Talk on Third-Party Risk.

Check out our new podcast, Third Party, where we unpack what actually works (and what doesn't) in TPRM.

Apple Podcasts
Follow Third Party on Apple Podcasts
Follow
Spotify
Follow Third Party on Spotify
Follow

Ready to get started?

Integrate risk intelligence into every part of your workflow so you can make more informed decisions with confidence.