Back

Episode 41: How to Evaluate AI Claims From Security Vendors

TL;DR

Every MSP security vendor now claims some kind of AI feature, and most buyers never get past the claim to check what's actually behind it. This episode walks through a practical framework for testing vendor AI claims: the numbers to ask for instead of the demo, the four signals that separate a real feature from a roadmap item dressed up as one, and why accountability for a wrong decision matters more than the AI story itself.

Episode notes

Adam

Okay, here's a game. Name a security vendor right now that isn't claiming some kind of AI feature.

Christina

That's genuinely hard. I don't think I can.

Adam

Right, because basically all of them are. Every single one has an AI story now, and a lot of them started telling that story within the same few months of each other.

Christina

Which is exactly the problem, isn't it. When everyone says the same thing at the same time, you can't tell who actually built something and who just updated their homepage.

Adam

Exactly. And to be clear, the issue isn't that vendors are lying, necessarily. It's that the claim and the actual feature are two completely different things, and most buyers never get past the claim to check.

Christina

There's actually a term for this now, isn't there. People call it AI washing, the security version of greenwashing, where the label gets slapped on something regardless of what's actually happening underneath it.

Adam

That's exactly the right comparison. And just like greenwashing, it isn't always dishonest on purpose. Sometimes a vendor genuinely believes a basic rules engine or a bit of statistical scoring counts as AI, because the bar for what counts as AI keeps shifting. That doesn't make the claim any more useful to the buyer trying to evaluate it though.

Christina

So let's fix that today. If somebody's evaluating a vendor's AI claim, what should they actually be measuring, instead of just nodding along at a demo?

Adam

There are basically three things that applied AI should show up as, in a security operations context. Faster triage, so somebody working through alerts is getting through them quicker than a person doing it alone. Fewer false positives reaching a human's queue. And measurable analyst time saved per incident.

Christina

Not saved in general, but per incident specifically.

Adam

Right, because "saves time" is vague enough to mean almost nothing. If a vendor can't point to a real number in at least one of those three places, what they're describing is probably a roadmap item, not something that's actually shipped and working today.

Christina

Okay, so beyond just asking for a number, what are the signals that tell you the difference between something real and something that's mostly packaging?

Adam

Four things I'd look at. First, roadmap specifics. Is there an actual date and an actual scope, or is it just "coming soon"?

Christina

Because "coming soon" can mean anything from next month to never.

Adam

Exactly. And a good way to press on this in a sales conversation is to ask what specifically ships in the next release, not the next year. A vendor with a real roadmap can answer that in one sentence. A vendor without one usually starts talking about vision instead of dates.

Christina

That's a useful tell. What's the second signal?

Adam

Second, integration depth. Does the AI feature actually reach every surface it needs data from to be useful, or is it only working off the one data source it happened to be built on first?

Christina

So if it's an endpoint tool, is the AI feature only looking at endpoint data, when the thing it's trying to solve really needs network or identity data too?

Adam

That's exactly the kind of gap to watch for. A feature that only sees one layer of the environment can only ever be as smart as that one layer, no matter how good the underlying model is.

Christina

What does a good answer to that actually sound like, versus a bad one? Because I imagine most vendors will just say yes when you ask if it covers everything.

Adam

A good answer is specific. Something like, it currently correlates endpoint and identity data, network visibility ships next quarter, and here's exactly what that gap means for you in the meantime. A bad answer is a flat yes with no detail, or worse, a pivot straight into a different feature entirely because they know the honest answer is thinner than the pitch.

Christina

So the dodge itself is basically the tell.

Adam

Every time. The vendors worth working with are usually the most upfront about what their AI doesn't cover yet, because they know the buyer is going to find the gap eventually anyway.

Christina

Which brings us to the third one, I'm guessing.

Adam

Third one, telemetry coverage generally. AI is only ever as good as what it can actually see. So the real question to ask a vendor isn't what their AI can do, it's what their AI cannot see.

Christina

That's a great reframe. Ask about the blind spot, not the feature.

Adam

And fourth, governance and audit. Can the vendor actually show you, after the fact, what the model decided and why? If they can't produce that, you're trusting a black box, and you have no way to check it later.

Christina

And that governance piece matters for more than just curiosity, doesn't it. If a client ever asks why a particular alert got closed automatically, or an insurer or an auditor asks the same question during a renewal, "the AI decided" isn't an answer anyone can actually use.

Adam

Right, you need something you can point to. A specific reason, tied to a specific decision, that you can retrieve later. Without that, every automated decision is basically unaccountable the moment it happens.

Adam

There's a pricing angle worth flagging here too, honestly. Some vendors bundle AI capability into the base product, and some sell it as a paid add-on layered on top. Neither approach is automatically wrong, but you want to know which one you're buying before you sign, because a feature you're paying extra for deserves a higher bar of proof than something included by default.

Christina

So ask the pricing question in the same breath as the capability question.

Adam

Exactly, otherwise you end up finding out at renewal time that the thing you thought was core to the platform was actually a bolt-on the whole time.

Christina

Okay, so let's turn this into something practical. If I'm sitting across from a vendor right now, what do I actually ask them, out loud, in the room?

Adam

Five questions. One, what decision does this feature make without a human involved at all, versus what does it only suggest for a human to approve?

Christina

Because there's a huge difference between "it flags this for you" and "it acts on this by itself."

Adam

Huge difference. Two, show me the false positive rate before this feature existed and after, on the same data. Not two different environments, the same one.

Christina

So they can't cherry-pick a flattering comparison.

Adam

Right. Three, what happens when the model gets it wrong? Who sees that, and how quickly?

Christina

Because every model gets things wrong sometimes. The question is what happens next.

Adam

Exactly. Four, can I go back and audit a specific decision six months from now? Not just today, but later, when it actually matters for a client conversation or a compliance question.

Christina

And the fifth?

Adam

Does this run across my whole environment, or only a subset of it? Because a lot of these features get demoed on a clean, narrow slice of data that makes them look more capable than they are in a messy real environment.

Christina

Okay, so if I ask all five of those and a vendor gets visibly uncomfortable, that's probably useful information in itself.

Adam

Very useful information. A vendor with a real feature will usually enjoy answering those questions, honestly, because they've thought about all of it already.

Christina

What about during an actual trial or a proof of concept, rather than just a sales call? Is there anything specific worth watching for once you're actually running the thing?

Adam

Good question. I'd say track the same three metrics we opened with, triage speed, false positive rate, and analyst time per incident, on your own real data rather than a demo environment, for at least a few weeks. Ask for those numbers in writing before you sign anything longer term, not just as a verbal promise during the pitch.

Christina

So the trial period is really where the roadmap claims either hold up or fall apart.

Adam

That's exactly it. A demo can be staged. A few weeks of your own alert volume is much harder to fake.

Christina

So does this mean AI is replacing security analysts? Because that's the fear a lot of people jump to.

Adam

No, and I think that's actually the most important point in all of this. It changes what the analyst spends their time on. The tools that hold up under real scrutiny use AI to handle volume, the repetitive first-pass work, and then route the genuine judgment calls to an actual person. They don't try to remove the person.

Christina

Which, honestly, is the model over at enhanced.io as well. Automation does the first pass, and then a named security director, an actual person assigned to that partner, interprets what comes out the other side and owns the outcome.

Adam

Right, that's really the standard worth holding every vendor to. If a vendor can't tell you clearly who owns the outcome when their automation gets something wrong, that's worth asking again, more directly, until you get a real answer.

Christina

Great place to end on. Ask for the number, not the story, and ask who's accountable when the number's wrong.

Christina

One last thought before we close. None of this means an MSP should be anti-AI, or treat every claim with suspicion by default. It just means the burden of proof sits with the vendor, not the buyer. If the feature is real, showing the numbers should be easy for them, not a favor they're doing you.

Adam

That's the whole test, really. Same five questions, every vendor, no exceptions, and you'll know a lot faster who actually built something.