Skip to main content
Contextual risk review · Published verdict vocabulary

Every vendor tells you they use AI. Almost none will tell you where they did not.

ShadowMap AI does not add a suggestion to a row. It restructures the queue: every finding it touches gets a verdict from a fixed four-term vocabulary, a score that expresses priority, a confidence percentage that expresses certainty, and a note in plain language explaining the call. It routes findings. It cannot remove them. And there is a published list of surfaces where it does not run at all, because those surfaces answer a question of fact and a verdict on a fact is noise wearing a confidence score.

4
Published verdicts, routing into three queues
0–1,000
AI score for priority — a separate field from confidence
0
Findings the AI can delete

Where we deliberately did not use it

AI Review does not run on these, and that was the decision

In a market this saturated with AI-washing, the honest signal is not the list of places a vendor put a model. It is the list of places it chose not to, and the reasoning. The surfaces below are inventory and configuration rather than triage queues: they answer a question of fact, and a probabilistic verdict on a fact adds nothing an operator would act on.

AI Review does not run on these, and that was the decision
StateWhat it meansWhat follows
Code repository monitoring High volume, and the difference between a test fixture and a live production secret is contextual rather than mechanical. This is where a verdict earns its place — both customer examples at the foot of this page come from this surface. ✓ AI Review runs
Open Ports The partial case, and worth stating precisely rather than rounding to a yes. The model scores and tags an open port so the list ranks usefully, but it is never permitted to hide one. Whether a port ought to be open is your policy, and policy is not a thing to infer. ◑ Score and tags only — it ranks, it never hides
Single Sign-On An identity provider is exposed on a given path or it is not. The finding is a configuration fact, and there is no reading of the evidence for a model to add to it. ✕ No AI Review
SSL Certificates Issuer, validity window and chain are deterministic properties of the certificate. A model cannot be more certain about them than the certificate already is. ✕ No AI Review
Cloud IAM A permission is granted or it is not. Reading it correctly is a parsing problem rather than a judgement one, and an operator needs the exact grant rather than an opinion about it. ✕ No AI Review
Web Applications Findings here arrive with their own evidence and their own severity. A second verdict on a different scale would only give two numbers to argue about in the same review meeting. ✕ No AI Review
JavaScript Trackers The question is which third-party scripts execute on your pages. That is an inventory, and an inventory does not need triaging — it needs to be right. ✕ No AI Review
Links and Redirects A redirect chain resolves the way it resolves. What an operator needs is the chain itself, in full, rather than a verdict summarising it. ✕ No AI Review
Network Services A reachable service and its banner are observations, and the inventory of them is either right or wrong. Whether a given service ought to be reachable at all is your policy, and we are not going to guess at your policy. ✕ No AI Review
Vendor Risk Management The whole of it, deliberately. Third-party risk decisions carry contractual consequences and belong to a named person in your organisation. We surface the evidence and stay out of the verdict. ✕ No AI Review
Universal Search · Action Center · Automated Mitigation Deterministic by design. Search matches what you asked for, the Action Center reflects what was decided, and mitigation executes a rule you set. A probabilistic layer in any of the three would make the product less predictable rather than more capable. ✕ Rules, not a model
Key
  • AI Review runs here
  • Ranking only — it never hides a finding
  • No AI Review, by decision
  • Deterministic — rules, not a model

Definition

What the model writes, and what it is not allowed to do

That list is only worth anything if the mechanism it withholds is itself described, and most of the time the word AI covers a scoring model nobody will describe. So here is the whole of ours, in the fields it writes to the record.

A finding arrives from a surface into one correlated exposure model, and ShadowMap AI evaluates it in that context before writing four separate values: a verdict drawn from a vocabulary of four terms, a priority score between 0 and 1,000, a confidence percentage covering how certain it is of that verdict, and a plain-language note giving the reasoning. Four distinct fields rather than one blended number, and all four are visible to you on every finding they touch. This is also where a model earns its place over a rules engine. A rules engine checks fixed conditions; the model reads context — likely ownership, what a file actually contains, whether a matched key belongs to a production console or a test fixture — and then reorders the working queue around what it read. What it cannot do is narrow that queue permanently. Nothing it dispositions is ever deleted, a human ruling always overrides it, and on the published list of surfaces above it does not run at all.

The published vocabulary

Four verdicts, and the three queues they route into

Almost no vendor publishes the vocabulary its model writes, which is what makes a verdict impossible to audit. Ours is four terms long and every finding carries exactly one of them. The verdict decides which queue a finding sits in. It never decides whether you can see it.

Four verdicts, and the three queues they route into
StateWhat it meansWhat follows
Confirmed Exposure The evidence supports a real exposure and the model is confident in that reading. Where the finding is also critical, it moves straight into Investigating rather than waiting on an analyst to sort the queue first.
Needs AI Review The model reached a reading it is not certain enough to act on alone — ambiguous context, a surface it has thin signal on, or evidence that argues both ways. Into the Needs Review queue for a human. Uncertainty is published as its own verdict rather than buried inside a score.
Likely False Positive The pattern matched but the context argues against it — a test fixture, a sanctioned service, a system your own teams stood up on purpose. Filtered by AI, with the reasoning attached. An analyst pulls it back into the working queue in one action.
Benign Clear junk against this surface. No exposure, and nothing ambiguous about the reading. Filtered by AI — out of the working set, still in your view, still carrying everything the model thought about it.
Key
  • Investigating
  • Needs Review — a human decides
  • Filtered by AI, and reversible
  • Filtered by AI, still in view

Two fields, not one number

Priority and certainty are different questions, so they are different fields

Most tools blend the two into a single severity number, which is why a "high" finding can mean either "this matters a great deal" or "we are fairly sure this is something". In ShadowMap they are separate columns on the record and you can sort on either one.

Worked example

One finding, three fields the model writes

AI score — 0 to 1,000 — how much this matters
A priority ranking against everything else in your queue. A leaked key reaching a production console and a leaked key for a decommissioned staging box can both be genuine findings; the score is what separates them, and it says nothing about how certain the model is.
AI confidence — a percentage — how sure the model is
Certainty about its own verdict, and nothing else. A finding scored 900 at low confidence is precisely the thing an analyst should open first — and a blended severity number is exactly what would have buried it in the middle of the list.
AI note and tags — why it said so
A plain-language explanation of the call, plus tags drawn from a controlled vocabulary defined per surface rather than typed as free text. Controlled tags are what make triage countable: filter on one and the number means the same thing this quarter as it did last quarter.

The contract

ShadowMap AI never deletes anything

This is a design constraint rather than a preference, and it is the reason a verdict is safe to publish at all. Material the model filters leaves your working queue and stays in your view, with everything the model thought about it still attached to it.

Filtered is a queue, not a bin

Findings the model dispositions land in a Filtered by AI queue that you open, sort and export like any other. Nothing is removed from your tenant, nothing sits behind a support request, and the count is on screen.

The reasoning travels with the finding

Verdict, score, confidence, tags and the plain-language note all stay attached after filtering. You can audit the call on its own evidence rather than asking us what the model was thinking that day.

Anything can be dispositioned back

An analyst returns a filtered finding to the workflow in one action, and from that point the human decision governs. The precedence runs one way only: a person overrides the model, never the reverse.

The model writes only to its own field

An AI verdict and an analyst disposition are separate values on the same record, and the model can write only the first of them. It has no path to overwriting a human decision, and an audit six months later can still tell which call was whose.

What it changes in practice

Two customers, two ratios that do not agree

Both rows below are code-repository monitoring, and they are printed side by side precisely because they disagree by a factor of thirty. A ratio like this is a property of an estate, not of a platform — which is why we publish two of them with their scope attached, refuse to average them into a third number, and do not offer either as a prediction of what a first scan of yours would return.

Customer exampleBrought into monitoringReached an analystRatio
Large private-sector bank ~10,000 repositories ~30 critical-risk 333 : 1
One anonymised customer estate, as of August 2026, scoped to code-repository monitoring alone. The denominator is repositories brought into monitoring; the numerator is what an analyst actually worked once AI Review had ordered the queue, not everything the platform observed. Not a platform average, and not combined with the row below.
Large manufacturing group ~2,000 repositories ~200 actionable 10 : 1
A second anonymised customer estate, as of August 2026, also repository-scoped. The thirty-fold gap against the row above is the reason both are printed: it reflects how each organisation writes and publishes code, not how the model performs. No false-positive rate and no accuracy figure sits beside either row, because we have not established one we would stand behind in front of a contract.

Large private-sector bank

Brought into monitoring
~10,000 repositories
Reached an analyst
~30 critical-risk
Ratio
333 : 1

One anonymised customer estate, as of August 2026, scoped to code-repository monitoring alone. The denominator is repositories brought into monitoring; the numerator is what an analyst actually worked once AI Review had ordered the queue, not everything the platform observed. Not a platform average, and not combined with the row below.

Large manufacturing group

Brought into monitoring
~2,000 repositories
Reached an analyst
~200 actionable
Ratio
10 : 1

A second anonymised customer estate, as of August 2026, also repository-scoped. The thirty-fold gap against the row above is the reason both are printed: it reflects how each organisation writes and publishes code, not how the model performs. No false-positive rate and no accuracy figure sits beside either row, because we have not established one we would stand behind in front of a contract.

Questions buyers actually ask

Before you evaluate this

How do I know the AI isn't hiding something?

Because it cannot. Material the model dispositions moves out of your working queue and into Filtered by AI, a view you can open, sort and export like any other, with the verdict, score, confidence, tags and note still attached to every finding in it. Two structural facts sit behind that. The model writes only to its own disposition value, so it has no path to overwriting an analyst decision or to changing what a finding is. And any analyst can send a filtered finding back into the workflow in one action, after which the human ruling governs. The model narrows what you look at first. It has no power over what you are permitted to see.

What was the model tuned against?

Live enterprise estates, worked by our own analysts. Security Brigade runs managed security services for the same kinds of estates ShadowMap monitors, so the people tuning the model are also the people working the queues it orders, every day. AI Review has been tuned and reviewed iteratively by those managed-service teams rather than fitted to a benchmark corpus, and findings stay auditable throughout — including the filtered ones, which remain available for human review. We are aware that "tuned by practitioners" is a softer claim than a benchmark score. It is the one we can substantiate, and the next question is why we do not publish the harder-looking number instead.

Why is there no false-positive rate on this page?

Because we have not established one we would stand behind, and a figure we could not defend is worse than none. A rate of that kind is only meaningful against a labelled ground truth, and on our own findings the labelling would be ours — so the number would largely measure our own marking. Vendors do publish these. The useful question to put to them is how the ground truth was established, who labelled it, and on whose estate. What we publish instead is the verdict vocabulary in full, two customer examples with their scope attached, and the list of surfaces where we decided not to use a model at all.

Can we turn AI Review off, or tune it?

You can work entirely from the raw queues if you want to — the verdict is an additional field on a finding, not a gate in front of it, so ignoring it costs you nothing but time. In practice the more common request is the opposite: teams ask which surfaces the model does not cover, then handle those with their own rules. The list is published above rather than kept for the implementation call, so you can make that plan before you sign anything.

See the verdicts on your own findings, including the ones we filtered

One apex domain, two business days, a written snapshot — with the Filtered by AI queue open so you can read the calls the model made and disagree with them.