Skip to main content
Reporting · An output of the programme

A security rating is an output. The findings underneath it are the work.

The rating summarises what ShadowMap observes on your external estate and attributes every movement to a category, so a board conversation starts from a reason rather than from a number that changed. It is deliberately not where the work happens: each category opens into the findings that moved it, and those findings are what gets fixed.

0–100
Score, rolled up to an A–F letter grade
One engine
Your estate and your vendors, scored identically
Published
Grade bands and the scoring arithmetic, on this page

What the number is for

A summary is useful precisely because it loses detail

Three jobs a rating does well: it travels into a board pack, it sets a threshold a supplier has to clear, and it draws a trend line long enough to argue with. None of those three is a security outcome.

A score compresses a large amount of observed evidence into one number so that the number can travel. That compression is the entire value, and it is also the entire risk: the score moves for reasons the score does not state. ShadowMap publishes a rating because those three jobs are real and a security team should not have to assemble them by hand each quarter. What it will not do is let the grade stand in for the work. Every category score opens into the findings that produced it — the exposed service, the leaked key, the credential that still authenticates — and the finding is where a response actually lives. If a rating ever starts reading as the answer rather than as the summary, it is being used wrongly, and we would rather say that on the page than after the sale.

The boundary

What a rating can answer, and what it cannot

Both halves of this table are honest. The first three questions are what a score is genuinely good at; the last three are questions only a finding can answer, which is why a rating is never where an investigation ends.

The questionRating answers it?What actually answers it
Is this vendor's posture trending down? Yes Score history across the selected period, one line per vendor, with material drift raised as an alert so degradation is visible before the renewal conversation rather than during it.
Vendor posture recalculates on the same recurring scan cycle as the rest of the platform, so the trend is continuous rather than a point-in-time reassessment someone has to commission.
Is it above the threshold we set for this tier of supplier? Yes The score itself, against fixed published bands and filterable by grade, by score range and by category, so a policy written as a number can be enforced as one.
A threshold is a commercial control, not a security control. It decides whether a contract proceeds. It does not decide whether anything gets fixed.
How does this compare with the rest of our portfolio? Yes Portfolio ranking, and the concentration underneath it — where several of your suppliers sit on the same infrastructure, the same upstream supplier or the same staff, which a vendor-by-vendor view is structurally unable to show.
Eleven suppliers behind one upstream provider is a different problem from eleven unrelated findings, and only a portfolio view tells them apart. It is also the exposure a questionnaire programme is least likely to have thought to ask about.
Is there a working credential for this organisation right now? No A stealer-log record carrying an explicit state. Probed where it is safe and authorised: a credential that authenticated says so, and one that was never tested says that too.
A leaked credential might move a category score by a point. Whether it still opens something is a different fact, and that fact is not in the number. The boundary sits here: probing runs on the estate you authorise us over. A vendor-side credential reaches you observed, dated and evidenced — never tested, because that authorisation is not yours to give.
Is that exposed key live? No The secret itself with its source, and validation where it is safe and authorised to attempt one. Not every finding is validated, and the ones that were not are labelled as such.
Detection and validation are separate steps with separate evidence. A score can only ever reflect the first of them.
Has anyone tested it — and can it be removed? No Continuous Automated Red-Teaming carries the verdict, the probe history and the evidence on the finding record. Removal is a takedown request raised from that same finding: unlimited, subject to fair use.
This is the honest competitive line, and it is not a claim about whose number is more accurate. A rating tells you something is wrong. It does not tell you whether it can be used, and it does not remove it.

Is this vendor's posture trending down?

Rating answers it?
Yes
What actually answers it
Score history across the selected period, one line per vendor, with material drift raised as an alert so degradation is visible before the renewal conversation rather than during it.

Vendor posture recalculates on the same recurring scan cycle as the rest of the platform, so the trend is continuous rather than a point-in-time reassessment someone has to commission.

Is it above the threshold we set for this tier of supplier?

Rating answers it?
Yes
What actually answers it
The score itself, against fixed published bands and filterable by grade, by score range and by category, so a policy written as a number can be enforced as one.

A threshold is a commercial control, not a security control. It decides whether a contract proceeds. It does not decide whether anything gets fixed.

How does this compare with the rest of our portfolio?

Rating answers it?
Yes
What actually answers it
Portfolio ranking, and the concentration underneath it — where several of your suppliers sit on the same infrastructure, the same upstream supplier or the same staff, which a vendor-by-vendor view is structurally unable to show.

Eleven suppliers behind one upstream provider is a different problem from eleven unrelated findings, and only a portfolio view tells them apart. It is also the exposure a questionnaire programme is least likely to have thought to ask about.

Is there a working credential for this organisation right now?

Rating answers it?
No
What actually answers it
A stealer-log record carrying an explicit state. Probed where it is safe and authorised: a credential that authenticated says so, and one that was never tested says that too.

A leaked credential might move a category score by a point. Whether it still opens something is a different fact, and that fact is not in the number. The boundary sits here: probing runs on the estate you authorise us over. A vendor-side credential reaches you observed, dated and evidenced — never tested, because that authorisation is not yours to give.

Is that exposed key live?

Rating answers it?
No
What actually answers it
The secret itself with its source, and validation where it is safe and authorised to attempt one. Not every finding is validated, and the ones that were not are labelled as such.

Detection and validation are separate steps with separate evidence. A score can only ever reflect the first of them.

Has anyone tested it — and can it be removed?

Rating answers it?
No
What actually answers it
Continuous Automated Red-Teaming carries the verdict, the probe history and the evidence on the finding record. Removal is a takedown request raised from that same finding: unlimited, subject to fair use.

This is the honest competitive line, and it is not a claim about whose number is more accurate. A rating tells you something is wrong. It does not tell you whether it can be used, and it does not remove it.

The published bands

What each grade means, and what it is legitimately used for

The bands are fixed, and identical on your own estate and on a vendor you have added. What is deliberately absent is a pass mark: we publish what a score means, and where the line sits for a tier of supplier is a decision for your policy rather than a property of our scale.

What each grade means, and what it is legitimately used for
StateWhat it meansWhat follows
A · 90–100 No material exposure observed in this category on the current scan cycle. Report it. Read it as the absence of observed external exposure, which is not the same statement as the absence of risk.
B · 80–89 Minor observed exposure, none of it severe. Normal remediation queue. The direction of travel matters more here than the value.
C · 70–79 Exposure is present, identified and attributed to a category. Work the findings rather than the grade — the distance to a B is a list, and the list is one click under the category.
D · 60–69 Sustained exposure, usually concentrated in one or two categories rather than spread evenly. The proportionate ask of a supplier here is a remediation plan and a re-score date, not a refused renewal. A grade is an opening position in a conversation; it is not a verdict.
F · 0–59 Severe or widespread observed exposure. A commercial decision, taken on the findings underneath it. The grade is the trigger for reading them; it is never the evidence.
Key
  • No material observed exposure
  • Observed exposure, identified and attributed
  • Severe or widespread observed exposure

The arithmetic

A grade you can reconstruct is a grade you can dispute

A black-box rating is a letter with no derivation attached, which is how a score turns into a procurement checkbox instead of a work queue. The derivation is below, along with what the method structurally cannot see.

From one finding to a letter grade As of Scoring engine, August 2026
  • Every externally observable finding is assigned to a rating category as it is discovered — Vulnerability Management, Network Security, Application Security, Encryption & Certificates, Email & DNS Security, Dark Web & Threat Intelligence, Data Exposure, Brand Protection.
  • Inside a category the score is the geometric mean of its sub-scores rather than a plain average, so one badly weak area drags the category down harder than an average would permit. Findings are weighted by severity, with critical and high dominating; by recency, so an ageing finding loses weight; and by remediation, so a closed finding stops counting entirely.
  • The overall score is the unweighted mean of the category scores, rounded. There is no inter-category weighting, and the consequence is worth stating rather than hiding: a single weak category is diluted by the average and is visible only in the rows. Read the rows for your weakest area and the header for your average.
  • Movement is attributed, not merely reported. Each category carries a plain-language day-over-day delta — "Network Security declined by four points, twelve new findings" — and moves below half a point are discarded as noise rather than drawn as a trend line.
  • Each category carries a ranked list of open recommendations with an estimated score impact, so "what moves this grade" is a list somebody can be handed rather than a question for support.
  • The score recalculates as scans finish rather than on a review cycle, and a short server-side cache means everyone in your organisation reads the same number at the same moment. The same engine produces group, subsidiary and business-unit rollups, and benchmarks against up to five peer organisations you select, overall and category by category.

Deliberately excluded

  • The rating sees what is externally observable and nothing else. It says nothing about your internal controls, your policies or your people, and we will not let it be read as though it did.
  • The estimated impact on a recommendation is directional. It comes from a different weighting model than the one that produces the grade, and it is published as an indication of order rather than a promised movement.
  • No accuracy figure and no false-positive figure sit anywhere on this page, and that is a policy rather than an omission. We have not established one we would defend in front of a contract.

One engine, both sides

The same score on your own estate and on your vendors

A vendor score is computed by the same category engine that produces your own rating — same categories, same arithmetic, same bands. That is what makes an internal target and a supplier threshold comparable numbers rather than two scales that happen to share a letter.

How it gets used

The same engine, seen from three positions

On your own estate
Here the rating is a reporting instrument for something you own and can act on. The board line is the exposure and its direction; the work is the category rows underneath it. Both fall out of the same scan rather than out of a quarterly assembly exercise somebody has to schedule.
On your vendors
Third parties are scored from publicly observable evidence only. They are never contacted, never sent a questionnaire and never asked to participate, so onboarding a supplier costs you nothing but the name. Twenty vendors are included in the base licence, and more can be added.
On the comparison between them
"Our internal target is 85" and "no supplier below 75" become statements on one scale, defensible to an auditor because the methodology did not change when the estate stopped being yours. The drill-down is shared too: a vendor category opens into the breach records, dark-web discussion and phishing pages behind the movement — not into an explanation of the score.

Questions buyers actually ask

Before you evaluate this

Is a good rating the same as being secure?

No, and treating it that way is the single most common misuse of the number. A rating is a summary of what was observed from the outside on a scan cycle; it is very good at travelling into a board pack and very poor at telling you whether a specific credential still authenticates. Use it to report, to set thresholds and to watch a trend. Act on the findings underneath it. The long-form version of this argument, including what a score can and cannot move, is at /blog/what-a-security-rating-does-not-tell-you/.

How does your rating compare with a dedicated ratings platform?

On the number itself it is a fair fight, and we do not pick it. The established ratings vendors have spent a decade on scoring methodology, some hold independent third-party certification of their attribution, and we publish no accuracy figure of our own — deliberately, because a figure like that means nothing without the methodology and the audit standing behind it. One concession beyond that: if a counterparty, an insurer or a regulator has written a specific vendor's score into a contract, ours will not substitute for it. That is a distribution position rather than a methodology argument, and it is real. What differs is what sits underneath the score. A ratings platform can tell you a supplier looks worse this quarter; it cannot hand you the stealer-log record and the compromised device behind that movement, or the page impersonating that supplier and collecting logins from your customers — and it cannot request that the page come down. On the estate you authorise us over it goes one step further again: Continuous Automated Red-Teaming establishes whether an exposed credential or key opens anything at all, which is a fact no score contains. The question worth asking any ratings vendor is what happens after the score.

What if we disagree with a score?

Raise the specific findings with support rather than the number. Every category score decomposes into the findings that produced it, so a dispute is always really about a finding — an asset that is not yours, a service decommissioned last quarter, a record attributed to the wrong entity. Resolve the finding and the score follows on the next cycle. That is the correct order, and it is the same order when one of your own suppliers disputes the score you are holding them to.

See the findings underneath the rating

One apex domain, two business days, a written snapshot. The findings, not a grade.