Skip to main content
Evaluation guide · Useful whichever platform you buy

How to evaluate an attack surface management platform in fourteen days.

Most external exposure evaluations test the things nobody doubted: whether the dashboard loads, whether alerts arrive, whether there is an API. A fortnight is enough to settle something harder — whether a platform’s picture of your organisation is correct, and whether a finding it sends you is worth anyone’s morning. Below is the scoping plan, the questions worth asking any vendor in this category, and an honest account of what two weeks cannot tell you.

None of it is specific to ShadowMap. We would rather be evaluated properly and lose than win on a demo.

The decision underneath

Three things an evaluation actually has to settle

A feature matrix will not separate two external exposure platforms. The datasheets have converged, every serious vendor in the category ticks every box, and the boxes were mostly written by vendors. What separates them is attribution, what a finding carries when it arrives, and what happens to it afterwards. Each of the three below can be tested inside a fortnight — and each needs its measure agreed before the trial starts rather than argued over at the readout.

01 Attribution

Is its picture of your organisation correct?

What you are testing
Whether the platform gets from the domains you name to the estate you did not name — subsidiaries on their own registrars, acquisitions still carrying old branding, the marketing property a business unit stood up last quarter — without dragging in hosts that merely resemble yours. Breadth and precision are one test, not two: a tool that finds everything by attributing loosely has found nothing.
How to measure it
Give it a seed list you have deliberately under-specified, and hold back two properties you already know about. Then walk the other direction: take twenty assets it attributed to you and ask, for each one, what evidence tied it to your organisation.
What a pass looks like
It surfaces the properties you held back without being told, and every asset it claims is yours carries attribution evidence you can audit rather than a confidence score you have to trust.
02 Finding quality

Is a finding worth acting on when it arrives?

What you are testing
Not how many findings arrive — a volume comparison rewards whichever platform is noisiest, and the backlog of a first scan flatters everyone. Whether each finding carries enough to act on: what it is, what evidence supports it, who owns the asset, and what the platform concluded rather than merely what it observed.
How to measure it
Triage the first thirty findings as though the trial were production. Record how long each took and what you had to go and find for yourself. Then pick three you believe are wrong and ask the vendor to defend them.
What a pass looks like
Median triage time falls between week one and week two rather than rising, and the defence of a disputed finding cites the evidence rather than the model that produced it.
03 Operational fit

Does anything reach the people who actually fix things?

What you are testing
Whether a finding leaves the vendor console and lands in the queue your team already works, with its evidence and its owner attached — and, where the answer is removal rather than a ticket, whether the platform does the removing or hands you a registrar abuse address.
How to measure it
Wire one integration on day two, not day thirteen, and route a real finding into your own ticketing system. Then read the ticket as the engineer who receives it: is the evidence in the ticket, or is the ticket a link back to a console that engineer has no account for?
What a pass looks like
Somebody who has never logged into the platform can action the ticket without asking anyone a question.

Scoping

How to scope a fourteen-day POC

Fourteen days is the right length: long enough for a discovery pass to settle and for a second week of findings to test whether the first week was luck, short enough that nobody’s attention has moved on. Almost everything determining whether the exercise produces an answer happens before day one.

The long version of this — the seed-list technique, the scoring sheet, and the two weeks broken down day by day — is written up in How to scope a fourteen-day external exposure proof of concept. Take it and run it against whoever you are evaluating.

If you would rather issue it than read it, the POC scoping and evaluation pack is the same method as a set of worksheets — the scope schedule, the seed list, and a weighted matrix for scoring what each vendor comes back with. It names no vendor, including this one.

Vendor questions

The questions to ask any external exposure vendor

These are fair questions rather than traps — we expect to answer all of them, and a good competitor will answer most. Each is here because it has an answer that separates a platform doing the work from one reselling somebody else’s data, and because none of them can be settled by watching a demo. They work as well pasted into an RFP requirements schedule as asked on a call.

Evaluation criteria for external exposure and attack surface management platforms
The questionWhy it separates vendorsWhat a weak answer sounds like
Can you reprocess your own historical data when your extraction improves? A platform that only parses at ingest is frozen at the quality of the day it collected. When its extraction gets better, everything already held stays exactly as bad as it was — and that corpus is usually the largest thing you are paying for. “Our extraction is already accurate.” Or a description of how new data will be parsed, which answers a different question.
Follow up with: when did you last reprocess, what changed, and did anything land in a customer tenant as a result?
Do you publish the vocabulary your system uses to classify a finding? A closed, readable verdict set is a commitment: it names the states a finding can be in and what each one licenses you to do next. An unpublished one lets “high” mean whatever this quarter of model output happens to mean. “Findings are scored one to a hundred.” A score is a ranking, not a vocabulary — ask what a 71 means that a 68 does not.
What happens to a finding your AI judged benign? Every platform in this category now has a model deciding what not to show you. The question is whether that decision is visible, reversible and reviewable, or whether the discarded pile is simply gone. Any statistic about how rarely it gets that wrong. That is a claim about the pile you cannot see, produced by the party that made it disappear.
Ask to see the suppressed queue during the trial, reverse one decision, and confirm the finding comes back with its evidence intact.
Does detection terminate in removal, or in an alert? Finding an impersonating domain and getting it removed are different businesses with different cost structures. Many platforms detect and hand you an abuse contact; some file on your behalf; fewer do it inside the subscription rather than per incident. “We support takedowns.” Ask who files, under whose authority, what evidence pack goes with it, and whether it is metered.
Is active validation bounded to inventory you attributed to me? Testing an asset that turns out not to be yours becomes your incident, not the vendor’s. The bound matters more than the technique: validation should run only against assets already tied to you with evidence, inside a scope somebody signed. “We test everything we discover.” Careful attribution and unbounded testing are in tension, and a vendor claiming both has usually not thought hard about the second.
Can I reconcile your activity against my own logs? You should be able to take a window of vendor testing and find it in your WAF, CDN and edge logs — source addresses, user agents, timing. If you cannot, you have no way to separate vendor traffic from a real attacker during the trial, and no way to audit the boundary you agreed. “We test from a rotating cloud pool.” Rotation is reasonable; declining to disclose the ranges and windows is not.
What does the second year cost, and which units are metered? Renewal is where a discounted first year gets recovered, usually by metering something the trial made look unlimited: assets, takedowns, seats, historical retention. Every one of those is cheap to agree in year one and expensive to argue about in year two. “We will look after you at renewal.” Ask instead for the uplift cap and the metered units, in the contract, before the first year starts.
What do I take with me if I leave? The findings, the evidence, the attribution decisions and the disposition history are your operational record, and in a regulated environment they are audit evidence. Most platforms export current state; fewer export the history that makes current state defensible. “Everything is on the API.” Ask specifically for evidence artefacts and state-change history, not the current finding list.
Who is on the other end when a finding is wrong? By month nine the thing deciding whether a platform is renewed is rarely detection quality. It is whether a disputed finding reaches somebody who can change the system, rather than somebody whose job is to log that you were unhappy. “You get a dedicated CSM.” Ask what that person is able to change — a detection rule, an attribution decision, a suppression — and how long it took last time.

Evaluation criteria for external exposure and attack surface management platforms

Can you reprocess your own historical data when your extraction improves?

Why it separates vendors
A platform that only parses at ingest is frozen at the quality of the day it collected. When its extraction gets better, everything already held stays exactly as bad as it was — and that corpus is usually the largest thing you are paying for.
What a weak answer sounds like
“Our extraction is already accurate.” Or a description of how new data will be parsed, which answers a different question.

Follow up with: when did you last reprocess, what changed, and did anything land in a customer tenant as a result?

Do you publish the vocabulary your system uses to classify a finding?

Why it separates vendors
A closed, readable verdict set is a commitment: it names the states a finding can be in and what each one licenses you to do next. An unpublished one lets “high” mean whatever this quarter of model output happens to mean.
What a weak answer sounds like
“Findings are scored one to a hundred.” A score is a ranking, not a vocabulary — ask what a 71 means that a 68 does not.

What happens to a finding your AI judged benign?

Why it separates vendors
Every platform in this category now has a model deciding what not to show you. The question is whether that decision is visible, reversible and reviewable, or whether the discarded pile is simply gone.
What a weak answer sounds like
Any statistic about how rarely it gets that wrong. That is a claim about the pile you cannot see, produced by the party that made it disappear.

Ask to see the suppressed queue during the trial, reverse one decision, and confirm the finding comes back with its evidence intact.

Does detection terminate in removal, or in an alert?

Why it separates vendors
Finding an impersonating domain and getting it removed are different businesses with different cost structures. Many platforms detect and hand you an abuse contact; some file on your behalf; fewer do it inside the subscription rather than per incident.
What a weak answer sounds like
“We support takedowns.” Ask who files, under whose authority, what evidence pack goes with it, and whether it is metered.

Is active validation bounded to inventory you attributed to me?

Why it separates vendors
Testing an asset that turns out not to be yours becomes your incident, not the vendor’s. The bound matters more than the technique: validation should run only against assets already tied to you with evidence, inside a scope somebody signed.
What a weak answer sounds like
“We test everything we discover.” Careful attribution and unbounded testing are in tension, and a vendor claiming both has usually not thought hard about the second.

Can I reconcile your activity against my own logs?

Why it separates vendors
You should be able to take a window of vendor testing and find it in your WAF, CDN and edge logs — source addresses, user agents, timing. If you cannot, you have no way to separate vendor traffic from a real attacker during the trial, and no way to audit the boundary you agreed.
What a weak answer sounds like
“We test from a rotating cloud pool.” Rotation is reasonable; declining to disclose the ranges and windows is not.

What does the second year cost, and which units are metered?

Why it separates vendors
Renewal is where a discounted first year gets recovered, usually by metering something the trial made look unlimited: assets, takedowns, seats, historical retention. Every one of those is cheap to agree in year one and expensive to argue about in year two.
What a weak answer sounds like
“We will look after you at renewal.” Ask instead for the uplift cap and the metered units, in the contract, before the first year starts.

What do I take with me if I leave?

Why it separates vendors
The findings, the evidence, the attribution decisions and the disposition history are your operational record, and in a regulated environment they are audit evidence. Most platforms export current state; fewer export the history that makes current state defensible.
What a weak answer sounds like
“Everything is on the API.” Ask specifically for evidence artefacts and state-change history, not the current finding list.

Who is on the other end when a finding is wrong?

Why it separates vendors
By month nine the thing deciding whether a platform is renewed is rarely detection quality. It is whether a disputed finding reaches somebody who can change the system, rather than somebody whose job is to log that you were unhappy.
What a weak answer sounds like
“You get a dedicated CSM.” Ask what that person is able to change — a detection rule, an attribution decision, a suppression — and how long it took last time.

Two of these are procurement questions wearing trial clothes. Metering and renewal do not become visible in a fortnight, so ask them in week one and put the answers in the contract — which is the reason our own pricing publishes a floor rather than making you request one.

The honest part

What a POC cannot tell you

Two weeks is a good instrument for a narrow set of questions and a poor one for everything else. It shows discovery breadth and finding quality, because both are visible immediately. It cannot show you how a vendor behaves in month nine, once the novelty has gone and the person who sold it to you has moved to another account. Knowing which is which stops you over-reading a good fortnight — and stops a vendor selling you one.

What fourteen days does and does not settle
What you are trying to learnThe question behind itCan a fortnight answer it?
Discovery breadth Whether the platform gets from your seed domains to the estate you never named. Settled. Visible on the first pass and confirmable on the second.
Attribution quality Whether what it claims is yours really is yours, with evidence you can check. Settled. Sample twenty assets and audit the evidence behind each.
Finding usefulness Whether a finding carries enough context for someone who does not use the platform to act on it. Settled, provided you triage the trial as though it were production.
Workflow fit Whether findings arrive intact in the queue your team already works. Only if you wire the integration in week one. Left to week two, it goes untested.
Noise in the steady state Whether volume settles once the initial backlog clears, or keeps arriving at week-one rates. Partly. A trial is mostly backlog; the steady state is what you will live with and you cannot see it yet.
Behaviour in month nine Whether support, tuning and escalation hold up once the trial attention has moved elsewhere. Not settled. Ask for two references at eighteen months rather than two at three months.
Renewal pricing What years two and three cost, and which units are metered. Not settled by any trial. It belongs in the contract before the first year starts.
How a dispute ends What happens to a finding you believe is wrong and the vendor believes is right. Not settled — a fortnight is too polite. Ask an existing customer for a worked example instead.
Key
  • Fourteen days settles this
  • Only under the condition stated in the row
  • No trial settles this — get it another way

If you run it against us

What this looks like with ShadowMap

Every step above is one we would want run against us, so here is where each lands. Discovery starts from an apex domain and fans out into the estate you did not name — Attack Surface Management does that work, and the evidence tying each asset back to your organisation sits on the asset itself rather than being summarised into a number.

Validation is Continuous Automated Red-Teaming. It answers the only question that changes what a team does on Monday — whether a discovered exposure can actually be used — and it does that where it is safe and authorised, against inventory already attributed to you, inside the scope agreed at onboarding. Checks are non-destructive, and recovered credentials are reported so you can revoke them rather than used to demonstrate access.

The cheapest way to begin is not a trial at all. Ask for an exposure snapshot: one apex domain, no call, and a written account of what is reachable from outside. That tests the first of the three questions — is the picture correct — with no procurement involvement, which is usually enough to decide whether a fourteen-day trial is worth anyone’s fortnight.

What we commit to in an evaluation, and what we will not claim As of August 2026
  • Discovery is outside-in. Nothing is installed, no agent, no credentials handed over, and nothing is touched that is not already reachable from the public internet.
  • Attribution evidence is shown per asset, so you can audit the ones you think are wrong instead of accepting a score.
  • Active validation runs only against inventory we have attributed to you, only where it is safe and authorised, and only inside the scope agreed at onboarding.
  • Every validation request carries an audit identifier, and each probe is recorded — what was sent, when, against which asset and under which authorisation — so you can reconcile our activity against your own edge and WAF logs.
  • Takedowns are unlimited, subject to fair use: notices filed for your own marks, domains, applications and data, at volumes consistent with the estate under monitoring, and not as a channel for filing on behalf of third parties. No per-notice charge and no monthly allowance.
  • Findings, evidence and the audit trail behind them are exportable through the console and the API, and anything you export is licensed to you perpetually — it stays yours, and stays usable with your auditors and regulators, after the subscription ends.

Deliberately excluded

  • We publish no takedown completion time. Removal depends on registrars, hosts and platforms nobody in this category controls, so a published figure is worth less than a commitment you can hold someone to. Ours is contractual rather than published: where managed service is part of the engagement, response and completion targets are set per abuse class in the agreement, with a penalty clause behind them.
  • We publish no accuracy or suppression statistic, and we would rather you distrusted anyone else’s: a figure of that shape describes the findings you were never shown, measured by whoever decided not to show them.
  • We do not claim that everything discovered gets tested. Validation is bounded to what is safe and authorised, which means some findings are reported as observed rather than as proven.

Run the evaluation against your own estate

Start with one apex domain and a written account of what is already reachable from outside. No call, no procurement, and the result is yours whether or not you take it any further.