Skip to main content
All posts
Attack Surface Management

How to scope a fourteen-day external exposure proof of concept

Most external exposure POCs test features that were never in doubt. What fourteen days can actually settle is whether a platform's picture of your organisation is correct — and how to measure that.

ShadowMap Research · April 21, 2026 · 11 min read

Most shortlists of attack surface management tools are assembled from a feature matrix, and most fourteen-day proofs of concept are then run as though the matrix were the thing under test. The vendor demonstrates the features. The features work. A fortnight later the evaluation concludes that the platform does what the vendor said it does, which was never seriously in question.

That is a demonstration, not a proof of concept. The questions a POC can genuinely settle are narrower and considerably more useful: is this platform's picture of my organisation correct, and can the work it produces be absorbed by the people who would have to do it? What follows is a method for answering both inside fourteen days. It is written to be run against any vendor on your shortlist, including us. If you are earlier in the process and still deciding which category of tool you need, start at the evaluation hub instead and come back when a shortlist exists.

What a proof of concept is actually testing

A feature comparison can be completed from documentation in an afternoon. Discovery, correlation, enumeration and prioritisation appear in every vendor's materials in nearly identical language, and by the time a shortlist has been assembled the differences on paper are small. What cannot be established from documentation is accuracy — specifically attribution accuracy — because that is a property of the vendor's method meeting your particular estate, and no two estates produce the same answer.

Attribution is the claim that a given asset belongs to you. Everything else rests on it, and it fails in two directions. A platform that under-attributes leaves assets out of the inventory, which is the problem you set out to solve. A platform that over-attributes fills your inventory with a shared-hosting neighbour's infrastructure, a competitor's domain that happens to contain your brand name, or a subsidiary you divested three years ago. Every one of those becomes an asset, then a finding, then an alert, and eventually a reason your team stops reading the queue.

Neither failure is visible in a demo, because a demo runs against the vendor's own reference environment where attribution was settled long ago. Both are visible within three days of a POC against your estate — provided you decide in advance how you intend to measure them. A POC without pre-agreed acceptance criteria becomes a fortnight of impressions, and impressions get settled by whichever dashboard was best designed.

The seed set: one apex, or the whole group?

The instinct is to hand over a single apex domain. It is clean, it is easy to authorise, and it produces a tidy result. It also tests the easy case, and the easy case is not where enterprise external exposure goes wrong.

Most large organisations have a group problem rather than a domain problem: subsidiaries trading under different names, acquisitions whose infrastructure was never migrated, regional entities with their own registrars, joint ventures, and engineering teams that register domains on departmental cards. A platform that discovers the estate hanging off your primary apex has done something a competent afternoon with certificate transparency logs would also have done.

The more informative scope is a minimum seed against a maximum truth. Give the vendor your company name and two or three apex domains, of which at least one belongs to an entity that does not share your brand name. Then withhold everything else — particularly your asset inventory.

That last point matters more than it sounds. Handing over a CMDB export on day one destroys the only clean measurement the POC will ever offer you. Once a vendor has your inventory, you can no longer distinguish what they discovered from what you told them. Keep the export in a locked file and produce it on day three, after the vendor's inventory has been submitted and timestamped.

Days 1–3: discovery and the reconciliation test

When the vendor's inventory arrives, reconcile it against yours into four buckets:

  • In your inventory and discovered. The baseline. Confirms the vendor can see what you already know about.
  • In your inventory and not discovered. A coverage gap. Ask why for each one — some will be legitimately invisible from outside, which is a fair answer, and some will not be.
  • Discovered, not in your inventory, and genuinely yours. This is what you are buying.
  • Discovered, not in your inventory, and not yours. False attribution. This is what will kill adoption in month four.

Across our own customer base, outside-in discovery typically surfaces 30–60% more external assets than the organisation's own inventory contained. Treat that as a sanity band rather than a target. A vendor returning a low single-digit percentage of new assets against a large decentralised estate is more likely reading your DNS zone than discovering anything — measure it against the 30–60% band above rather than against zero. A vendor returning several times your known estate is not necessarily better — it is more likely attributing loosely, and the fourth bucket will tell you which.

Measure the fourth bucket properly. Take a random sample of fifty assets from the discovered-and-unknown set — chosen by you, not offered by the vendor — and adjudicate each one as yours or not yours. That sample is your false-attribution rate, and it is the single most useful number the POC will produce. No vendor volunteers it.

Then run the removal test. Reject a wrongly attributed asset, and check after the next discovery cycle that it has stayed rejected rather than reappearing as new. Attribution correction that does not persist is not correction; it is a fortnight of the same conversation, repeated indefinitely.

Days 4–7: measure the queue, not the catalogue

Every vendor will report what the platform found. What you need to measure is what reached a human.

Record two numbers. First, the total volume of findings generated across all categories. Second, the number presented to your team as requiring action. The relationship between those two is the product. A platform that generates a large catalogue and passes all of it through has moved the work to you and charged for the privilege.

Then interrogate the suppression mechanism, because this is where evaluations are usually won on faith rather than evidence. Ask what happened to everything that did not reach the queue. In ShadowMap, findings carry an explicit verdict — Confirmed Exposure, Needs AI Review, Likely False Positive or Benign — and anything the model sets aside lands in a Filtered by AI queue that never deletes. Whatever the vendor's equivalent is called, the test is the same: can you open the suppressed set, audit a sample of it, and satisfy yourself that nothing important was quietly discarded? A platform that silently drops findings is asking for trust it has not yet earned. One that files its suppressions in an auditable queue is making a claim you can check.

A calibration note for the categories beyond infrastructure, which is where a fortnight's data is easiest to misread. In a typical first thirty days we surface between five and fifteen active secrets and somewhere between two and eight hundred stealer-log credentials per customer. If two weeks against a large estate returns none of either, the explanation is more likely to be collection coverage than good fortune, and it is a fair question to put to the vendor.

Finally, do not assess noise by counting the vendor's own severity labels. Have one analyst work the live queue for two hours a day and record how many items they would actually have escalated. That number, not the dashboard's, predicts the next twelve months.

Days 8–11: validation and the evidence you received

The test for this phase is a single question, applied to every high-severity finding: could you have acted on this without opening a second tool?

Acting requires evidence, and evidence has a specific shape. For an exposed service it is the request and response that established the condition, with a timestamp and the attribution chain for the asset. For an exposed secret it is what the key is, whether it is live, and what it opens. For a credential it is which authentication surface it was tested against, under what authorisation, and what the successful authentication reached. A severity score and a screenshot are not evidence; they are a request that your team go and reproduce the finding themselves, which is precisely the work you were trying to remove.

Ask each vendor what proportion of the exposures they surfaced during the POC they are prepared to describe as confirmed exploitable, and on what basis. Across our own customer base, 8–15% of surfaced exposures validate as genuinely exploitable. That is a shape rather than a benchmark — estates differ enormously — but it is a useful reference point. A vendor telling you that most of your external exposure is exploitable is counting something other than exploitability.

Authorisation deserves equal attention. Ask what validation requires before it runs: a signed scope, named targets, an agreed window, an escalation contact. A vendor willing to validate against your production estate without asking is a vendor who will eventually do that at a moment you would not have chosen. The differences between the validation approaches on the market — control simulation, autonomous attack-path testing, scheduled human testing, contextual validation of discovered artefacts — are a longer subject, treated in what actually validates external exposure.

Days 12–14: workflow, export and the exit test

The last three days test whether a finding can become a completed task without leaving the platform. Assign one to a real owner. Apply an SLA. Move it through to closure. Confirm that what remains afterwards is an audit trail somebody outside the security team would accept.

Then push a finding into the system your team actually works in — the ticketing platform, not the one the vendor demonstrates best — and check that the evidence survives the journey. Integration counts are worth very little in the abstract; the only integration that matters in a POC is the one you will use every day. Verify that one, and treat the rest of the catalogue as context.

Where this pays off is time-to-action. Our customers see high-severity findings actioned 40–60% faster than their previous process, and the mechanism is not that findings arrive sooner. It is that they arrive with an owner, an SLA and the evidence already attached, so the hours normally spent establishing whether a finding is real and whose it is have already been spent.

Finish with the export test. Request a complete export of every finding, every asset and every piece of evidence in a machine-readable format, and confirm you can obtain it yourself without raising a support request. Not because you are planning to leave — because a platform you cannot export from is a platform you cannot audit, and because the answer you get during an evaluation is the most favourable answer you will ever get.

The seven questions to ask every vendor

Put all seven to every vendor on the shortlist, including your incumbent.

  1. What in this inventory did I give you, and what did you derive? Separates discovery from ingestion, and it is best asked before you hand over your inventory.
  2. Show me the attribution evidence for this specific asset. A certificate transparency record, a registration record, a passive DNS observation. "Our attribution engine" is a description of a process, not evidence.
  3. What is your false-attribution rate against my estate, on a sample I select? The sample selection is the whole question.
  4. What did you suppress, and can I inspect it? Establishes whether the noise reduction is auditable or merely asserted.
  5. Which finding classes do you validate, which do you not, and what authorisation does validation require? The credible answer names the classes the vendor does not validate.
  6. What happens to a finding when it leaves your platform? Covers the integration, the export and the evidence that either survives or does not.
  7. What does this cost at three times this scope? Discovery will expand your scope — that is the point of it — and the price band structure is easier to negotiate during an evaluation than after one.

Published rankings of the best attack surface management tools cannot answer any of these, and it is not a failing on their part. The answers are properties of your estate, and they only exist once someone has run the method against it.

What fourteen days cannot prove

Be explicit about this in whatever you put in front of an approval committee, because a fourteen-day result read as a twelve-month prediction is how evaluations produce regret.

A fortnight samples the categories that depend on collection over time — credential exposure, dark-web material, impersonation activity — rather than measuring them. It tells you nothing about how the platform behaves once the novelty has worn off and the alerts have become routine, which is the real determinant of whether a monitoring programme survives its first renewal. It cannot show you support quality under pressure, and it cannot show you the outcome that actually matters, which is whether your externally exposed surface shrinks over the following year.

Two mitigations are worth the effort. Extend the pilot to thirty days where the vendor and your own calendar allow it; the second month is where a monitoring platform's real behaviour becomes visible. And ask to speak to a customer eighteen months in rather than a reference in their first quarter.

Then write down what remained untested, and hand that to the committee alongside the result. An evaluation that is honest about its own limits is considerably more persuasive than one that is not — which, as it happens, is also the standard you should be holding the vendors to.

Running an evaluation now? The External Exposure Platform RFP & Evaluation Pack contains the scope definition template, a day-by-day fourteen-day runbook with acceptance tests, the attribution-accuracy test in full, a weighted scoring matrix and a like-for-like pricing comparison. → Get the evaluation pack

When the POC concludes, the conversation moves to commercials, and the scope you validated is the scope you should be quoted against — our pricing model is built around exactly the estate dimensions a POC establishes.

Related: Attack surface management · What actually validates external exposure · Evaluating an external exposure platform

Related to

Attack Surface Management Vendor Evaluation Exposure Management EASM Procurement

Ask what ShadowMap would find on your assets.

A 30-minute live walk-through with a ShadowMap engineer on your own domains. We map you live; you keep the report whether or not you choose to engage.