Skip to main content
All posts
cyber security ratings

What a security rating does and does not tell you

A security rating compresses your external posture into a single letter. Four things the grade structurally cannot see, and how to use it as an index rather than an outcome.

ShadowMap Research · July 29, 2026 · 7 min read

A security rating is a compression algorithm. It reduces a broad set of externally observable signals about an organisation to a number between zero and a hundred, and then to a single letter. Compression is useful — it is the only reason a risk committee can hold two hundred suppliers in its head at all. But every compression discards something, and what these ratings discard is consistent from one provider to the next.

Knowing what gets discarded is the difference between a rating that runs your programme and one that flatters it. For anyone evaluating an external exposure platform, it is also the most useful question to put to a vendor whose demo opens with a grade.

What a rating is actually computed from

Every cyber security rating on the market is built from the same raw material: what can be observed about you from the public internet without your cooperation. TLS and certificate configuration, DNS and email authentication posture, open ports and the services behind them, software versions inferred from banners, IP reputation, breach mentions matched to your domains.

Those signals are sorted into categories, weighted by severity and aggregated. ShadowMap's Security Rating works this way too, and publishes its method rather than treating it as proprietary: eight equally weighted categories — Vulnerability Management, Network Security, Application Security, Encryption and Certificates, Email and DNS Security, Dark Web and Threat Intelligence, Data Exposure, and Brand Protection. Each is scored on the count and severity of its open findings, with critical and high issues weighing disproportionately, and the eight average into a 0–100 score and an A–F grade — per entity and per subsidiary, trended weekly, benchmarked against up to five peers.

The important property of that pipeline is not the arithmetic. It is that every input is an inference drawn from outside the perimeter: no rating provider has an agent on your endpoints, a session in your identity provider, or a row in your ticketing system. A rating measures the shadow your organisation casts on the internet — informative about shape, silent about contents.

What a rating is genuinely good at

Board communication. A directional number that moves slowly and consistently is a legitimate instrument for a quarterly conversation, and for most organisations it is the closest thing they have to cyber risk quantification. Security teams underrate this. The alternative — a slide assembled by hand from four tools that disagree — is unreproducible and therefore unarguable.

Portfolio trend. Across subsidiaries, acquisitions and business units, a common scoring method shows which parts of a group are diverging. A newly acquired entity two grades below the parent is a real finding, arrived at cheaply.

Vendor triage at scale. With four hundred suppliers and a team of three, a rating is the only practical way to decide which forty deserve a human. As a filter it is excellent. As a verdict it is dangerous.

Four things a rating cannot see

A working credential. Ratings register credential exposure as a category signal: records observed, volume trended. What no rating does is test. A first pass across our customer base typically surfaces 200–800 stealer-log credentials per organisation, of which the urgent set is a small fraction — the ones that still authenticate against something you own. A grade cannot separate a password rotated in 2019 from one that opens your VPN this morning, because separating them requires an authorised authentication attempt that rating providers do not make.

A live key. A secret committed to a public repository is either an inert string or an active credential with a blast radius, and the difference is everything. In the first thirty days of monitoring we typically surface 5–15 active secrets per customer. A rating may note that code exposure exists in your category mix; it will not tell you what the key opens.

An origin behind the WAF. Ratings grade the edge they can resolve. If your application is fronted by a WAF and the origin IP is independently reachable, the hostname's observable posture stays clean while the actual path stays open — a consequence of attributing findings to hostnames rather than to the infrastructure underneath them.

Whether anyone acted. The largest omission. A rating measures state, never motion. It cannot tell you whether a critical finding was assigned an owner, breached an SLA, or was closed rather than quietly reclassified. Two organisations with identical B grades can have entirely different security functions — one closing high-severity external findings in days, the other in quarters — and no rating separates them.

Why an A grade and an active compromise coexist comfortably

Consider an organisation with a strong external posture: certificates current, DMARC enforced, no unpatched internet-facing services, no exposed management interfaces. The rating is an A and it is honestly earned. Meanwhile a contractor's personal laptop — never enrolled, outside every inventory — is running information-stealing malware, and the browser has been emptied: saved credentials, live session cookies for the corporate identity provider, authentication tokens, internal hostnames in the history.

Not one input to that A grade changed. The compromise sits on a machine no rating provider can see, in a data class no rating provider tests, and it reaches your estate through a valid session rather than a vulnerability. The grade is not lying; it is answering a different question from the one you needed answered.

Use the rating as an index, not an outcome

An outcome is something you optimise. An index is something you navigate by. The failure mode of cyber security ratings — and every serious provider knows this — is that once the grade becomes the objective, the cheapest route to a better one is to remediate what is visible and inexpensive rather than what is dangerous.

The structural defence is traceability: a rating is safe to manage by when every point of it decomposes into findings a team can work. Three tests, against whatever rating you have today.

  1. Can you get from a category score to the open findings that produced it in one step, without raising a support request?
  2. Does the score decompose by subsidiary, geography or business unit, so that it belongs to an accountable owner rather than a group-level abstraction?
  3. When a finding is validated and closed, does the score move, and can you see the delta?

A rating that fails all three is a procurement artefact. One that passes all three is a work queue with a headline attached, which is the only version worth operating.

The same method, applied to your estate and your vendors'

The argument holds symmetrically: the eight categories that grade your external posture grade a supplier's on identical evidence, which is why ratings became a third-party risk instrument first.

Outside-in evidence and questionnaires answer different questions, and neither replaces the other. A questionnaire captures what a vendor asserts about controls you cannot observe — policy, training, segmentation, joiner-mover-leaver process. Outside-in observation captures what is true of their perimeter continuously rather than annually. The gap shows up sharpest during an incident: when a supplier appears in the news, teams working from continuous external evidence reach a defensible exposure decision 60–80% faster than those starting a questionnaire round.

One concession, because it is real. If an insurer, a regulator or a customer's procurement team names a specific ratings provider in a contract, an alternative rating does not substitute for it. That is a distribution position, not a methodology argument.

What to ask when a vendor leads with the grade

How many assets did you attribute to me, and by what method? A rating scores only what a provider attributed to you, and attribution is hard — outside-in discovery against our own customers routinely finds 30–60% more internet-facing assets than the internal inventory contained. Ask for the discovery evidence, not the discovery claim.

Which findings behind this score are observed, and which are validated? Observation says a service responds in a way consistent with a weakness; validation establishes that it can be used. In our own data, 8–15% of surfaced exposures are confirmed genuinely exploitable — worth knowing before a team spends a quarter on the rest.

What does this grade deliberately not include? A vendor who cannot answer crisply either has not thought about it or would rather you did not.

Is this number for me, or for someone else? An internal instrument and a procurement currency are different products, and confusing them is how a programme ends up optimising a number never designed to run it.

A rating is a good index and a poor conclusion. Treat it as the first question in an external exposure programme rather than the last, and make sure whatever sits underneath it can tell you which credential still works, which key is live, and whether anyone did anything about either.

Evaluating a ratings or exposure platform? The External Exposure Platform RFP & Evaluation Pack includes the ratings section: attribution-accuracy testing, the observed-versus-validated question set, and a weighted scoring matrix procurement can use directly. → Get the evaluation pack or request a quote.

Related: Security Rating · Third-party risk management · Switching external exposure vendors: the first thirty days

Related to

cyber security ratings security ratings third-party risk management exposure management cyber risk quantification vendor risk

Ask what ShadowMap would find on your assets.

A 30-minute live walk-through with a ShadowMap engineer on your own domains. We map you live; you keep the report whether or not you choose to engage.