Skip to main content
All posts
continuous penetration testing

Continuous pentesting, BAS, autonomous pentesting and CART: what actually validates external exposure

Breach and attack simulation, autonomous pentesting, continuous penetration testing and CART all promise proof. What each one validates, and the part of the external surface it cannot see.

ShadowMap Research · May 13, 2026 · 9 min read

Four product categories now compete to answer the same question, and they describe themselves in almost the same words. Breach and attack simulation, autonomous pentesting, continuous penetration testing and continuous automated red-teaming each promise evidence that a weakness is real, reachable and usable, not a scanner's opinion that it might be.

They are not substitutes. Each starts somewhere different, tests something different, and has its own blind spot. The expensive mistake is rarely a bad product; it is a good product bought to answer a question you did not have. A team that buys control-efficacy testing when it needed external discovery ends up with clean dashboards and an incident on an asset that was never in scope.

What follows compares categories, not named products. If you are further along and building the evaluation itself, the platform evaluation hub covers scoring, scope and the questions to ask.

Validation is four different questions

Analysts have begun grouping all four under a single heading, adversarial exposure validation. It is a fair budget line and a poor buying guide. The word "validation" now covers four distinct enquiries, and most arguments about these tools are really arguments about which one was being asked:

  1. Do my controls detect and stop known attacker techniques?
  2. If an attacker had a foothold, how far could they get?
  3. Can a skilled human find a way in, repeatedly, on a schedule?
  4. Of everything visible from outside my perimeter, what is actually usable today?

Each category below answers one of those well and the others poorly.

Breach and attack simulation: control efficacy

Breach and attack simulation executes a library of known adversary techniques, typically mapped to MITRE ATT&CK, inside your environment in a controlled way, and measures what your controls did about it. Was the technique blocked? Detected? Logged and forwarded to the SIEM? Did the alert actually fire, and did it reach anyone?

This is a genuinely valuable instrument, and it is the only one of the four that measures your defensive stack rather than your attack surface. Teams that run it seriously find configuration drift they had no other way to see: an EDR policy that silently reverted, a log source that stopped shipping in March, a detection rule that matched on a field the vendor renamed.

The boundary is built into the method. Breach and attack simulation tests the estate you point it at, using the techniques its library contains, and it discovers nothing on its own. Coverage is bounded twice over: by the technique library, and by your asset inventory.

That second bound matters more than most buyers expect. Across ShadowMap deployments we typically surface 30 to 60 per cent more external assets than the customer's own inventory contained. A control-efficacy programme running against the inventory you already had is silent about all of them, and the dashboard gives no sign that anything is missing.

Autonomous pentesting: attack paths from a foothold

Autonomous pentesting automates the work a network penetration tester does: start from a position, chain misconfigurations, credential reuse and known vulnerabilities together, and reach an objective, whether that is domain administrator, a sensitive data store or a privileged service account.

At this it is very good, and it beats a human at one thing: repeating the exercise every week without fatigue. Path-finding through Active Directory, over-permissioned service accounts and reused local administrator passwords is combinatorial work, and combinatorial work is what automation is for.

Everything depends on where the starting position sits. Most of the value in this category is inside the perimeter, because that is where the chains are. An automated red teaming engine pointed only at a hardened external estate has little to chain: a handful of exposed services, most of them patched, and no lateral movement to reason about. The category is often sold as external validation and delivers most of its value internally.

Put one question directly to anyone selling it: what does the engine actually execute against production, and what does it infer? Real exploitation carries real risk, and vendors draw that line in different places. Settle it in the contract, not in the middle of an outage.

Continuous penetration testing: humans, on a cadence

Continuous penetration testing replaces the annual engagement with recurring human-led testing, usually delivered through a platform that manages scoping, findings, retest and evidence. The testing is still done by people; the "continuous" refers to cadence and workflow, not to automation.

This is the only category of the four that finds business-logic flaws. No technique library contains "the password-reset flow accepts a user ID from the request body", or "the tenant identifier in this API call is trusted from the client", or a chained authorisation failure that requires understanding what the application is for. Those findings come from a human who read the application, and they are often the most serious issues in an external estate.

Its limit is scope, and no amount of tester skill changes that. A penetration test assesses what you asked it to assess. Finding the targets is your job. If the target list is incomplete, the test is rigorous against the wrong estate. The forgotten subsidiary domain, the marketing microsite a contractor registered, the staging environment a certificate transparency log has been publishing for two years: none of these are findings in a scoped engagement, because none were in the scope document.

Cadence is the second constraint. Human testing is expensive, so it is bounded, so it happens in windows. External exposure does not respect windows. A repository goes public on a Tuesday; a credential appears in a stealer log the following week.

Continuous Automated Red-Teaming, as we build it

The term is used loosely across the market, so here is precisely what it means in ShadowMap.

We validate artefacts the platform found on its own, from outside, with no prior knowledge, credentials or agents. The sequence matters: discover the estate, attribute it to you, surface the exposures (assets, leaked source code, exposed keys, credentials and stealer-log material, brand abuse, dark-web mentions, vendor posture), then validate the classes where validation is safe, authorised and decisive.

In practice that means establishing whether an exposed key is live and what it opens; whether a leaked credential still authenticates against an in-scope authentication surface, and what access it carries when it does; whether an exposed administrative interface is genuinely reachable; whether a misconfiguration actually yields data or merely looks as though it might.

Several categories run validation of some kind. Ours runs inside a platform that already holds the discovered assets, the leaked code, the exposed keys, the credential and stealer-log corpus, the brand-abuse signals and the vendor context. A validated finding arrives already correlated to everything else known about that identity, that host and that supplier. A validated credential comes with the named service it reached, the host discovered three weeks earlier, and the device that also leaked live session cookies.

Across deployments, between 8 and 15 per cent of surfaced exposures validate as exploitable. The other 85 to 92 per cent is real exposure that does not need to be worked this week, and knowing which is which is what moves time-to-action on high-severity findings by 40 to 60 per cent. In the first thirty days we typically surface between 5 and 15 active secrets: keys and tokens that are live, not expired credentials in an old commit.

The comparison

Starting point What it proves What it requires Where it is blind
Breach and attack simulation Assets you nominate, inside your environment Your controls detect, block and log known techniques Deployment inside the estate; a maintained technique library Assets not in your inventory; anything outside the perimeter
Autonomous pentesting A foothold, network range or credential set An attacker with that position could reach a defined objective Internal access; agreed exploitation boundaries External-only estates; exposure that is not a network path
Continuous penetration testing A scope document you supply A skilled human could or could not get in, including business-logic routes Human time; a defined and current scope Anything outside scope; anything that changes between windows
Continuous Automated Red-Teaming (external) Nothing; discovery is the first step Which externally discovered exposures are actually usable Contractual authorisation and progressive scoping Internal estate; anything requiring agents or endpoint telemetry

What none of these replaces

A scoped human red team. Automated validation confirms the classes of exposure it was built for. An objective-driven red team tests the organisation against a goal, using whatever route presents itself, including ones nobody anticipated. It puts people, process, detection, escalation and decision-making under pressure, and no automation produces that.

ShadowMap is built by Security Brigade, which has delivered more than 6,700 security assessments since 2006 and has been CERT-In empanelled since 2008. We sell both human testing and automated validation.

Internal vulnerability management. None of the external categories tells you the patch level of a server inside your network.

The hard boundary: what our validation does not do

Validation is not enabled at go-live. Onboarding begins with discovery and enumeration; your inventory is built and confirmed first, and validation is introduced progressively within the scope set by the agreement and the onboarding configuration. Authorisation is contractual, not assumed.

We do not attempt authentication against systems outside your authorised scope. We do not run public exploit code against your production estate. We do not test credentials belonging to your customers. They remain exposure you need to know about, and we report them as untested.

We do not validate every category. Phishing infrastructure, brand mentions, threat-actor reporting and media monitoring cannot all be validated.

If a vendor quotes you a single accuracy percentage across every capability and every kind of finding, ask what the denominator was and who chose it.

We are external. No agents, no endpoint telemetry, no Active Directory path-finding. If your open question is internal lateral movement, this is not the category you need.

Choosing between them

If the uncertainty is your controls, and you have the inventory and the stack but do not know whether it works, breach and attack simulation answers that and the others do not. If your concern is what an attacker does after landing inside, autonomous pentesting is the discipline built for it. If you need business-logic assurance on commercially significant applications, that remains human work on a cadence.

If you cannot state with confidence what you own on the internet, all three are scoped against an inventory you have not verified, and that is the problem external validation solves first. Most mature programmes run two of these, not one. The mistake to avoid is buying the second and assuming it covers the first.

When you get to the evaluation itself, test discovery before you test validation, because validation quality is meaningless on an estate the platform never found. Our guide to scoping a fourteen-day proof of concept sets out how to measure that in practice, and the alternative-stack cost comparison covers what the same coverage costs assembled from separate tools. When you are ready to specify what you need, start here.

See what a validated finding actually looks like. Sample CART output: the evidence attached to a confirmed exposure, the state model behind it, and the scope boundaries it was produced within. → See a CART validation sample

Related: Continuous Automated Red-Teaming · How to scope a 14-day exposure POC · Evaluate an external exposure platform

Related to

continuous penetration testing breach and attack simulation adversarial exposure validation continuous automated red-teaming autonomous pentesting external attack surface management security validation vendor evaluation

Ask what ShadowMap would find on your assets.

A 30-minute live walk-through with a ShadowMap engineer on your own domains. We map you live; you keep the report whether or not you choose to engage.