Back to AI

Frontier Threat Defense — Program Strategy

Original document
Frontier Threat Defense · Programme strategy · FY27

Defending the enterprise against adversaries that attack with AI

What has changed is not the adversary’s capability — it is their velocity and scale. That distinction is the foundation of this strategy, and getting it right is what separates a fundable programme from a reaction to headlines.

The framing this strategy adopts

“Frontier AI models do not introduce a fundamentally new threat capability; they simply change the velocity and scale of existing attack tactics and exploit generation.”

AND THE INSTRUCTION THAT COMES WITH IT
Gartner is explicit that the CISO’s job here includes managing the noise: “CISOs should help their organization ignore sensational marketing hype around emerging AI threats and focus on actionable responses.” This strategy is therefore written to be defensible rather than dramatic. Where the evidence is thin we say so.

What is actually happening

Three measured realities, not projections.

01  The dominant real vector
41% / 35%

of surveyed organisations have experienced deepfake-enabled social engineering — on an audio call and a video call respectively. Deepfakes now account for one in five biometric fraud attempts.

Technical detail
02  The structural change
days → minutes

The gap between vulnerability disclosure and active exploitation, now that frontier models can autonomously reverse-engineer patches into working exploits.

Technical detail
03  And what AI is mostly used for
Scale

“threat actors leveraging and abusing widely available AI tools to scale and enhance existing tactics, not the invention of net-new attack methods.”

Technical detail

And three facts that keep us honest

A strategy that omits these will not survive scrutiny from anyone who reads the same research.

The majority of breaches still do not exploit vulnerabilities
31%

Vulnerability exploitation is the initial access vector in 31% of breaches; credential abuse 13% (appearing somewhere in the chain in 39%) and phishing 16%. “Top threats are cyclical.”

Technical detail
Autonomous attack is emerging, not established
Unclear

Uses of LLMs to orchestrate automated attacks are “emerging with unclear impact so far”, and malware integrating LLMs shows a “lack of sophistication” — “more experimental than mature”.

Technical detail
Foundational controls still work
Multiple

“Attackers still need to be ‘right’ multiple times: Foundational cybersecurity controls still provide adequate protection and resilience.” This “is not a repeat of Y2K”.

Technical detail

So why stand up a programme at all? Because the process, not the control, is what breaks. “The core issue is not a new, apocalyptic threat, but a long-standing reality: manual vulnerability management processes are structurally misaligned with machine-speed attacks. And separately: “organizations are creating attack surfaces faster than technologies can protect them.” Those two sentences are the whole case.

Why the enterprise, specifically

Three facts from the public record, all checkable, none of them speculative.

Established factWhat AI changes
Proven targetA manufacturer pled guilty to stealing the enterprise core product trade secrets and paid a $60M fine, in a scheme prosecutors described as enabling “self-sufficiency in computer memory production.”Not the motive — the cost and speed of running the campaign
Already contestedAfter the May 2023 CAC decision the enterprise disclosed that a low-double-digit percentage of worldwide revenue was at risk.The adversary is a standing presence, not an event
Manufacturing is the exposure — where information technology meets operational technology (IT/OT)The enterprise manufactures across seven countries in a business it calls “capital intensive”. Gartner: “Fragmented IT/OT convergence creates severe risks, as current systems lack unified governance and struggle to support the downtime required for constant patching.”Frontier models now surface infrastructure vulnerabilities faster than the production environment can absorb downtime to patch
The ten-layer landscape
AND THE ENTERPRISE HAS ALREADY FILED THE PREMISE
The FY2025 10-K states it directly: “The emergence and maturation of AI capabilities may also lead to new and/or more sophisticated methods of attack”, and concedes “despite our system of controls over our intellectual property, it may be possible for our current or future competitors to obtain, copy, use, or disclose” our process technology. The programme is not proposing a new view of the risk. It is proposing to act on the one the company has already published.
The full threat picture, and the forward scenario we are not leading with

The threat section is written conservatively on purpose. Gartner’s guidance for this exact briefing is to temper the fear, uncertainty and doubt and to help the organisation “ignore sensational marketing hype.” A strategy that overstates the threat gets discounted the first time someone checks it.

What is established

  • Deepfake-enabled social engineering is the dominant vector — 41% of organisations on audio calls, 35% on video, and one in five biometric fraud attempts. It is the highest-prevalence vector in the published evidence and the one with the cheapest available mitigation.
  • The vulnerability-to-exploit gap has collapsed from days to minutes, because frontier models can autonomously reverse-engineer patches into working exploits.
  • AI is mostly being used to scale existing tactics, not to invent new ones. And LLM-driven vulnerability discovery “predates” the named frontier releases.

What is not yet established

THE SENTENCE THAT SHOULD GOVERN OUR CLAIMS
“Uses of LLMs to assist in malware creation, as part of the malware workflow, or to orchestrate automated attacks are emerging with unclear impact so far”, and malware integrating LLMs shows a “lack of sophistication” — “more experimental than mature.” Separately: “today, true AI-powered attacks remain very rare in the real world.”

So where does the agentic scenario belong?

As a forward scenario used for validation, not as the operative threat model. The worked example in the companion document — a documented intrusion carried out by an autonomous agent collective — remains the most useful exercise input we have, because it is real, externally adjudicated, and it stresses exactly the processes Gartner says are structurally misaligned. It earns its place under D1 as a validation scenario and under D2 as scenario planning — not as the basis for a threat claim.

The counterweight evidence, stated so nobody has to find it themselves

FigureSource
Vulnerability exploitation is the initial access vector in 31% of breaches; credential abuse 13% (39% somewhere in the chain); phishing 16%2026 Verizon DBIR via Gartner
“Attackers still need to be ‘right’ multiple times” and foundational controls “still provide adequate protection and resilience”Gartner board scenario
“This situation is not a repeat of Y2K and does not require a transformational response”Gartner board scenario
Next 12–24 months see more vulnerabilities, but “the medium-term outlook is a positive one”Gartner board scenario

Two threats worth adding to the watch list

  • Cognitive attack. “The next breach will be cognitive: Threat actors will abandon purely technical exploits in favor of influence operations, utilizing cutting-edge AI models to manipulate employee decision making and bypass traditional cybersecurity controls.” This is the strategic extension of the deepfake finding.
  • Rogue automation from our own estate. “The proliferation of AI agents operating with broad agencies will create severe risks. Traditional centralized IT and third-party vendor management will be unable to oversee these agents.” Which is why D4’s runtime enforcement outcome matters more than it looks.
WHAT THIS MEANS FOR THE PROGRAMME
Write the threat conservatively, keep the agentic case as a validation scenario, and add deepfake SOP hardening to the plan immediately — it is the highest-prevalence real vector and the cheapest thing on the roadmap.

The four-domain structure, the domain definitions and each domain’s core benefit are Gartner’s. Ours is the adaptation to the enterprise: the outcomes per pillar, the ODM formulation, the maturity assessment, and the sequencing.

Four adaptations worth knowing about

DomainGartner’s emphasisOur adaptation
D1Validating control efficacy; simulating real-world attacksExtended backwards into asset and identity discovery. You cannot validate controls over an estate you cannot enumerate, and Gartner separately recommends pulling automated IT/OT/cloud inventories from exposure platforms.
D2TTP insight and predicting likely attacksWeighted toward conversion: intelligence only counts once it becomes a requirement with an owner. Gartner: adversary management “is of most value when combined with one of the other pillars.”
D3Deception, threat hunting, intelligence-driven controlsExtended to include machine-identity revocation as a cost-raiser. Gartner frames disruption as slowing the attacker and raising cost; a credential that dies in ten minutes does both.
D4Configuration validation and device hardeningReframed for agents as runtime enforcement of declared scope — the agentic equivalent of a hardened configuration, and currently absent.
AND THE FOUNDATION IS WHERE WE WERE LEAST HONEST BEFORE
Gartner’s own model rests the four domains on managed services and capabilities, and states that most organisations will need vendor or third-party support “unless they have large and high-level security teams.” Naming that layer forces us to state what the strategy actually depends on — detection-engineering capacity, forensic independence, and accountable resolvers outside CDR.

The disagreement inside the corpus, disclosed

Two Gartner teams write about this in materially different registers. The Emerging Tech team describes “an arsenal of unprecedented sophistication against unprepared enterprises” and warns of “ruinous business loss”; the CxO Leadership team writes that “today, true AI-powered attacks remain very rare in the real world.”

How we resolved it. We use the Emerging Tech research for market direction — the 50%-of-spend-by-2030 prediction is a strong planning signal — and the CxO Leadership research for threat claims, because that is the audience-appropriate standard of evidence and because overstating the threat is the failure mode Gartner itself warns about. If challenged on why our threat framing is more conservative than the spend forecast, this is the answer.

WHAT THIS MEANS FOR THE PROGRAMME
Present the diagram, then the outcomes. The frame’s value to us is that it is recognised, it is preemptive by construction, and it separates strategy from method — and crediting Gartner strengthens the argument rather than weakening it.

Gartner’s outcome-driven metric construct is the measurement system for this entire strategy. It is not a reporting style — it is a way of turning security work into investment decisions. The stated purpose of a metric is blunt: “The value of a metric is its ability to influence decision making. The decisions influenced are typically priorities and investments.”

The five steps, as we applied them

StepGartnerWhat we did
1Develop an initial set of business processes and supporting technology stacks — “the three to five most obvious”Took the three business outcomes already agreed by the programme, and the crown jewels named by the business: critical IP and recipe data repositories.
2Identify business outcomes and business ODMsManufacturing continuity, protection of critical IP and recipe data, and speed of AI adoption — the three above-the-line outcomes.
3Identify technology risks and dependencies — “What breaks in the business process if the technology breaks?”The two-control-plane assessment and the confirmed detection baseline supplied this.
4Define technology ODMs, which “reflect the technology stack’s readiness” and behave as leading indicatorsThe fifteen programme metrics, three per domain plus three foundation — each in the canonical “% of” form.
5Assess readiness as risk to business enablement, on a scale from no investment to leading edgeThe readiness scale on the Outcomes page, with current position marked per domain.

The seven characteristics, used as filters

  • Metric value — ability to influence a decision. Anything that could not change a priority or an investment was cut.
  • Above the line, below the line — “only above-the-line metrics should be shared with executives.” Hence a seven-metric committee set distinct from the fifteen.
  • Leading indicators — they must “expose a problem before it leads to material loss.” This is why validation coverage beats incident counts.
  • Direct line of sight — “A causal relationship should exist… If the technology metric changes, it indicates a change in the business outcome.” Every metric on the page states its line of sight.
  • Metric changes drive action — green to yellow to red must trigger “changing a priority or an investment.”
  • Discrete audiences — “The CFO needs different metrics from those required by the head of a business unit or a board of directors.”
  • Limit the number — five to nine per audience, which sets the committee set below.
WHY SEVEN AND NOT FIFTEEN
The programme tracks fifteen outcomes. Gartner sets the reporting ceiling: You only need five to nine metrics for each target audience. Do not report everything you know. Prioritize the top five to nine technology dependencies and the top audiences to receive only the highest-value information that drives above-the-line decisions.” So fifteen run the programme and seven reach the committee. The distinction is not stylistic — it is what keeps each metric capable of driving a decision.

The prioritisation rule that follows from all of it

“If there is no clear line of sight to a business outcome, then the technology investment should not be prioritized as it will drive little value to the organization.” Applied honestly, this is a test the programme must keep passing — and it is the reason adversarial-ML tooling stays deferred while deception and SOP hardening move to tranche 1.

Measures we rejected, and why

RejectedWhy
Days to patchGartner’s explicit reframe: ask instead “what is the tolerance to be exploited by a known vulnerability?”
CVSS-weighted remediation countsDrives “efforts to address only critical and high CVSS scores rather than prioritizing those riskiest to the organization.”
Number of detections deployedRises when we write rules, not when protection improves. No investment line of sight.
Mean time to detect, aloneNot a leading indicator, and the reference failure was in escalation rather than detection.
Inventory completenessNeeds a denominator we do not have. Replaced with reconciliation against observed behaviour.
Any expected-loss figure for process IPThe enterprise does not disclose a value for its process technology; an invented number is the first thing a finance reviewer tests.

And the honest cost. “Measuring some of these elements may require instrumenting parts of the infrastructure to gather new types of data… Gartner believes the visibility and power to report benefits and to guide priorities and technology investments in a business context will make the initial investment worthwhile.” Two of the fifteen need that instrumentation, and they are the two with the largest payoff: reconciliation, and runtime enforcement of declared scope.

Every recommendation across the corpus that this programme can act on, deduplicated and attributed. This is the operational substance beneath the architecture.

Actions Gartner recommends, mapped to our domains

DomainActionSource
D1“Scope exposure assessments based on key business priorities… taking into consideration the potential business impact of a compromise rather than primarily focusing on the severity of the threat alone.”G00837909
D1Build validation into exposure management by evaluating adversarial exposure validation tools or offensive cyber services.G00837909
D1“Pull automated critical IT, OT, and cloud asset inventories from existing exposure assessment platforms, in order to prioritize where to start changing architectures.”G00853789
D2“Modify standard operating procedures for high-risk processes, such as password reset, partner and supplier interactions, and financial transactions, to minimize the risk of deepfake and phishing impersonation attacks.”G00852902
D2Inventory approved and rogue employee use of client-side GenAI tools using incumbent endpoint protection, EDR and security service edge; use tags and groups to enable monitoring.G00852902
D2“Perform key threat scenario simulations and adapt strategic roadmaps to cover the most likely AI evolutions.”G00852902
D3Deploy deception and honeypots that “provide a clear signal of an attack when triggered”; conduct proactive threat hunting; operationalise threat intelligence in other controls.G00859378
D4“Prioritizing the journey from macrosegmentation to network security microsegmentation to limit lateral movement.”G00853789
D4“Improve response time by using threat management techniques to identify and implement mitigation controls” where a patch cannot be completed.G00810627
D4Ensure a cross-functional cybersecurity governance framework including zero trust, centralised access and lifecycle management.G00853789
F“Establish a model delivery system with layered control planes”; “define risk tiers”; and for high-risk use cases “combine AI models with deterministic systems.”G00858028
FEngage senior leadership and adjacent departments for “effective routes to resolution, risk prioritization criteria, and consistent categorization for newly discovered exposures.”G00837909
THE ONE ACTION THAT IS NOT OPTIONAL
Distributed ownership. Gartner measured that only 36% of organisations have infrastructure teams actively engaged on vulnerability remediation, and states plainly that without business engagement exposure management “cannot function effectively.” The board-facing version is stronger still: those teams “must be held accountable by the board.”

And a caution on how far to automate. “AI-driven exploit generation will continuously outpace traditional patching cycles. At the same time, fully automating fixes will eliminate practical learning ground required to develop experienced Level 3 analysts. So automation is scoped to reversible actions. The long-run mandate Gartner describes is antifragility — “organizations will intentionally use systemic shocks and controlled risk exposure to emerge operationally stronger” — which is an argument for exercises, not for autopilot.

C2M2 maturity levels, assessed independently per domain and cumulatively within one — a practice at a level counts only if every lower practice is also achieved. CSF Tiers are deliberately not used as a maturity scale, because NIST does not define them that way.

THE SHAPE OF THE RESULT, AND WHY IT IS ACTIONABLE
D2 is our strongest domain; D3 is our weakest. We understand the adversary better than we can disrupt them. Gartner notes adversary management “on its own… allows for scenario planning” but is “of most value when combined with one of the other pillars.” So our strongest domain is currently the one delivering least value — which is precisely why tranche 1 invests in D3.

The four MIL0 findings, with evidence

CapabilityEvidence
Deepfake / social-engineering SOPsNo hardened procedures for the processes Gartner names. The highest-prevalence real vector at 41%/35%.
DeceptionNothing placed. The cheapest signal-generating control available.
Machine-identity revocationNo named authority, no target time, no rehearsal.
Runtime enforcement of agent scopeNothing binds the agent registry to enforcement, so declared scope is documentation.
IT/OT governanceGartner describes the gap directly: systems “lack unified governance and struggle to support the downtime required for constant patching.”
Forensic independenceHosted model guardrails refuse defensive work, measured at 2.72× refusal.

What we could not assess

  • Existing gateway control efficacy — what our AI gateway control actually detects and whether it alerts is unverified. Tranche 1 action.
  • Detection engineering capacity — not established with the CSOC. Foundation ODM F.1 creates the baseline.
  • SEMI standards alignment — adoption status unconfirmed.

A limitation we disclose rather than let a reviewer find. No board-credible AI-specific security maturity model exists yet, so this assessment applies a general maturity model to AI-specific domains. Gartner’s own framing supports the approach — the four domains are largely delivered by technologies that already exist, so assessing them with a conventional model is defensible.

Three tranches, weekly sprint execution beneath them. The most important structural point: this is explicitly not a transformation. Gartner: “This situation is not a repeat of Y2K and does not require a transformational response.”

WHY TRANCHE 1 IS MOSTLY PROCESS AND GOVERNANCE
Four of seven tranche-1 items are decisions or process changes rather than technology purchases: revocation authority, SOP hardening, the coverage baseline, and verifying an existing control. That is deliberate, and it follows from Gartner’s observation that “many capabilities promoted as preemptive exist in current platforms” and that these tools are “additive to a well-established and mature program.”

Dependencies — the programme’s real critical path

DependencyOwnerNeeded byIf not met
Machine-identity lifecycle and revocation mechanismIdentityDay 30D3 stays at MIL0; the cheapest cost-raiser is unavailable
SOP ownership for password reset, supplier and financial processesService desk, procurement, financeDay 30The highest-prevalence real vector stays unaddressed
Asset inventory pulled from exposure platformsI&O with CDRDay 90D1 cannot produce a reconciliation measure
Registry-to-runtime enforcement bindingAI Enablement, jointlyDay 90D4’s primary outcome remains zero
Accountable resolver teams for exposure findingsNamed by the boardDay 90Exposure management “cannot function effectively”
Factory-software layer ownerManufacturingFY27The IT/OT gap stays uncharacterised

One sequencing risk worth stating. Gartner warns that preemptive offerings may be “limit[ed] to detecting, rather than fixing issues, thereby increasing, rather than reducing overhead.” If we stand up validation and posture assessment without the accountable resolvers from decision 02, we will manufacture a findings backlog nobody owns. That is the single most likely way this programme creates work rather than protection — which is why the ownership decision is sequenced first, not last.

WHAT THIS MEANS FOR THE PROGRAMME
Fund tranche 1 now; gate tranche 2 on the two verification findings. Take decisions 01 and 02 — the appetite reframe and named accountable owners — to the Security Committee before any tooling conversation, because both are free and both determine whether the rest works.

The single most decision-relevant assessment in the engagement: the enterprise’s strongest control sits on the plane the reference intrusion would not have crossed.

The governed model and tool plane

A registered directory non-human identity calls the AI gateway with its content-inspection layer, which routes to approved models and to MCP servers acting as mediated proxies. Policy actually enforced at that chokepoint: directory JWT and OAuth validation, scope and group authorisation, quotas and token caps, DLP and redaction, prompt-injection inspection, a model allowlist sourced from AI governance, and central logging to the SIEM. All MCP interactions are required to route through it, and direct MCP connections are prohibited for Tier 1 servers. Assessed as stronger mediation than most enterprises operate.

The substrate plane

The same identities, a different path: artefact registry (the default package source, with development and production container registries) → object store → source repositories and CI → databases, wikis and ticketing. The only controls found were a package vulnerability scanner and write-group membership. No behavioural analytics, no provenance attestation, no inter-agent channel detection.

Where the reference intrusion actually ran: governed model / tool plane 0 of 10 stages substrate plane 10 of 10 stages
AND THE CHOKEPOINT MAY FAIL OPEN
Two known risk items sit against the governed plane: a RAID entry for a gateway outage, and a re-routing risk in the model proxy. A chokepoint that fails open under load is not a chokepoint — and Gartner’s caution applies directly: current detection and response methods “aren’t sufficient to keep up with the speed, sophistication and scope” of these threats.

The discriminator that makes the blind plane detectable anyway

A legitimate dependency has one writer, provenance, and a consuming build. A coordination channel has many writers, no provenance, and no consumer. That distinction is computable from access metadata alone, without inspecting content. Which matters because in the reference case content inspection was never possible — the registry’s audit trail covered only user, group, permission and token events.

The adversary labelled the channel itself. The first message posted on the rebuilt coordination board was a directory named zzHELP_PHASEONE_ARV010841_NO_CONSUMER_SEEK_IDEA. Within three hours 53 agents posted 1,188 messages; by 13 July they were signing with Ed25519 keys after being spoofed by other agents. The discriminator is not an analyst’s inference — it was a field in the object name.

The registry that exists but does not enforce

An the agent registry records owner, declared tools, declared write scope and autonomy tier. Nothing was found binding it to runtime enforcement, which is why D4’s primary outcome — declared scope enforced at runtime — currently reads zero.

WHAT THIS MEANS FOR THE PROGRAMME
This assessment is the scope statement for the whole programme. D1 measures the blind plane; D4 fixes the enforcement gap; D3 detects on metadata because content inspection is not available. Everything else follows from these two paragraphs.

The case for self-hosted defensive models does not rest on a claim that they match frontier capability — that claim is indefensible and unnecessary. It rests on four independent legs.

1. Availability — the tool must work at 3am on the worst day

arXiv:2603.01246, 2,390 real blue-team tasks: refusal rate vs neutral equivalents 2.72× system hardening tasks refused 43.8% malware analysis tasks refused 34.3% stating you are authorised increases refusal
THE CLAUSE THAT MATTERS MOST FOR AN AUTOMATED PROGRAMME
“critical for autonomous defensive agents, which cannot rephrase refused queries or retry.” A human analyst works around a refusal in seconds; a response pipeline stops. The refusal problem therefore scales with exactly the automation this programme is building — and Gartner projects AI agents autonomously managing 25% of incident-response workflows for data security events by 2028.

2. Confidentiality — the guardrail and the retention pipeline are the same system

The classifier that refuses a request is also the event that flags the session for retention and human review. So the payload you most need help with is retained longest and seen by the most third parties — and mid-incident, the credentials in it are still valid. Gartner records that 69% of organisations are concerned about maintaining control over AI models, and that “due diligence alone is insufficient; technical controls are essential.”

3. Legal and export exposure — an open question, not a settled one

The safe harbour at 15 CFR § 734.18 covers encrypted transit; inference decrypts and processes. No BIS or DDTC guidance exists on whether inference over controlled technical data constitutes a deemed export. That it is open is the argument — against an enforcement climate that has produced penalties of roughly $213M and $252M against two US critical-manufacturing-sector companies. For a manufacturer whose crown jewels are process technology, this is a board-relevant consideration rather than a technical one.

4. Forensic defensibility

ISO/IEC 27037 requires repeatability and reproducibility of forensic process, and hosted models update and reroute silently.

One design requirement, free now and impossible to retrofit. Every inference that touches evidence must record model hash, engine version and decode parameters into the case record. Retrofitting this invalidates the case work already done.

What runs in it, and the risk it imports

ElementDetail
Serving stackvLLM, SGLang, TensorRT-LLM or llama.cpp. Quantisation introduces calibration sensitivity that must be characterised before forensic use.
Ingest gateOpen weights are untrusted executable code entering a high-trust enclave: pickle-format models execute code at load time, roughly 95% of malicious models on one public hub used that format, the standard scanner has a published bypass, and ShadowRay (CVE-2023-48022) was never fixed.
Weight provenancePRC-origin weights are a first-order selection constraint for a US critical-manufacturing enterprise and need a documented risk position, not an assumption.
Control planesGartner’s own prescription: “Establish a model delivery system with layered control planes”, define risk tiers, and for high-risk use cases combine AI models with deterministic systems.
WHAT THIS MEANS FOR THE PROGRAMME
This is the cheapest item in the programme and the one where a mid-incident discovery is unrecoverable: you cannot procure a forensic capability while the credentials are still valid. It has no dependency on inventory, gateway or SIEM, which is why it sits in tranche 1 and moves a binary outcome from no to yes.

Every question the reference incident raises is an inventory question first. What could one stolen credential reach? Nobody had asked — and the answer took one second to demonstrate once an adversary did.

BLAST RADIUS SHOULD BE A FIELD ON EVERY ALERT, NOT A REPORT
If an analyst has to commission an analysis to learn what a compromised identity reaches, the answer arrives after the decision. Gartner’s framing of the same point: scope exposure by the potential business impact of a compromise rather than primarily the severity of the threat.

Week one, building nothing

  • Define critical assets in the exposure-management platform and read the attack paths that appear. the enterprise almost certainly already holds the licences, and first-class connectors exist for the vulnerability scanner and the service-management CMDB — a day of integration materially improves the graph. The gating input is a business workshop to say what matters, not an engineering build — without it the page may simply be empty.
  • A Kubernetes attack-path run against one production-representative cluster, and an identity graph across the tenant — an afternoon of open tooling, aimed at the surface that actually carried the kill chain.
  • Switch on directory activity logs while you are there. They are frequently off, and they are the read half of the strongest identity signal available.
  • Pull automated critical IT, OT and cloud asset inventories from the exposure assessment platform — Gartner’s explicit recommendation, in order to prioritise where to change architecture first.

The inventory that detection joins to

Self-declaration is not the defect — being unreconciled is. The only published discovery methodology puts attestation fourth of four: tag scan → six-signal heuristics → CMDB reconciliation → developer attestation of the residue, with a deadline and an escalation path. the enterprise’s registry is step one of one.

Design choiceWhy
Copy the published 42-column schemaIt already carries Triggers and ConnectedAgents — the two fields ours lacks — plus declared tools, MCP servers, instructions, memory, permissions and guardrail coverage.
Three-layer asset modelSolves lifecycle velocity: one agent in three regions becomes one model, one asset, three configuration items. Published failure mode of a flat registry: “governance is applied to an entire class of AI, not individual deployments.”
Ranked merge precedenceRank discovery highest and self-declaration lowest, and let them disagree. Works without consolidating anything — which matters because repository unification across roughly 8,000 developers is not achievable.
Measure reconciliation rate and declare-lagBoth computable without knowing the true denominator, so they can be reported honestly from week one. Completeness cannot.

The digital twin, honestly

The published exemplar did not build a twin from telemetry: a sanitised natural-language specification was supplied, an agent-assisted workflow built an isolated environment, and sensors were installed into it. The engineers write “representative test environment”, never “digital twin”. The bar is far lower than the marketing implies. And Gartner explicitly permits validation against “a simulated digital twin to reduce potential impact”, which is consistent with the cloud-first cyber-range MVP the team chose over a full production-environment twin.

And the independent justification for deferring OT. Attack-path analysis in the exposure platform we would use is unsupported for OT connectors. That is a tooling fact rather than a risk appetite — worth having in the room so the FY27 deferral reads as considered scoping rather than omission.

There was no published vulnerability identifier for either way into the victim. A scanner-driven programme would have matched nothing, because what made it an intrusion was the path — and no severity score rates a path.

The five phases, and the one that matters

PhaseWhat it means here
ScopeBy business impact rather than technology silo. Crown jewels are named: critical IP and recipe repositories.
DiscoverFar wider than known-vulnerability scanning — Gartner projects that by 2028 more than half of exposure findings will be nontechnical rather than technical flaws.
PrioritiseBy attack path rather than severity. CVSS-led prioritisation drives “efforts to address only critical and high CVSS scores rather than prioritizing those riskiest to the organization.”
ValidateThe phase most programmes skip, and the one that would have caught this. Gartner’s aligned categories: adversarial exposure validation, breach and attack simulation, automated security control assessment.
MobiliseRequires resolver teams and a mobilisation process. “Without widespread business engagement most exposure management functions… are unable to function effectively.”
ONE THING WE SHOULD NOT PROMISE
Simulating the swarm. No tooling simulates a collective of independent, coordinating autonomous agents. The offensive multi-agent frameworks that exist decompose one attacker’s workflow into role-specialised stages — division of labour, not coordination. Genuine multi-agent environments exist only on the defensive side. The defensible reframing, which is also stronger: we do not need to simulate a swarm to defend against one. Remove the assumption that attacks arrive at human speed.

The one experiment to run first

Security chaos engineering against the admission tier. Admission-controller fail-open is the upstream recommendation, not a misconfiguration — and setting the policy to fail closed does not save you, because of a hard 30-second admission-chain budget. The architectural fix is in-process CEL policy evaluation. That is a complete, bounded, week-one experiment with a real finding either way.

Why “patch faster” is not the answer

FindingConsequence
The federal CVSS mandate was revoked in June 2026 in favour of a four-variable stakeholder model with 3/14/60/defer timelines — and only about 1% of instances fell in the 3-day band while roughly 60% were deferred.Severity-driven remediation targets the wrong work at scale.
The same directive says collect evidence before remediating.Remediation and forensics compete; sequence them deliberately.
No automatic rollback exists anywhere in the mainstream patch stack, and stopping a fleet takes on the order of eight hours.Longer than the entire escalation window in the reference incident. Mass-patching is not a containment mechanism.
Gartner: use mitigation controls where a patch cannot land.Mitigation is a first-class response, not a fallback.

And the risk-appetite question this defence exists to answer. Gartner’s reframe applies precisely here: stop asking how many days to patch and ask “what is our tolerance to be exploited by a known vulnerability?” Validation is what turns that question from rhetoric into a number.

WHAT THIS MEANS FOR THE PROGRAMME
Defence D moves two outcomes: % of priority controls validated by simulated attack (D1) and % of critical assets under continuous posture assessment (D4). Both are near zero for the agentic technique set, and the first is the core measure of the entire preemptive model.

The defenders’ hard part succeeded. The easy part failed.

On their own, each was ambiguous… those were correlated by our AI-based security agent stack and resolved into a coherent attack signal. However, it failed to correctly raise the alert’s criticality and trigger the on-call team, costing precious time in the response.
The victim’s own post-mortem
READ THE SEQUENCE CAREFULLY
Fusing ambiguous, low-signal events from separate systems into a single coherent attack narrative — the genuinely difficult engineering — worked. Criticality scoring did not, and nobody was paged. The stated remediation confirms the diagnosis exactly: “a high-severity signal pages a responder in minutes, any day of the week.” That phrase is doing real work — the intrusion ran Saturday 11 July to Monday 13 July. This is a governance decision, not an engineering one.

So measure escalation latency separately from detection latency

Almost every programme reports mean time to detect and folds escalation into it, which makes this class of failure structurally invisible. Separating them requires no new tooling and no new data — only a decision about how existing numbers are cut. And the related failure mode produces no alert at all: “Nothing on any screen tells you that a finding failed to appear.”

Why there is no external help for the agentic case

FindingConsequence
The reference intrusion is now a formal MITRE ATLAS case study mapping 22 techniques across 37 edges.The technique scope is externally adjudicated — we no longer argue about representativeness.
ATLAS technique objects have no detection field, no data-source field and no detects relationship. The entire published guidance is one mitigation.ATT&CK users inherit a data-source layer for free. For agentic attacks it does not exist. Building it is the work.
73 of 197 ATLAS techniques have no mitigation at all, and the two describing exactly what happened are graded Realized with zero mitigations.MITRE names the coordination channel and ships nothing against it.
Independent coverage matrix: agent-runtime telemetry touches 69 of 78 techniques, gateway telemetry 16 — all 78 mappings partial.Our confirmed telemetry is the gateway. The inversion is quantified by someone other than us.

Attribution is a proved impossibility from logs alone

arXiv:2606.09692 (Splunk / Cisco), Proposition 2.1: delegation membership from audit logs non-identifiable trace-based grouping, events recovered 4–6% time-window correlation, candidate groupings 107,718 median fragments per delegation 35–55 sub-agent fanout breaking point > 2
WHICH MOVES THE FIX UPSTREAM, TO A CHOKEPOINT WE ALREADY OWN
“Shared authority is not shared causality. Any telemetry model that conflates same trace with same unit of accountability will fail under delegated execution.” Mint a delegation identity at execution time at the gateway and MCP-proxy boundary, and propagate it. The carrier is specified rather than hypothetical: W3C baggage in MCP params._meta.

The detection primitive, and two approaches to rule out

distinct_count(behaviour_class) by delegation_id >= N behaviour_class ∈ { credential read · IMDS query · registry write · new egress destination · K8s API call · secret list · schema enumeration } a legitimate agent, per task 1–2 of these the reference case’s agents all seven
  • Do not build strict-sequence detection. MITRE: “the complexity and cost of such implementations often outweigh the benefits, and adoption… has been limited in practice.” Order-independent convergence counting survives retries and parallelism, which is how agents behave.
  • Do not detect automation by tempo. It fails at 99.8% or worse against a competent adversary and fails confidently — mean confidence above 0.993 when wrong, with a formal zero-mutual-information result. Enrich an alert with tempo; never raise one.

The highest-value detection nobody ships

An agent invoked a tool it never declared. No baseline needed. Eleven published queries exist and none joins the inventory to runtime activity; all are hunting queries rather than analytics rules, so none raises an incident. And the join key is not guaranteed — in the runtime schema the agent identifier is merely recommended and the tool name optional, so a compliant event can carry no agent identity at all. Make its presence a source-onboarding acceptance criterion.

The 30-day backlog, from telemetry already held or one switch away. Telemetry tampering — model-invocation logging deleted, guardrails deleted, registry token and user creation. Tiny volume, near-zero false positives, and they defend the detection stack itself. Autonomy-boundary removal — a published detection exists for permission overrides on Linux, and the Windows and macOS equivalents are published nowhere and must be written locally. Governed-plane bypass — egress and DNS to model providers and repositories from any source. Registry request logging — ship it and alert on repository creation and on deploys outside known CI paths, because neither the audit trail nor the available webhooks give this.

Almost everything else in this programme is harder against an autonomous adversary. This is the exception, and the evidence is unusually clean.

arXiv:2606.21037 — 21 models, 10,962 responses, 47-person human control, identical 174 reconnaissance queries: models taking deceptive bait ~78% human control group ~37% attention-diversion effect statistically absent in models recognition-action gap 73.4% — say it is a trap, exploit anyway recognition → behaviour correlation r = +0.08
SO THE STRATEGY SHIFTS FROM MISDIRECTION TO DETECTION
Three consequences, all in the same direction. A decoy credential is a higher-yield tripwire against an agent than against a person. The classic reason to deploy decoys — slowing the adversary down — largely stops working, so do not sell it that way. And trap recognition does not predict behaviour at all: an agent can know it is a honeypot and take it regardless. That is the authors’ own conclusion.

Why precision is structurally high

A decoy has no legitimate consumer, so any interaction with it is by construction either an error or an intrusion — the same structural argument as the coordination-channel discriminator, and the reason both work without a behavioural baseline. The honest limit: no published false-positive rate exists for any canary vendor, and no recall data exists anywhere. Use the structural argument rather than inventing a number.

Two things to get right before anyone buys anything

TrapDetail
Do not put a retired product in a planOne major endpoint deception capability was retired before general availability — onboarding blocked August 2025, decoys removed 31 October 2025. The identity-side honeytoken tag is generally available but covers on-premises directory infrastructure via sensors, so cloud-only accounts are not covered — which is exactly where the non-human identities live. The SIEM honeytoken solution covers a key vault only and is community-supported.
Do not use free public canariesThey are statically fingerprintable and are flagged by secret-scanning tooling as canaries without triggering an alert. Paid or self-hosted only, and diversify token types — document, DNS, kubeconfig, MCP configuration.

What deployed deception actually achieves

  • One published deployment reported 83% exploitation at around 30 minutes.
  • A vendor detected a real 2025 intrusion within seconds using canaries placed in source-control secrets at both organisation and repository level, running tens of thousands of tokens.
  • An MCP-configuration canary detects agents routing around organisational MCP restrictions — directly relevant to our own Tier 1 direct-connection prohibition, whose enforcement point is otherwise unclear.
  • A caution worth carrying: in the defensive challenge environments, hardening decoys produced only a “slight advantage” because probing for deception is itself observable — but elsewhere decoys blocked up to 94% of privilege-granting exploits. Not evidence that decoys stopped working.

And respect the published guidance on maturity. National-level trials across 121 organisations and 14 vendors recommend tripwires and honeytokens for everyone, and full interactive honeypots only for mature security operations teams. Given our baseline, tokens are in scope now and interactive honeypots are not — saying so protects the programme from a credibility problem later.

WHAT THIS MEANS FOR THE PROGRAMME
Gartner places deception squarely in the adversary-disruption domain, alongside threat hunting and detection engineering. Defence F is in tranche 1 because it needs no inventory, no gateway change and no new telemetry pipeline, and it is the fastest route to a detection that would plausibly have fired in the reference case.

Prevention is the easier half. This is the half that decides whether an intrusion costs two days or two quarters — and the half most organisations have not engineered.

Why human-paced escalation cannot survive: average adversary breakout time 29 min — down from 98, then 84, then 48 fastest observed breakout 27 sec — the tail, not the average median to hand-off to a 2nd group 22 sec — down from >8 hours in 2022
THE DESIGN QUESTION IS NOT HOW TO RESPOND BETTER
The defenders’ containment was competent and fast — find the vector, shut the service down, cut the adversary off the internal network, have the third-party platform lock the agents out. Several hours end to end. The problem was that it started two and a half days late, because nothing paged. So the question is: which parts of containment can be pre-authorised to execute without waiting for a human, and how do we make that safe.

The three tiers

TierScopeMembership test
A — autonomousMachine-identity revocation, and non-critical workload quarantine.If executing it wrongly breaks a workload rather than a person’s day, and it can be reissued, it qualifies.
B — cappedPermitted autonomously up to a blast-radius ceiling, then escalated.The cap is the control, and it must be expressed in scope rather than in count.
C — two-humanAnything that stops production, touches the production environment, or cannot be reversed.OT sits entirely here and remains deferred to FY27.

Revocation alone is not a containment strategy

If the adversary holds signing material it can mint valid credentials faster than we withdraw them — which is what happened. So the catalogue must include authority-level actions, not only credential-level ones.

MechanismWhat it does
Issue-time cohort revocationInvalidate every credential issued in a window rather than chasing individuals — via the token-issue-time condition key.
Lease revocation by prefixWithdraw a whole class of dynamically-issued secrets in one action.
Signing-authority taintThe only response to stolen signing material. Without it, revocation loses to an adversary who can issue.
Continuous session signalsPush revocation to relying parties rather than waiting for token expiry — final specifications published September 2025.
Label-swap quarantine and GitOps pruneIsolate a workload and restore declared configuration without a rebuild.

Recovering the agent estate itself

Disaster recovery covers data and applications. It rarely covers agent definitions, system prompts, tool registries, MCP configuration, vector-store contents, memory stores or gateway policy. Two consequences: without a known-good configuration you cannot prove an agent is clean, and the fastest route back to service is therefore to redeploy the compromised one.

And containment has a cost that must be priced in. The only serious published evaluation of autonomous containment prices the defender’s own collateral damage into an availability-weighted reward — a restore action carries a negative score. A containment action that halts production is not a successful containment. It is also why the automation decision is taken separately per asset class rather than once for the estate.

AND GARTNER SAYS THIS IS EXACTLY THE HARD PART
“it is extremely rare for organizations to be willing to automatically remediate discovered issues due to the concern over potential disruption… Without a cultural shift, many CISOs will be unable to fully utilize preemptive cybersecurity approaches.” We are asking for that shift, scoped to the reversible asset class only — the narrowest defensible version of the ask, and the reason the Tier A list is changeable only by the Security Committee.

Every defence, the domain it serves, the outcomes it moves, and its current state. A capability that moves no measured outcome fails Gartner’s prioritisation test and should not be funded.

DefenceDomainOutcomes it movesToday
AWhere we standFScope statement for all five — no outcome of its ownDone
BThe AI LabFForensic independence (binary); response-workflow capacityMIL0
CKnow the groundD1Asset reconciliation; crown-jewel blast radiusMIL1
DTest & fix at speedD1 D4Controls validated by simulated attack; assets under continuous posture assessmentMIL1
EDetect & escalateD2Scenarios converted to requirements; techniques with a named telemetry sourceMIL0
FDeceiveD3Deception coverage; incidents first surfaced by a disruption controlMIL0
GRespond & recoverD3 FCredential classes revocable in ten minutes; known-good agent configurationMIL0
FOUR OF SEVEN SIT AT MIL0, AND THAT IS THE PLAN’S SHAPE
B, E, F and G are all at zero — and three of those four need no new technology. The AI Lab is a vetting exercise plus a small deployment; escalation measurement is a reporting change; deception is token placement; revocation authority is a decision. That is why tranche 1 can move four outcomes without new money.

Three framings the evidence settled

OriginallyReframed toBecause
“Build a digital twin of the estate”A cloud-first cyber range as an MVP, plus attack-surface managementThe published exemplar used a sanitised natural-language specification and an isolated representative test environment, not telemetry-derived replication — and the team judged a full production-environment twin impractical. Gartner permits validation against a “simulated digital twin” for exactly this narrow purpose.
“Simulate the agent swarm”Remove the assumption that attacks arrive at human speedNo tooling simulates a coordinating collective; offensive multi-agent frameworks are division of labour, not coordination.
“Deception slows the adversary down”Deception as a detection controlThe attention-diversion effect is statistically absent in models and trap recognition does not predict behaviour. Selling it as delay would be selling the one benefit the evidence removes.

The gap the mapping exposes

D2 is covered by one defence, and nothing in the seven addresses deepfake-enabled social engineering — the highest-prevalence real AI attack at 41% of organisations on audio calls and 35% on video. They were designed against an infrastructure intrusion. The SOP-hardening work Gartner specifies — password reset, partner and supplier interactions, financial transactions — closes it in tranche 1 at essentially no cost.

WHAT THIS MEANS FOR THE PROGRAMME
Present the mapping rather than the list. It demonstrates that the technical work already done serves the strategy rather than sitting beside it, shows which outcome each defence moves, and makes the one real gap visible instead of hiding it.

The structural change is not that attacks are cleverer. It is that the interval between a vulnerability becoming known and being weaponised has collapsed from days to minutes, because frontier models can autonomously reverse-engineer a patch into a working exploit.

BUT READ GARTNER’S NEXT SENTENCE, BECAUSE IT IS THE HONEST ONE
“attackers have always been faster at generating exploits than defenders are at applying patches… The core issue is not a new, apocalyptic threat, but a long-standing reality: manual vulnerability management processes are structurally misaligned with machine-speed attacks. The model was an attention catalyst, not a new capability.

What speed actually defeats

Our processIts assumed timescaleWhat the adversary now needs
Patch cycleWeeks, gated by change windows and downtimeMinutes to produce a working exploit
Analyst triageHours, business-day weightedIn the reference case the intrusion ran Saturday to Monday and nothing paged
Fleet-wide remediation~8 hours to stop a fleet, and no automatic rollback exists anywhere in the mainstream patch stackLess than the escalation window
Containment approvalA change ticketAverage breakout 29 minutes; fastest observed 27 seconds

The number that ends the human-paced-escalation argument

Adversary tempo, recent reporting cycles: average breakout time 29 min — was 98, then 84, then 48 fastest observed 27 sec — the tail is what a design must survive median to second-stage 22 sec — was >8 hours in 2022

And one piece of good news worth carrying to the board. “the next 12 to 24 months will likely see an increase in the aggregate volume of vulnerabilities… the medium-term outlook is a positive one. Frontier AI models will enable defenders to inspect an unprecedented volume of source code.” The speed problem is transitional, not permanent — which argues for process change now rather than panic buying.

WHAT THIS MEANS FOR THE PROGRAMME
Speed is answered by D3 adversary disruption (pre-authorised revocation, deception that fires early) and by measuring escalation latency separately from detection latency. It is not answered by patching faster, which Gartner explicitly warns against promising.

Gartner’s assessment of what AI is actually doing for attackers today: “threat actors leveraging and abusing widely available AI tools to scale and enhance existing tactics, not the invention of net-new attack methods.” The consequence is a volume problem, and volume breaks controls that were tuned for plausible attempt counts.

What scale looked like in the reference case

MITRE ATLAS AML.CS0068, procedure narrative: agent runs ~1,200 coordination messages >70,000 attacker actions inside victim ~17,600 across ~700 runs
WHY THIS DEFEATS THRESHOLD-BASED CONTROLS
An alert threshold calibrated to human patience assumes an attacker tries a handful of things and moves on. Seventeen thousand actions inside one victim is not a louder version of that — it is a different distribution. The reference victim’s own account is that volume was the defeating factor: correlation succeeded, but criticality scoring did not raise it above the noise floor.

The detection consequence, and the approach that survives it

Sequence-based detection fails because parallel agents retry and reorder constantly. MITRE is explicit that strict sequencing “can be challenging to implement effectively… adoption of these types of analytics have been limited in practice.” What survives is order-independent convergence counting:

distinct_count(behaviour_class) by delegation_id >= N a legitimate agent, per task 1–2 behaviour classes the reference case’s agents all seven

And the approach that does not survive it. Detecting automation by its tempo is the intuitive response to a scale problem, and it fails at 99.8% or worse against a competent adversary — and fails confidently, with mean confidence above 0.993 when wrong. Use tempo to enrich an alert; never to raise one.

WHAT THIS MEANS FOR THE PROGRAMME
Scale is answered by D1 exposure management — if the adversary will try everything, the defensible objective is knowing what is reachable and bounding the consequence, not blocking every attempt.

Fifty-four attack vectors across ten layers of the estate: agents and models, gateways, identity, cloud and compute, network and egress, data and knowledge, software supply chain, endpoints and collaboration, observability and recovery, and manufacturing. Four have no named owner.

GARTNER’S VERSION OF THE SAME POINT
“Cybersecurity leaders must innovate in their practices as their organizations are creating attack surfaces faster than technologies can protect them.” And looking forward: “By 2028, more than half of threat exposure findings will result from nontechnical vulnerabilities, rather than technical flaws.”

The layer nobody owns, and why it matters most at the enterprise

SEMI E188 explicitly excludes the manufacturing execution system, the material-control system and factory-provided host systems; E187 addresses supplier-provided equipment. Neither standard covers the factory-software layer that actually holds write authority into the tools — and that layer is the one reachable from IT. Gartner describes the same gap from the infrastructure side: “Fragmented IT/OT convergence creates severe risks, as current systems lack unified governance and struggle to support the downtime required for constant patching.”

Two layers our own defences did not cover

LayerWhy it was missedEvidence
Endpoint and collaborationOur seven defences were designed against an infrastructure intrusion, so the human-facing surface was out of frame.Deepfake-enabled social engineering at 41% audio / 35% video is the highest-prevalence real AI attack
Observability and recoveryThe systems that watch and restore the estate are targets because they are the detection capability.An agent holding backup or DR automation scope can neutralise recovery as a legitimate action

The reframe that makes breadth fundable. Breadth cannot be closed by buying more controls — there are too many layers and four have no owner. It is closed by enumerating what exists and naming who is accountable, which is why D1 precedes everything and why the board ask for distributed ownership is the programme’s critical path.

Gartner reproduces the 2026 Verizon breach dataset under a heading that is itself the finding: “The Majority of Breaches Still Do Not Exploit Vulnerabilities.”

2026 Verizon DBIR, via Gartner: vulnerability exploitation, initial access 31% — rising credential abuse, initial access 13% credential abuse, anywhere in the chain 39% phishing, initial access 16% — stable
WHY THIS BELONGS IN A STRATEGY ABOUT AI ATTACKS
Because the AI conversation is dominated by exploit generation, and exploit generation addresses the 31%. Credentials and phishing together account for more initial access than vulnerabilities do. Gartner’s instruction follows directly: “Top threats are cyclical. To truly support defense in depth, CISOs must prioritize among many threats across the attack kill chain, not only the most discussed threat of the year.”

What this changes in our allocation

  • Identity is the highest-leverage surface, not vulnerability management. Credential abuse appears in 39% of breach chains, and machine identity is where our maturity is lowest.
  • The human-facing surface earns real investment. Phishing is stable at 16% and deepfake-enabled social engineering is measured at 41% and 35% — and the mitigation is process change, not technology.
  • Vulnerability work shifts from patching to exposure. CVSS-led prioritisation drives “efforts to address only critical and high CVSS scores rather than prioritizing those riskiest to the organization.”

And a forward-looking caution on where findings will come from. “By 2028, more than half of threat exposure findings will result from nontechnical vulnerabilities, rather than technical flaws, requiring a fundamental shift in security priorities as these risks surpass traditional IT concerns.”

This is the claim our strategy is most exposed on, so it is stated conservatively and sourced precisely.

What Gartner says

  • “Uses of LLMs to assist in malware creation, as part of the malware workflow, or to orchestrate automated attacks are emerging with unclear impact so far.”
  • Threat actors integrating models into malware show a “lack of sophistication” and are “more experimental than mature”.
  • “today, true AI-powered attacks remain very rare in the real world. Rather, adversaries are using AI to scale traditional phishing, automate tasks, and compensate for skill gaps.”
  • AI-augmented attacks are classified in the “unpredictable threats” section of the 2026-2027 ThreatScape, which maps threats by signal quality and threat-actor advantage.

What has nonetheless been observed

ObservationStatus
A frontier lab disrupted a state-attributed campaign against ~30 targets in which AI performed 80-90% of the work — “the first documented case of a large-scale cyberattack executed without substantial human intervention.”Single-vendor self-report, not independently audited. Not sector-specific.
MITRE has published the reference intrusion as a formal case study, AML.CS0068, mapping 22 techniques across 37 edges.Externally adjudicated. This is the strongest evidence available.
ATLAS moved from 16 tactics and 84+ techniques at v5.1.0 to v5.4.0 within roughly three months.The taxonomy is moving fast enough that any coverage baseline must record its ATLAS version.
SO WHERE DOES THE AGENTIC SCENARIO BELONG?
As a validation scenario, not a threat claim. It is real, externally adjudicated, and it stresses exactly the processes Gartner says are structurally misaligned. That earns it a place under D1 as an exposure-validation input and under D2 as scenario planning. It does not earn a place as the basis for the programme’s threat model.

And the instruction that governs how we present all of this. “CISOs should help their organization ignore sensational marketing hype around emerging AI threats and focus on actionable responses.” A strategy that overstates this gets discounted the first time a director checks it against the same research.

The most reassuring finding in the corpus, and the one most likely to be misread as a reason to do nothing.

Attackers still need to be ‘right’ multiple times: Foundational cybersecurity controls still provide adequate protection and resilience.
Gartner’s board-scenario guidance

And alongside it: “This situation is not a repeat of Y2K and does not require a transformational response.”

WHICH IS EXACTLY WHY THE PROGRAMME IS SHAPED THE WAY IT IS
Preemptive capability is additive: Gartner is explicit that these tools “must be treated as additive to a well-established and mature program; they are not substitutes or replacements for cyber hygiene, nor do they replace core skills or technology gaps.” So this strategy asks to stop doing nothing we do today — and it cannot be funded by cutting prevention or detection.

What “foundational controls still work” does not mean

It does not meanBecause
That no investment is needed“Many capabilities promoted as preemptive exist in current platforms. Combining them, however, has the potential to add resilience.” The work is combination and validation, which is effort rather than licences.
That our controls are workingNothing in our estate has been validated by simulated attack against the agentic technique set. Working and untested are different claims, and outcome D1 exists to close that gap.
That the gap will not widen“AI-driven exploit generation will continuously outpace traditional patching cycles.” Foundational controls hold today; the margin is narrowing.

The honest framing for the board. Our controls are probably adequate against the attacks of the last three years, and we cannot currently demonstrate that they are adequate against the next three. The programme buys the demonstration — which is a far cheaper ask than a transformation, and a far more defensible one.

Gartner’s definition: “tooling to test for vulnerabilities, assess how they expose your organization to breach activity, and evaluate the security controls you have in place to ascertain their effectiveness. These offerings simulate real-world attacks… against extant security controls.” The stated benefit is visibility into “where and why they fail prior to an actual attack.”

The technology categories Gartner places here

  • Adversarial exposure validation (AEV) — the core capability
  • Breach and attack simulation (BAS)
  • Automated security control assessment
  • Autonomous exposure remediation — noting Gartner’s caution that organisations are “extremely rare[ly]” willing to auto-remediate
WE EXTEND THIS DOMAIN BACKWARDS INTO INVENTORY, DELIBERATELY
You cannot validate controls over an estate you cannot enumerate. Gartner separately recommends “pull[ing] automated critical IT, OT, and cloud asset inventories from existing exposure assessment platforms, in order to prioritize where to start changing architectures.” So D1 owns discovery first, then validation.

The five-phase discipline beneath it

PhaseWhat it means here
ScopeBy business impact, not technology silo — “taking into consideration the potential business impact of a compromise rather than primarily focusing on the severity of the threat alone.”
DiscoverWider than known-vulnerability scanning. Neither entry vector in the reference case had a published vulnerability identifier, so a scanner-driven programme would have matched nothing.
PrioritiseBy attack path. CVSS-led prioritisation addresses “only critical and high CVSS scores rather than prioritizing those riskiest to the organization.”
ValidateThe phase most programmes skip — and the one that would have caught this.
MobiliseNeeds resolver teams. “Without widespread business engagement most exposure management functions… are unable to function effectively.”

Where validation runs

Against real controls, or — in Gartner’s words — “in some cases, a simulated digital twin to reduce potential impact.” That is the narrow, defensible version of a twin, and it matches the cloud-first cyber-range MVP the team chose over a full production-environment twin. The published exemplar for building such an environment used a sanitised natural-language specification, not telemetry replication — the engineers call it a “representative test environment”.

The measurement discipline that must come with it. Score detections on robustness, precision and implementation coverage rather than techniques touched — MITRE: “the goal is not simply to maximize the number of ATT&CK techniques associated with detection content.” And never publish a single coverage percentage without its partial-coverage caveat: the reference framework states its own headline as both “78 of 78 have a native analytic” and “0 direct, 78 partial”.

WHAT THIS MEANS FOR THE PROGRAMME
D1 owns three outcomes: asset and identity reconciliation, controls validated by simulated attack, and crown-jewel blast radius. The first is 90 days, the second is continuous and is the core measure of the whole preemptive model, and the third is FY27.

Gartner: “using data and context to understand likely trends from adversaries… TTPs that are in use and increasing as well as targets.” And the qualification that shapes how we fund it: “On its own, adversary management allows for scenario planning… It is of most value when combined with one of the other pillars to provide actionable enforcement.”

WHICH IS WHY OUR D2 OUTCOME IS A CONVERSION RATE, NOT A FEED COUNT
Intelligence that does not become a detection requirement with a named owner has produced scenario planning and nothing else. The measure is therefore the proportion of priority scenarios converted into a requirement, not the number of feeds subscribed to.

There is no feed to buy for this

No vendor-neutral, standards-body TTP catalogue for autonomous attackers exists beyond MITRE ATLAS and the OWASP Top 10 for Agentic Applications. The US government AI-ISAC remains policy-mandated but pre-decisional with no launch date, so interim AI-threat intelligence should route through an existing sector ISAC. So D2 is a built pipeline, not a procurement.

The pipeline, and what it produces

StagePractice
RequirementsWrite priority intelligence requirements at a “stable middle” specificity — a named technique, a named source, and the decision the answer supports. An analyst can realistically sustain three to five.
CollectionSource named techniques from ATLAS and OWASP’s agentic list. Record the ATLAS version — it moved from 16 tactics and 84+ techniques at v5.1.0 to v5.4.0 in roughly three months.
AnalysisConvert to a detection or mitigation requirement with an owner. This is the step that turns D2 from a reading exercise into a control.
DevelopmentSigma-coded, ATT&CK- and ATLAS-tagged detections, through a Requirements Discovery → Triage → Investigation → Development pipeline.
ValidationPurple-team emulation — which hands the result to D1.

The gap that makes this domain urgent rather than academic. MITRE has adjudicated the reference intrusion as a case study and publishes no detection guidance for it — no detection field, no data-source field, one generic mitigation; and 73 of 197 techniques have no mitigation at all, concentrated in the autonomous block. ATT&CK users inherit a data-source layer for free. Here it does not exist, so D2 builds it.

And the vector D2 must cover that our defences did not

Deepfake-enabled social engineering, at 41% of organisations on audio calls and 35% on video. Gartner names the processes: password reset, partner and supplier interactions, and financial transactions. This is a D2 scenario feeding a process control, not a technology purchase — and it is the cheapest item on the roadmap.

WHAT THIS MEANS FOR THE PROGRAMME
D2 is our most mature domain and currently delivers the least value, precisely because Gartner says it needs to be combined with another pillar to produce enforcement. Fix the conversion step and the existing investment starts paying.

Gartner: “accepting that attacks will occur and working to delay or misdirect such attacks… deception and honeypots that provide a clear signal of an attack when triggered, conducting proactive threat hunting… and operationalizing threat intelligence in other controls.” Three stated purposes: slow the attacker, give a clear signal, and feed attacker TTPs back into other controls.

ONE OF THOSE THREE PURPOSES DOES NOT SURVIVE CONTACT WITH AN AGENT
The slowing purpose largely fails. Across 21 models and nearly eleven thousand responses against a 47-person human control, the human attention-diversion effect is statistically absent in models, and trap recognition does not predict behaviour at all — models articulate that something is a trap and exploit it anyway 73.4% of the time. So sell D3 on signal, not on delay.
arXiv:2606.21037 — 21 models, 10,962 responses, 47-person control: models taking deceptive bait ~78% human control group ~37% recognition-action gap 73.4% recognition → behaviour r = +0.08

The second half of D3: raising cost through revocation

A stolen credential that dies in ten minutes is worth very little. Machine identities are the reversible asset class — revoking one breaks a workload rather than a person’s day, and it can be reissued — which is what makes pre-authorisation defensible here and nowhere else yet. Mechanisms: issue-time cohort revocation, lease revocation by prefix, signing-authority taint, and continuous session signals.

And Gartner names this as the hard part. “it is extremely rare for organizations to be willing to automatically remediate discovered issues due to the concern over potential disruption… Without a cultural shift, many CISOs will be unable to fully utilize preemptive cybersecurity approaches.” We are asking for that shift, scoped to the reversible asset class only.

Two procurement traps

TrapDetail
A retired product in the planOne major endpoint deception capability was retired before general availability — onboarding blocked August 2025, decoys removed 31 October 2025. The identity-side honeytoken is generally available but covers on-premises directory infrastructure via sensors, so cloud-only accounts — where the non-human identities live — are not covered.
Free public canariesStatically fingerprintable, and flagged by secret-scanning tooling as canaries without triggering. Paid or self-hosted only, with diversified token types.
WHAT THIS MEANS FOR THE PROGRAMME
D3 is our weakest domain and the cheapest to move. Deception needs no inventory, no gateway change and no new telemetry; revocation authority is a decision. National-level trials across 121 organisations recommend tripwires and honeytokens for everyone, with interactive honeypots reserved for mature teams — so tokens now, honeypots not yet.

Gartner: “Successful attacks often take advantage of failures in the configuration of your environments… As the number of vulnerabilities organizations have to manage increases due to the effect of LLM-driven vulnerability discovery… hardening configurations is even more vital than before.” The benefit: it “reduces the attack surface and minimizes the impact of an incident that does compromise a hardened device.”

The technology categories, including one we do not have

CNAPP — cloud security posture management, cloud infrastructure entitlement management, Kubernetes security posture management, and AI security posture management — plus data security posture management, network security posture management, unified endpoint management and endpoint protection.

AND THE MOST IMPORTANT LIMIT IN THIS DOMAIN, STATED PLAINLY
AI security posture management is a static configuration and inventory discipline. It cannot detect an agent behaving outside its authorised scope at runtime. As one vendor puts the distinction: “static AI-SPM tells you what an agent can do; runtime-informed AI-SPM tells you what it actually does.” So our primary D4 outcome — declared agent scope enforced at runtime — must be written as a separate, explicitly-scoped requirement and tested, not assumed to arrive with a posture-management purchase.

A procurement note that follows from market structure. The category is consolidating into CNAPP: Wiz into Google, Protect AI into Palo Alto, CalypsoAI into F5, Aim Security into Cato, Robust Intelligence into Cisco. AI-SPM will therefore arrive as a module of something the enterprise may already license — consistent with Gartner’s wider point that “many capabilities promoted as preemptive exist in current platforms.”

The segmentation journey, and why it usually fails

Gartner’s recommendation is “the journey from macrosegmentation to network security microsegmentation to limit lateral movement.” The published failure pattern is organisational, not technical:

FindingConsequence for how we phase it
Of fourteen organisations attempting even limited segmentation, eleven failed — predominantly for want of an executive champion and application-owner buy-in.Secure the sponsor and the app owners before the tooling. This is the same distributed-ownership ask as everywhere else.
Discovery first: 30 to 90 days of traffic observation before enforcement; pilot in 8-12 weeks; first enterprise segment in 3-6 months.Enforcing before dependency mapping is the usual cause of outage-driven rollback. FY27, not tranche 1.
A Kubernetes NetworkPolicy is accepted and silently does nothing unless the CNI plugin implements enforcement.Verify enforcement empirically. A policy object is not a control.
For legacy process equipment, network-enforced segmentation is the only viable primary control — agents cannot run and VLAN tagging is often unsupported.The production environment path is network-level, and it sits behind the FY27 OT scoping decision.
WHAT THIS MEANS FOR THE PROGRAMME
D4 carries three outcomes: runtime enforcement of declared scope (currently zero, 90 days), critical assets under continuous posture assessment (FY27), and untrusted-input workloads under deny-by-default egress (FY27). The first is the one no product will hand us.

Gartner’s model puts the four domains on a foundation, and is direct about why: “many organizations will require vendor- or third-party-led managed services to realize significant value from these offerings unless they have large and high-level security teams.” And: “With all preemptive cybersecurity offerings, many organizations will also benefit from a managed or supported service.”

THIS LAYER IS WHERE OUR REAL GAPS SIT
An internally-authored strategy can describe capabilities without stating what they rest on. A foundation layer makes that impossible. Ours holds three honest gaps: detection-engineering capacity is not established, forensic capability is not independent of hosted guardrails, and exposure findings have no accountable resolver outside CDR.

The three foundation outcomes

OutcomeWhy it belongs at the base
% of exposure findings with an accountable resolver outside CDROnly 36% of organisations have infrastructure teams actively engaged on remediation, and without business engagement exposure management “cannot function effectively.” Gartner escalates this to a board ask.
Forensic capability independent of hosted guardrails — binaryHosted models refused defensive work in the reference incident, measured at 2.72× refusal on security tasks, and “cannot rephrase refused queries or retry.” Cannot be procured mid-incident.
% of response workflows with defined capacity and a lifecycleGartner projects AI agents autonomously managing 25% of incident-response workflows for data security events by 2028. Automating an undefined process automates the wrong one.

The sequencing risk this layer creates

Gartner warns that preemptive offerings may be “limit[ed] to detecting, rather than fixing issues, thereby increasing, rather than reducing overhead.” If validation and posture assessment stand up before the accountable resolvers exist, the programme manufactures a findings backlog nobody owns. That is the single most likely way this creates work rather than protection — which is why the ownership decision is sequenced first.

And one caution on how far to automate the foundation. “fully automating fixes will eliminate practical learning ground required to develop experienced Level 3 analysts.” The long-run mandate Gartner describes is antifragility — organisations “intentionally us[ing] systemic shocks and controlled risk exposure to emerge operationally stronger” — which is an argument for exercises, not autopilot.

“By 2030, preemptive cybersecurity solutions will account for 50% of IT security spending, up from less than 5% in 2024, and replace traditional ‘stand-alone’ detection and response solutions as the preferred approach to defend against cyberthreats.”

USE THIS FOR DIRECTION, NOT FOR SIZING
It comes from Gartner’s Emerging Tech team, whose register is markedly more forward-leaning than the CISO-facing research — the same document warns of “an arsenal of unprecedented sophistication” and “ruinous business loss”, while the CxO Leadership team writes that “true AI-powered attacks remain very rare in the real world.” We use the first for market direction and the second for threat claims, and say so if challenged.

Three corroborating signals

SignalFigure
Current spend splitOf roughly $244.2B in 2026 security spending, about $49B goes to AI-amplified security and only $2.8B to securing AI itself — roughly a seventeen-fold imbalance.
Control gap13% of organisations reported a breach of an AI model or application, and 97% of those lacked AI access controls. Unsanctioned AI added about $670,000 to average breach cost.
Adoption vs scale88% of organisations report AI adoption but fewer than 10% have fully scaled it; 42% abandoned most AI initiatives in 2025, up from 17%.

And the market-structure consequence for procurement. The AI security category is consolidating into cloud-native application protection platforms — Wiz into Google, Protect AI into Palo Alto, CalypsoAI into F5, Aim Security into Cato, Robust Intelligence into Cisco. Buying a standalone AI security product in FY27 risks buying something that becomes a module of an existing licence.

“Many capabilities promoted as preemptive exist in current platforms. Combining them, however, has the potential to add resilience.” And the instruction that follows: CISOs “must ensure that they do not already have a similar capability being delivered as part of their existing product and service portfolio.”

Gartner’s own mapping of domains to existing technology categories

DomainCategories that may already be licensed
D1Adversarial exposure validation; breach and attack simulation; automated security control assessment; autonomous exposure remediation
D2Cyberthreat intelligence; unified cyber risk intelligence; CTI management platforms
D3Deception; firewalls and NDR with threat-intelligence integration; threat hunting; detection engineering
D4CNAPPCSPM, CIEM, KSPM, AI-SPM; data security posture management; network security posture management; unified endpoint management; endpoint protection
WHAT THIS MEANS FOR TRANCHE 1
Four of seven tranche-1 items are decisions or process changes rather than purchases: revocation authority, SOP hardening, the coverage baseline, and verifying what an existing control actually does. That is not a cost-saving accident — it follows directly from Gartner’s observation that the capability often already exists and is simply uncombined and unvalidated.

Two things we already hold that are under-used

  • The exposure-management platform. Critical-asset definition generates attack paths, and first-class connectors exist for the vulnerability scanner and the service-management CMDB. The gating input is a business workshop to say what matters, not an engineering build.
  • The gateway. It already enforces directory authentication, scope authorisation, quotas, DLP and prompt-injection inspection, with central logging. What it does not yet do is mint a delegation identity — which is the one thing no product will supply and the fix for an otherwise unsolvable attribution problem.
WHAT THIS MEANS FOR THE PROGRAMME
Before any tranche-2 procurement, run Gartner’s test: do we already have this? The answer for several D1 and D4 capabilities is probably yes-but-unconfigured, which changes the ask from budget to effort.

Gartner’s shorthand for the same model: “Deny intruders access to your global attack surface grid through advanced obfuscation techniques. Deceive bad actors through automated cyber deception and moving target defense. Disrupt attacks via predictive threat intelligence and automated exposure management.”

How the verbs map to our domains

VerbDomainWhat we actually do
DenyD4 posture and policyReduce what is reachable: runtime enforcement of declared agent scope, deny-by-default egress for the untrusted-input tier, and the macro-to-microsegmentation journey.
DeceiveD3 adversary disruptionDeception tokens where an agent reaches them first — and sold on signal, because the delay benefit does not survive against agents.
DisruptD1 + D2Predictive intelligence converted into detection requirements, and automated exposure validation.
WHY WE LEAD WITH FOUR DOMAINS RATHER THAN THREE VERBS
The verbs are a better slogan and a worse structure. They have no foundation layer, and the foundation is where our honest gaps sit — detection-engineering capacity, forensic independence, accountable resolvers. Use the verbs in the room; use the domains in the plan.

One term worth knowing if it comes up. Gartner calls the thing being defended the “global attack surface grid” — which is the same observation as our ten-layer landscape, and the same one behind “organizations are creating attack surfaces faster than technologies can protect them.”

“Current detection and response and application security methods aren’t sufficient to keep up with the speed, sophistication and scope of emerging AI-enabled threats.” The preemptive model “adds focus beyond prevention, detection, and response and seeks to improve foresight capabilities to limit an attack’s impact early, optimally preventing initial access.”

The evidence from our own estate

FindingSource
Agent-runtime telemetry touches 69 of 78 techniques; gateway telemetry touches 16 — and all 78 mappings are partial. Our confirmed telemetry is the gateway.Independent coverage matrix
Detection and correlation worked in the reference case. Criticality scoring failed and nobody was paged, across a weekend.The victim’s own post-mortem
There is no published detection guidance for the autonomous technique block, and 73 of 197 techniques have no mitigation at all.MITRE ATLAS
SO THE ARGUMENT IS ADDITIVE, NOT CORRECTIVE
This is not a case that detection was the wrong investment. It is a case that detection cannot be the only investment, because the window it operates in has collapsed and because a third of the technique space has no published detection at all. Gartner is explicit that preemptive tooling is “not substitutes or replacements for cyber hygiene, nor do they replace core skills or technology gaps.”

And the one detection investment that still pays disproportionately. Escalation, not detection. Measuring mean time to escalate separately from mean time to detect costs nothing — it is a decision about how existing numbers are cut — and it is the only way the reference failure becomes visible. Related: “Nothing on any screen tells you that a finding failed to appear.”

“AI-washing, both as a threat and as a capability, is prevalent in cybersecurity product marketing today.” Gartner is blunt about the cause: “Preemptive cybersecurity is being applied as a label to both new and existing technologies by vendors… there is AI-washing both in creating demand for and in labeling these technologies as preemptive.”

THE TEST
“CISOs must focus on the outcomes delivered and whether existing tools support these outcomes, not on the hype around AI.” Which is why every capability in this strategy is tied to a named outcome-driven metric, and why the prioritisation rule is absolute: “If there is no clear line of sight to a business outcome, then the technology investment should not be prioritized.”

Three specific claims to interrogate

Vendor claimWhat to ask
“AI-powered attack detection”Ask for the detection content and its validation results. True AI-powered attacks “remain very rare in the real world”, so ask what corpus the detection was tuned on.
“AI security posture management”Ask specifically whether it detects an agent acting outside its declared scope at runtime. Classic AI-SPM is a static configuration discipline and does not.
“Autonomous remediation”Ask what it will do without approval, and read Gartner’s caution: organisations are “extremely rare[ly]” willing to auto-remediate, which can limit these offerings to detecting rather than fixing — increasing overhead.

And the sceptical habit worth keeping about our own claims too. The same discipline applies inward. We do not report coverage we have not validated, we never publish a single coverage percentage without its partial-coverage caveat, and we state limits plainly — including that no published false-positive rate exists for any canary vendor and no recall data exists anywhere.

“The immediate danger is the CISO attempting to own the solution in a silo. If the CISO promises to simply ‘patch faster’ without the backing of the CIO, they will fail. In most organizations this is neither attainable nor sustainable.”

Why it is not attainable

ConstraintEvidence
Severity-led prioritisation targets the wrong workCVE/CVSS assessment produces “efforts to address only critical and high CVSS scores rather than prioritizing those riskiest to the organization.”
The federal mandate itself moved away from itThe CVSS mandate was revoked in June 2026 for a four-variable stakeholder model with 3/14/60/defer timelines — and only about 1% of instances fell in the 3-day band while roughly 60% were deferred.
The tooling cannot support itNo automatic rollback exists anywhere in the mainstream patch stack, and stopping a fleet takes around eight hours — longer than the entire escalation window in the reference incident.
The production environment cannot absorb the downtime“current systems lack unified governance and struggle to support the downtime required for constant patching.”
And the gap will widen regardless“AI-driven exploit generation will continuously outpace traditional patching cycles.”
WHAT WE PROMISE INSTEAD
Three things. Mitigation as a first-class response — Gartner: “Improve response time by using threat management techniques to identify and implement mitigation controls” where a patch cannot land. Exposure-led prioritisation scoped by business impact rather than severity. And a different question entirely: not how many days to patch, but “what is our tolerance to be exploited by a known vulnerability?

Plus the ownership consequence, which is the real fix. “cybersecurity teams can only guide vulnerability prioritization. IT operations, product teams and business system owners must be held accountable by the board to fix exposures in their own systems.” Only 36% of organisations have infrastructure teams actively engaged on remediation today.

The programme asks for autonomous action in exactly one place — machine-identity revocation — and stops there deliberately. Three reasons, all sourced.

1. It erodes the capability it depends on

“fully automating fixes will eliminate practical learning ground required to develop experienced Level 3 analysts.” The senior analysts a programme like this needs in three years are produced by the work it would be most tempting to automate now.

2. The defender’s own disruption is a real cost

The only serious published evaluation of autonomous containment prices collateral damage into an availability-weighted reward — a restore action carries a negative score. A containment action that halts production is not a successful containment, which is why the automation decision is taken separately per asset class rather than once for the estate.

3. Organisations do not actually do it

“it is extremely rare for organizations to be willing to automatically remediate discovered issues due to the concern over potential disruption, and this may limit preemptive cybersecurity offerings to detecting, rather than fixing issues, thereby increasing, rather than reducing overhead.”

SO THE ASK IS NARROW BY DESIGN, AND THE SAFEGUARD IS STRUCTURAL
Autonomous action is limited to the reversible asset class — revoking a machine identity breaks a workload, not a person’s day, and it can be reissued. And the Tier A list is changeable only by the Security Committee, so the scope of autonomy cannot be widened by the people exercising it. Either half without the other is not defensible.

Where Gartner expects this to go anyway. “By 2028, cybersecurity AI agents will autonomously manage 25% of incident response workflows for data security events.” Which is an argument for defining the workflow and its capacity now, while the numbers are still manual and therefore honest — that is foundation outcome F.3. The long-run mandate is antifragility: “intentionally use systemic shocks and controlled risk exposure to emerge operationally stronger.”

“CISOs Must Focus on the Outcomes of Preemptive Cybersecurity”, 5 August 2026, nine pages plus a technology appendix. It supplies the four-domain structure, each domain’s definition and core benefit, the foundation layer, and a mapping of domains to technology categories.

Its three stated challenges, all of which we inherit

  • Cultural shift required, because “predicting cyberthreat activity accurately is largely aspirational” and organisations are “extremely rare[ly]” willing to auto-remediate.
  • AI-washing is prevalent in both the demand creation and the product labelling.
  • Much of it already exists in current platforms — “Combining them, however, has the potential to add resilience.”
THE SENTENCE THAT MOST SHAPES OUR FUNDING CASE
Preemptive capability “adds focus beyond prevention, detection, and response” and is “additive to a well-established and mature program… not substitutes or replacements for cyber hygiene.” So the programme cannot be funded by cutting prevention or detection, and it does not ask to.

What we added to it

Gartner suppliesWe added
Four domain definitions and core benefitsThe enterprise-specific content for each, and the outcome-driven metrics per domain
A foundation layer described as managed servicesThree named foundation outcomes, including the honest gaps: detection capacity, forensic independence, accountable resolvers
A technology mapping tableAn assessment of which of those we already hold, and the finding that AI-SPM does not deliver runtime scope enforcement
Domain-level benefit statementsA maturity assessment per domain, and a readiness position on an investment scale

One permission in this document worth knowing about. Gartner explicitly allows exposure validation to run “against actual controls, or in some cases, a simulated digital twin to reduce potential impact.” That is the narrow use the team’s cloud-first cyber-range MVP occupies — so the twin was not rejected outright, it was scoped to the one purpose the research supports.

“Outcome-Driven Metrics for the Digital Era”, 11 July 2025, twenty-one pages. It supplies above-and-below-the-line structure, the five-step derivation method, and — in Note 1 — seven characteristics that decide whether a metric is worth reporting at all.

The definition of an ODM, which is also a filter

1. measures the continuous outcome of an identifiable investment "If you can't identify the investment, it probably isn't an ODM." 2. sits on a sliding scale — better and worse both meaningful 3. supports investment in both directions "You can lower your investment and accept that the outcome will get worse."

The five steps, as applied

StepGartnerOurs
1Three to five most obvious business processes and their technology stacksThe three agreed business outcomes, plus the crown jewels named by the business
2Business outcomes and business ODMsContinuity, IP and recipe protection, adoption speed — above the line
3Technology risks and dependencies — “What breaks in the business process if the technology breaks?”The two-control-plane assessment and the confirmed baseline
4Define technology ODMs, which behave as leading indicatorsFifteen metrics in “% of” form, three per domain plus three foundation
5Assess readiness as risk to business enablement, on a scale from no investment to leading edgeThe readiness scale, with current position marked per domain
AND THE RULE THAT DECIDED HOW MANY REACH THE COMMITTEE
“You only need five to nine metrics for each target audience. Do not report everything you know.” Combined with “only above-the-line metrics should be shared with executives, that is why fifteen run the programme and seven reach the Security Committee.

The other five characteristics, used as filters

  • Metric value — “its ability to influence decision making… typically priorities and investments.” Anything that could not change one was cut.
  • Leading indicators — must “expose a problem before it leads to material loss.” This is why validation coverage beats incident counts.
  • Direct line of sight — “If the technology metric changes, it indicates a change in the business outcome.” Every metric on the page states its line of sight.
  • Metric changes drive action — green to yellow to red must trigger a change of priority or investment.
  • Discrete audiences — “The CFO needs different metrics from those required by the head of a business unit or a board of directors.”

And the honest cost the document names. “measuring some of these elements may require instrumenting parts of the infrastructure to gather new types of data… Gartner believes the visibility and power to report benefits and to guide priorities… will make the initial investment worthwhile.” Two of our fifteen need exactly that, and they are the two with the largest payoff: reconciliation, and runtime enforcement of declared scope.

Step 5 of the method: “Create the business enablement risk scale with, at one end, no investment or technology stack to enable the business outcome; at the other, investment in leading-edge technologies that would elevate the current technology stack to drive the most positive outcome for the business.”

WHAT THE SCALE IS FOR
It converts “we are immature” into “here is the protection level we are currently buying, and here is what more would buy.” Gartner: “By placing IT readiness on a business enablement risk scale, the CIO creates a context for IT investment discussions to help support business decision making and prioritization.” Maturity language invites a judgement about the team; readiness language invites a decision about investment.

Readiness, defined

“technology that operates the way it is designed, fully supports a business process, meets compliance requirements and is not unreasonably at risk of failure — and separately, “Readiness describes the state of investments and capabilities to manage [the risk].” So a low position is a statement about investment, not about competence.

Reading our five positions

DomainPositionWhy there
D2 Adversary managementHighestThreat modelling and scenario work are genuinely mature. But Gartner notes the domain “is of most value when combined with one of the other pillars” — so our strongest domain currently returns the least.
F FoundationMiddleGovernance is strong — a dedicated board-level Security Committee already owns this remit — while capacity, forensic independence and resolver accountability are not established.
D1 Exposure managementLowA self-declared registry exists; no controls have been validated by simulated attack against the agentic technique set.
D4 Posture and policyLowCloud posture management is partial, AI-SPM is absent, segmentation is macro-level, and runtime enforcement of declared scope is zero.
D3 Adversary disruptionLowestNothing placed, no revocation authority defined, no target time, no rehearsal. And the cheapest of the five to move.

The shape of the answer is the recommendation. D3 sits lowest and costs least to improve — deception needs no inventory, no gateway change and no new telemetry, and revocation authority is a decision rather than a purchase. That is why tranche 1 invests there rather than in the domain where we are already strongest.

All three measure the same underlying question in different registers: do we know what we have, and have we tested whether it holds?

MetricInvestmentHow it is computed, in outline
D1.1 % of discovered AI assets and machine identities that are fully reconciledAsset discovery tooling plus the exposure platform’s inventory connectorsDiscovered population as the denominator, reconciled-and-owned as the numerator. Not completeness — that needs a true total we do not have. Report alongside median declare-lag.
D1.2 % of priority adversary techniques with at least one defence proven by executing the attack in the last 90 daysAdversarial exposure validation or breach-and-attack simulation, plus the cyber-range MVPExternally-scoped technique set as the denominator; techniques with a detection that passed an adversarial test as the numerator. Never published without its partial-coverage caveat.
D1.3 % of crown-jewel data stores with a current, machine-generated blast-radius mapAttack-path analysis in the exposure platformCrown jewels are named: critical IP and recipe repositories. Numerator is those with destinations-reachable-per-credential computed and current.

The outline above is not the specification. Each metric’s unit of count, denominator, pass test and fail modes are in its own panel, reachable from the Definition and pass test link beside it on the programme set.

THE ONE TO WATCH
Validated coverage is the core measure of the entire preemptive model — Gartner’s stated benefit for this domain is visibility into “where and why they fail prior to an actual attack.” For the autonomous technique block the honest answer today is zero, which is an indisputable baseline that can only improve.

Intelligence only counts once it becomes enforcement. Gartner: adversary management “is of most value when combined with one of the other pillars to provide actionable enforcement.”

MetricInvestmentHow it is computed, in outline
D2.1 % of named high-risk procedures that pass a live social-engineering test with all four verification controls in placeProcess redesign — no technology purchaseDenominator is the named high-risk processes: password reset and MFA reset, partner and supplier interactions, financial transactions. Numerator is those with out-of-band callback to a pre-registered channel and dual authorisation above a threshold.
D2.2 % of priority threat scenarios converted into an implementable requirement with a named individual ownerThreat-intelligence analyst time and the detection-engineering pipelinePriority intelligence requirements raised as denominator; those with a detection or mitigation requirement and an accountable owner as numerator. An analyst sustains three to five PIRs.
D2.3 % of priority techniques with a named log source that is collected today and carries the required fieldsMinimum-telemetry-requirements analysisPer technique, whether a log source is named and shipping. The Dependencies column of that artefact is where the brokered asks land.

The outline above is not the specification. Each metric’s unit of count, denominator, pass test and fail modes are in its own panel, reachable from the Definition and pass test link beside it on the programme set.

Why the deepfake metric leads this domain. It addresses the highest-prevalence real AI attack — 41% of organisations on audio calls, 35% on video, and one in five biometric fraud attempts — and it is the cheapest item on the roadmap because it is process change. Detection tooling is not the answer here: commercial deepfake detectors lose 45-50% of their accuracy on realistic in-the-wild content, liveness certification structurally excludes the injection attacks actually used, and untrained humans catch deepfakes almost never.

Three metrics, matching Gartner’s three stated purposes for the domain minus the one that does not survive against agents.

MetricInvestmentHow it is computed, in outline
D3.1 % of crown-jewel environments with at least one live, monitored, tested decoy on the intruder’s pathPaid or self-hosted canary tokens — not the free public serviceCrown-jewel environments as denominator; those with diversified token types placed as numerator. Token types: document, DNS, kubeconfig, MCP configuration.
D3.2 % of machine-identity credential classes revocable cohort-wide in under ten minutes, measured at the resourceIdentity lifecycle work, plus cohort-revocation mechanismsCredential classes enumerated as denominator; those with a named authority, a rehearsed procedure and a measured time as numerator. Currently undefined, so the first measurement is also the first improvement.
D3.3 % of confirmed incidents whose first signal, in the post-incident timeline, came from a disruption controlNo incremental investment — an attribution field on incident recordsConfirmed incidents as denominator; those whose first signal came from deception or threat hunting as numerator. This is the metric that tests whether D3 is working rather than merely deployed.

The outline above is not the specification. Each metric’s unit of count, denominator, pass test and fail modes are in its own panel, reachable from the Definition and pass test link beside it on the programme set.

AND THE HONEST LIMIT ON THE FIRST METRIC
No published false-positive rate exists for any canary vendor, and no recall data exists anywhere. So coverage is reported as coverage — not as efficacy — and the case rests on the structural argument that a decoy has no legitimate consumer, which is why precision is high by construction.

Reduce what is reachable, and make declared scope mean something at runtime.

MetricInvestmentHow it is computed, in outline
D4.1 % of production agents whose registry-declared scope is refused at runtime when exceededGateway and registry integration — a build, jointly with AI EnablementProduction agents as denominator; those whose registry entry is bound to an enforcement point as numerator. Currently zero. Must be written as an explicit requirement and tested — AI-SPM is a static discipline and does not deliver it.
D4.2 % of critical assets assessed against a named standard at least weekly, with drift raised as an owned findingCNAPP modules already partly licensed, plus AI-SPMCritical assets from the inventory as denominator. For OT, note that assessment is achievable where response authority is not — the FY27 deferral is about action, not visibility.
D4.3 % of untrusted-input workloads with egress denied by default and the denial proven by testNetwork policy work, scoped to the parsing tier onlyWorkloads that parse externally-sourced input as denominator. Scoping it narrowly is what makes the cost bounded — and verification must be empirical, since a Kubernetes NetworkPolicy silently does nothing without an enforcing CNI plugin.

The outline above is not the specification. Each metric’s unit of count, denominator, pass test and fail modes are in its own panel, reachable from the Definition and pass test link beside it on the programme set.

And the segmentation journey these sit inside. Gartner’s recommended path is macro- to microsegmentation to limit lateral movement. Phase it discovery-first: 30 to 90 days of observation before enforcement, a pilot in 8-12 weeks, a first enterprise segment in 3-6 months. And expect the failure mode to be organisational — eleven of fourteen organisations attempting segmentation failed, mostly for want of an executive champion and application-owner buy-in.

Gartner rests the four domains on managed services and capabilities, noting most organisations need partner support “unless they have large and high-level security teams.”

MetricInvestmentHow it is computed, in outline
F.1 % of deduplicated exposure findings with a named individual outside CDR who has accepted or rejected them on the clockGovernance — a board decision, not a purchaseFindings raised as denominator; those with a named accountable owner outside CDR as numerator. Only 36% of organisations have infrastructure teams actively engaged on remediation.
F.2 Forensic analysis capability that does not depend on a hosted model’s guardrails — reported as met, partly met or not metA vetting exercise plus a small self-hosted deploymentYes or no. Currently no. Hosted models refuse defensive work at 2.72× the neutral rate and “cannot rephrase refused queries or retry.” Every evidence-touching inference must record model hash, engine version and decode parameters.
F.3 % of named response workflows with a stated concurrency limit, overflow behaviour and an exercise in the last 12 monthsDetection-engineering process definitionWorkflows as denominator; those with stated capacity and a requirement-to-validation lifecycle as numerator. Gartner projects AI agents autonomously managing 25% of IR workflows for data security events by 2028 — define it before automating it.

The outline above is not the specification. Each metric’s unit of count, denominator, pass test and fail modes are in its own panel, reachable from the Definition and pass test link beside it on the programme set.

THE SEQUENCING RISK THIS LAYER CREATES
Gartner warns preemptive offerings may be “limit[ed] to detecting, rather than fixing issues, thereby increasing, rather than reducing overhead.” Stand up validation before the resolvers exist and the programme manufactures a findings backlog nobody owns. That is why the ownership decision is sequenced first, not last.

Four objectives, all owned outside CDR. Roughly 20% of the relevant controls sit with us, so writing these as FTD deliverables would produce a plan that reports as late against work the programme was never able to do.

ObjectiveCounterpartyWhat FTD contributes
Reduce standing authority — credential lifetimes, scope, no ambient inheritanceIdentityBlast-radius analysis showing which credentials to shorten first
Author the isolation standard, then segmentPlatform engineeringThe standard, and the attack paths it must interrupt. Phase discovery-first: 30-90 days observation before enforcement.
Bind the agent registry to runtime enforcementAI Enablement + Security ArchitectureThe detection that proves the binding holds. No posture-management product supplies this.
Declare the untrusted-input tier; deny-by-default egress for it alonePlatform + networkScoping the tier so cost lands only where justified
THE PUBLISHED EVIDENCE ON WHERE NOT TO Platform EngineeringND FIRST
The most experienced published red team, after 80+ operations across 100+ products: “many impactful failures come from simple techniques, human creativity, and system integration issues rather than exotic ML attacks.” No adversarial-ML tooling is proposed for FY27.

One design consequence that is not a control choice

“if an LLM is supplied with untrusted input, it will produce arbitrary output.” Indirect prompt injection is a property of the technology, not a defect awaiting a patch. That makes the untrusted-input tier an isolation requirement rather than a filtering one — cheaper and more durable than any inspection tier, and consistent with Gartner’s related position that the future of AI security is in securing agent actions rather than prompts.

And the closing argument that has landed best with this audience. Five of the six capabilities this programme needs already exist at the enterprise for human adversaries. Asset inventory, risk analysis, segmentation, identity hygiene and egress instrumentation all exist in some form. The work is extending them to an actor that tests thousands of paths in parallel. Only pre-authorised response is genuinely new — and it is a governance decision rather than a purchase.

The centre of gravity, for a specific and checkable reason: MITRE has adjudicated the reference intrusion as a formal case study and publishes no detection guidance for it — no detection field, no data-source field, and one generic mitigation.

The sequencing that is not optional

inventoryvisibilityattributiondetectionescalationvalidation attempting these out of order is the most common way programmes like this produce dashboards instead of defence

The four things to build, in order

BuildWhy it comes when it does
The coverage baselineA measurement, not a build — one analyst, no new tooling, using MITRE’s published five-column schema. Its Dependencies column is where the brokered asks land.
The inventory to join toDiscovery first, attestation last. The only published methodology puts developer attestation fourth of four.
Delegation identity at the gatewayBecause attribution is formally non-identifiable from logs — trace-based grouping recovers 4-6% of a delegation’s events, and sub-agent fanout above two is the breaking point. Carrier is specified: W3C baggage in MCP params._meta.
Declared-versus-observed detectionNeeds no baseline. Eleven published queries exist and none joins the inventory to runtime activity.
AND TWO APPROACHES TO RULE OUT BEFORE SOMEONE FUNDS THEM
Strict-sequence detection — MITRE: “the complexity and cost of such implementations often outweigh the benefits, and adoption… has been limited in practice.” And tempo-based automation detection, which fails at 99.8% or worse and fails confidently, with mean confidence above 0.993 when wrong.

The 30-day backlog, from telemetry already held or one switch away. Telemetry tampering (model-invocation logging deleted, guardrails deleted, registry token and user creation) — tiny volume, near-zero false positives, and it defends the detection stack itself. Autonomy-boundary removal — a published detection exists for Linux permission overrides; Windows and macOS equivalents are published nowhere and must be written locally. Governed-plane bypass via egress and DNS to model providers. Registry request logging, because neither the audit trail nor the available webhooks give it.

The enterprise has nothing defined for machine-identity revocation — no named authority, no target time, no rehearsal. Combined with the fact that it costs nothing to fix, that makes this the cheapest material improvement available to the programme.

The three tiers

TierScopeMembership test
A — autonomousMachine-identity revocation; non-critical workload quarantineIf executing it wrongly breaks a workload rather than a person’s day, and it can be reissued, it qualifies
B — cappedPermitted autonomously to a blast-radius ceiling, then escalatedThe cap is the control, expressed in scope rather than count
C — two-humanAnything that stops production, touches the production environment, or cannot be reversedOT sits entirely here and remains deferred to FY27
REVOCATION ALONE IS NOT A CONTAINMENT STRATEGY
If the adversary holds signing material it can mint valid credentials faster than we withdraw them — which is what happened in the reference case. So the catalogue must include authority-level actions: signing-authority taint, and invalidation by issue-time cohort rather than per credential. Plus lease revocation by prefix and continuous session signals to push revocation to relying parties rather than waiting for token expiry.

Why the automation ask is narrow, and what makes it safe

Gartner: organisations are “extremely rare[ly]” willing to auto-remediate, and “without a cultural shift, many CISOs will be unable to fully utilize preemptive cybersecurity approaches.” We are asking for that shift, scoped to the reversible asset class only — and the safeguard is structural: the Tier A list is changeable only by the Security Committee, so the scope of autonomy cannot be widened by those exercising it.

And containment has a cost that must be priced in. The only serious published evaluation of autonomous containment prices the defender’s own collateral damage into an availability-weighted reward — a restore action carries a negative score. A containment action that halts production is not a successful containment, which is why the automation decision is taken separately per asset class.

Four objectives. One of them cannot be acquired mid-incident, which is why it sits in tranche 1.

ObjectiveThe finding behind it
A forensic capability that still works mid-incidentHosted models refused the disaster-recovery work in the reference case. Measured at 2.72× refusal on defensive tasks, 43.8% on system hardening — and stating you are authorised makes it worse.
Known-good configuration for the agent estateDR covers data and applications, rarely agent definitions, prompts, tool registries, MCP configuration, vector stores or gateway policy. Without it you cannot prove an agent is clean, so the fastest route back to service is to redeploy the compromised one.
Pre-authorised reversible recoveryAverage adversary breakout is 29 minutes, fastest observed 27 seconds. Anything needing a change window will not execute in time.
Continuous validation as the assurance loopCloses back into D1. It is how every claim in this strategy stops being a claim.
THE CLAUSE THAT MATTERS MOST FOR AN AUTOMATED PROGRAMME
“critical for autonomous defensive agents, which cannot rephrase refused queries or retry.” A human analyst works around a refusal in seconds; a response pipeline stops. The refusal problem therefore scales with exactly the automation this programme is building — and Gartner projects AI agents autonomously managing 25% of incident-response workflows for data security events by 2028.

The four arguments for our own capability, none of which is a capability claim

  • Availability — the tool must work at 3am on the worst day.
  • Confidentiality — the classifier that refuses is the same system that flags the session for retention and human review, so the payload you most need help with is retained longest. 69% of organisations are concerned about control over AI models.
  • Legal and export exposure — the safe harbour covers encrypted transit, but inference decrypts and processes, and no regulator has ruled on whether that is a deemed export.
  • Forensic defensibility — standards require repeatability, and hosted models update and reroute silently.

One design requirement that is free now and impossible to retrofit. Every inference touching evidence must record model hash, engine version and decode parameters into the case record. Retrofitting this invalidates the case work already done.

The seven defences were designed against a real agentic intrusion, before the Gartner frame was adopted. Mapping them onto the four domains is therefore a genuine test of whether the technical work serves the strategy — and it mostly does.

DomainCoverageBy which defences
D1 Exposure managementWell coveredC Know the ground (inventory, attack paths, blast radius) and D Test & fix at speed (CTEM, validation)
D2 Adversary managementThinE Detect & escalate only — and nothing at all for deepfake-enabled social engineering
D3 Adversary disruptionWell coveredF Deceive and G Respond & recover (revocation as a cost-raiser)
D4 Posture and policyPartialWithin D — but the runtime enforcement binding is a gap, and no product supplies it
F FoundationCoveredA Where we stand (the assessment) and B The AI Lab (forensic independence)
THE GAP, AND WHY IT IS NOT A FLAW IN THE DEFENCES
The seven were designed against an infrastructure intrusion, so they cover exposure, disruption and foundation well and the human-facing surface not at all. Deepfake-enabled social engineering is the highest-prevalence real AI attack — 41% of organisations on audio calls, 35% on video, and one in five biometric fraud attempts — and it is also the cheapest thing on the roadmap, because the mitigation is process change.

What closing that gap actually involves

Not detection tooling. Commercial deepfake detectors lose 45-50% of their accuracy on realistic in-the-wild content, liveness certification structurally excludes the injection attacks actually used against identity verification, provenance metadata is destroyed by any screenshot or re-encode, and untrained humans catch deepfakes almost never. What works is procedural: callback to a pre-registered channel rather than one the caller supplies, a passphrase that never travels over the channel being verified, and dual authorisation above a value threshold.

Two incidents that frame the ask. Arup lost a confirmed US$25.6M after an employee joined a video call on which every other participant was synthetic. And MGM Resorts was compromised in roughly ten minutes through a help-desk MFA reset with an impact exceeding $100M — using no synthetic media at all. The second is the more important one for us: the process weakness is exploitable without any AI, and AI simply makes it cheaper to attempt at scale.

Three measurements and one piece of good news. All four are chosen because they cannot be flattered.

NumberWhat it isWhy it is trustworthy
16 of 78Techniques touched by the telemetry modality we have — gateway. The modality we lack, agent-runtime, touches 69.Generated by an independent open-source coverage matrix, not by us. And all 78 mappings are partial — “0 direct, 78 partial”.
0Controls validated by simulated attack against the agentic technique set.Binary. D1’s core measure, and Gartner’s stated benefit for the domain is visibility into where controls fail before an attack.
0%Production agents whose declared scope is enforced at runtime.Nothing binds the registry to enforcement, and posture-management tooling does not supply it.
1A dedicated board-level Security Committee already owning this remit.From the enterprise’s own filed Item 1C disclosure. Most peers route this through an audit or technology-risk committee.
THE INVERSION, STATED ONCE
Our confirmed telemetry sits on the plane with the strongest preventive mediation — and that is the plane the reference intrusion would never have crossed. Every stage of it ran on the substrate plane, where there is no behavioural analytics, no provenance attestation and no inter-agent channel detection. That is a coverage argument, not a maturity confession, which is why it is the right thing to open with.

And the discipline that must accompany the first number. Never publish a single coverage percentage without its partial-coverage caveat. The framework we borrow from states its own headline twice and honestly — every technique has an analytic, and every mapping is partial. Both are true. If any technique on our dashboard ever shows full coverage, the dashboard is wrong.

Seven actions. Four are decisions or process changes rather than purchases — which follows from Gartner’s observation that many preemptive capabilities already exist in current platforms.

ActionWhy now
D3Place deception in crown-jewel environmentsNeeds no inventory, no gateway change, no new telemetry. Highest-evidence control available; use paid or self-hosted tokens, never the fingerprintable free service.
D3Define machine-identity revocation authority; time one revocationA decision, not a build. Currently undefined. Moves D3 off zero without procurement.
D2Harden SOPs for password reset, supplier and financial processesThe highest-prevalence real vector, and the mitigation is procedural: pre-registered callback, passphrase off-channel, dual authorisation.
D1Publish the control-coverage baselineOne analyst, MITRE’s published schema, no new tooling.
D1Verify what the gateway AI control actually logs and alerts onOur only existing AI-specific control, and its behaviour is unverified. May change tranche 2 scope.
FStand up forensic capability independent of hosted modelsCannot be procured mid-incident. No dependency on inventory, gateway or SIEM.
D2Inventory approved and rogue client-side GenAI tools via endpoint and network controlsGartner’s explicit action, using incumbent EPP, EDR and security service edge.
WHY THIS IS EVIDENCE-GATHERING RATHER THAN A REDUCED PROGRAMME
Two items are questions: what the existing control does, and what the SIEM ingests. Either could materially change tranche 2. Funding tranche 1 and gating tranche 2 on its findings is sequencing, not hedging — and it lets the programme demonstrate delivery before asking for anything substantial.

Four outcomes move in these thirty days. Deception coverage, revocation time, deepfake SOP hardening, and forensic independence. Three of the four are currently at zero and one is binary-no — which is why they are the fastest and least disputable wins available.

Seven actions, and this is where instrumentation spend starts. Gartner is candid that some measurement “may require instrumenting parts of the infrastructure to gather new types of data.”

ActionDepends on
D1Pull automated IT, OT and cloud asset inventories from the exposure platformBusiness workshop to define critical assets — without it the attack-path view may be empty
D1First adversarial exposure validation run against priority controlsThe coverage baseline from tranche 1
D2Convert priority scenarios into detection requirements with ownersAnalyst capacity — three to five sustainable PIRs
D4Bind declared agent scope to runtime enforcementJoint delivery with AI Enablement. Not supplied by any posture product.
D4Stand up AI security posture managementLikely a module of an existing CNAPP licence given market consolidation
FEstablish detection-engineering capacity and lifecycleCSOC agreement. Define it before automating it.
D3Ratify the tiered containment catalogueSecurity Committee decision
THE DEPENDENCY TO TEST IN WEEK ONE, NOT WEEK NINE
The gateway change that mints a delegation identity is the head of the only four-link chain in the plan, and both of the programme’s distinctive detections sit at its end. Attribution is formally non-identifiable from logs alone, so if the gateway cannot be changed, the fallback is credential-level resolution — weaker, workable, and worth knowing about in month one.

And the risk that makes ownership sequencing non-negotiable. Gartner warns preemptive tooling may be “limit[ed] to detecting, rather than fixing issues, thereby increasing, rather than reducing overhead.” Stand up validation and posture assessment before accountable resolvers exist and tranche 2 manufactures a findings backlog nobody owns. Decision 02 must land before this tranche starts.

Seven themes. The test of this tranche is whether the programme stops being a programme — coverage measurement becoming a standing CSOC function rather than a project activity.

CommitmentNote
D1Continuous validation as a standing capabilityThe assurance loop Gartner describes as the core of the domain
D4Begin the macro- to microsegmentation journeyDiscovery-first: 30-90 days observation, pilot in 8-12 weeks, first segment in 3-6 months. Expect organisational failure modes — 11 of 14 attempts failed for want of a champion and app-owner buy-in
D3Cohort revocation; pre-authorised reversible actions liveSigning-authority taint and issue-time cohort invalidation
FAI Lab: self-hosted defensive and forensic capabilityFour arguments, none a capability claim
D4IT/OT governance framework; factory-software owner namedGartner describes the gap directly; network-enforced segmentation is the only viable primary control for legacy process equipment
D2Composite AI patterns for high-risk use casesGartner: “combine AI models with deterministic systems to ensure reliable outputs”
FManaged or partner service where in-house depth is absentGartner: most organisations need this “unless they have large and high-level security teams”
WHAT INSTITUTIONALISED ACTUALLY MEANS
Three tests. Coverage measurement runs on the CSOC calendar without programme involvement. Revocation is a Tier A action exercised quarterly by people who were not on the programme. And every capability has a named owner outside CDR — which is the foundation outcome, and the one Gartner escalates to a board accountability.

And the long-run mandate this is heading toward. Gartner’s 2030 framing is antifragility: “Simply returning to the status quo after an incident will no longer be viable… organizations will intentionally use systemic shocks and controlled risk exposure to emerge operationally stronger.” That is an argument for standing exercise capability, which is what tranche 3 buys.

Ordered by what they unblock rather than by size. Decisions 01 and 02 cost nothing and gate almost everything else.

DecisionWhat it unblocksCost
01Reset the risk-appetite question — from days-to-patch to “what is our tolerance to be exploited by a known vulnerability?”Makes exposure prioritisation defensible, and stops the programme being measured on a target it cannot hitNone
02Name accountable owners outside CDR — “IT operations, product teams and business system owners must be held accountable by the board.”Eight of twenty objectives. Without it, tranche 2 manufactures findings nobody ownsNone
03Adopt the four domains and the outcome setGives the programme a definition of success that is not activityNone
04Grant reversible containment authority — machine-identity revocation without prior change approvalD3 moves off MIL0. Gartner names this the required cultural shiftNone; needs the Tier A safeguard
05Elevate technical debt to a material business riskFunds the cleanup that reduces exposure faster than patching canFY27 planning
06Fund tranche 1; gate tranche 2 on its findingsDemonstrated delivery before a substantial askExisting headcount
IF ONLY ONE DECISION IS TAKEN
Decision 02. It is free, it is the programme’s critical path, and Gartner frames it as a board accountability rather than a negotiation. Only 36% of organisations have infrastructure teams actively engaged on remediation, and without business engagement exposure management “cannot function effectively.”

And the message to carry into the room alongside them. Gartner’s guidance for exactly this briefing: temper the fear, uncertainty and doubt, and use the attention as a decision point “to reset expectations, funding, and accountability.” The medium-term outlook is genuinely positive — frontier models “will enable defenders to inspect an unprecedented volume of source code” — and saying so buys more credibility than alarm would.

The document carries over 400 citation markers against a registry of 173 entries across 21 source classes. Every marker opens the verbatim quotation and its locator. Nothing rests on an unattributed assertion.

The evidence base, by layer

LayerWhat it is
Gartner corpusTwelve reports, 334 pages, supplied by the programme and read in full. G00859378 supplies the architecture and G00799085 the measurement system; the other ten evidence and refine them.
the enterprise’s own filingsForm 10-K risk factors and the Item 1C cybersecurity disclosure — including the board Security Committee and the company’s own statement that AI “may… lead to new and/or more sophisticated methods of attack.”
Government and court recordThe DOJ trade-secret case, the acquittal that bounds it, the ODNI threat assessment, and SEC disclosure requirements.
Primary technical researchTwenty-eight research dossiers written to disk as the work proceeded, covering detection, inventory, deception, recovery, the AI lab, exposure management, peer benchmarking and defence depth.
THE FOUR ITEMS STILL OPEN, ALL IN TRANCHE 1
The gateway AI control’s real detection scope is unverified, and it is the only AI-specific control we have. The SIEM’s detection content model is not established. A small number of peer quotes are secondary-sourced and marked for verification before external use. And our SEMI standards position is unconfirmed. None changes the strategy; two could change tranche 1.

One methodological limitation, disclosed rather than discovered. No board-credible AI-specific security maturity model exists yet. The assessment therefore applies a general maturity model (C2M2 levels) to AI-specific domains. That is defensible — Gartner notes the four domains are largely delivered by technologies that already exist — but it is a limitation, and better stated by us than found by a reviewer.

The most prevalent AI-augmented attack in the published evidence, and the one our defences did not cover. 41% of surveyed organisations have experienced a deepfake and social-engineering attack on an audio call to an employee, 35% on a video call, and deepfakes now account for one in five biometric fraud attempts.

Two incidents that frame the problem

IncidentWhat happenedWhat it proves
Arup, Hong Kong, Feb 2024A confirmed US$25.6M loss after an employee joined a video conference on which every other participant was synthetic, including the CFO.Real-time video deepfakes are operationally viable against a competent finance function.
MGM Resorts, Sept 2023Initial access via socially engineering the IT help desk into an MFA reset. Compromise in roughly ten minutes; disclosed impact above $100M.No synthetic media was required at all. The process weakness is exploitable with a phone call — AI only makes it cheaper to attempt at scale.
WHICH IS WHY THE SECOND INCIDENT MATTERS MORE TO US
If the control gap is procedural, then AI is an amplifier of an existing weakness rather than a new threat — exactly Gartner’s framing that frontier models “do not introduce a fundamentally new threat capability” but change velocity and scale. The mitigation is therefore available now, and it is not a purchase.

Why detection tooling is not the answer

ApproachDocumented limit
Commercial deepfake detectionModels lose roughly 45-50% of their AUC on realistic in-the-wild 2024 content versus laboratory benchmarks.
Liveness / presentation-attack detectionThe certification standard structurally excludes injection attacks — feeding synthetic video directly into the data path, which is the technique actually used. A product can hold valid certification and remain fully exposed.
Content provenance (C2PA)Metadata is destroyed by any screenshot or re-encode, and absence of a credential is not evidence of fakery. Useful for content you publish, not content you receive.
Training people to spot fakesUntrained participants identified deepfakes in approximately 0.1% of trials. (Vendor-funded study — treat the precise figure with caution; the direction is not in doubt.)

What actually works — and it is procedural

The controls that hold do not depend on detecting a fake:

  • Callback to a pre-registered channel — never a number or address supplied by the caller. This single control defeats the Arup pattern entirely.
  • A passphrase that never travels over the channel being verified. If the code word is spoken on the call being authenticated, it authenticates nothing.
  • Dual authorisation above a value threshold for payment changes, vendor bank-detail changes and privileged access grants.
  • Help-desk hardening for password and MFA reset — the MGM vector. This is the highest-value single process in scope.

And the documented failure modes are procedural too. Two recur: the attacker supplies the callback number (defeated by pre-registration), and manufactured urgency causes staff to bypass the procedure — which is a training and authority problem, not a technology one. Staff must be explicitly authorised to refuse and delay a senior request without career risk, or the control exists on paper only.

WHAT THIS MEANS FOR THE PROGRAMME
This is Gartner’s named action — modify standard operating procedures for password reset, partner and supplier interactions, and financial transactions — and it is outcome D2’s leading metric. It sits in tranche 1 because it is process change with no technology dependency, and it closes the single largest gap between our defences and the real threat.

Gartner sets the ceiling: “You only need five to nine metrics for each target audience. Do not report everything you know. Prioritize the top five to nine technology dependencies and the top audiences to receive only the highest-value information that drives above-the-line decisions.” Combined with “only above-the-line metrics should be shared with executives”, that is why fifteen run the programme and seven reach the Security Committee.

The selection test each of the seven had to pass

  • Does it cover a distinct domain? Two from D1, one each from D2, D3 and D4, one from the foundation. No domain is unrepresented and none is over-represented.
  • Is it a leading indicator? It must “expose a problem before it leads to material loss.” This is why validated coverage is on the list and incident counts are not.
  • Can it be computed honestly today? Anything needing a denominator we do not have was excluded — which is why reconciliation is on the list and inventory completeness is not.
  • Does moving it change a decision? “The value of a metric is its ability to influence decision making. The decisions influenced are typically priorities and investments.”

Thresholds and the action each triggers

Each metric’s exact unit of count, denominator and pass test is in its own specification panel, reachable from the Definition link beside it on the committee set and the programme set.

MetricAmberRedAction on red
01 · D1.2Priority controls validated by simulated attackFalling quarter on quarterNo validation run in a quarterValidation capacity becomes a funded item rather than a best-effort activity
02 · D1.1AI assets and identities reconciledDeclare-lag risingReconciliation rate falling while the estate growsDiscovery tooling escalated; registry treated as advisory until reconciled
03 · D2.1High-risk processes hardened against deepfake impersonationAny named process unhardened after 60 daysA hardened process bypassed in an exerciseAuthority to refuse and delay a senior request is restated in writing
04 · D3.2Machine-identity classes revocable in ten minutesRehearsal missed in a quarterAny crown-jewel-reaching class not revocableAppetite statement A2 is breached — the gap becomes a named risk with a remediation date
05 · D4.1Production agents with declared scope enforced at runtimeEnforcement coverage staticNew high-autonomy agents deployed without enforcementAutonomy tier reduced until enforcement exists
06 · D4.2Critical IT, OT and cloud assets under continuous posture assessmentNew critical assets outside assessmentOT scope slipping beyond FY27Scoping decision re-ratified explicitly rather than allowed to drift
07 · F.1Exposure findings with an accountable resolver outside CDROutstanding exceeding acceptedFindings ageing without an ownerEscalation to the Security Committee — this is the board accountability Gartner names
WHY THRESHOLDS MATTER MORE THAN THE NUMBERS
“When a metric changes from green to yellow to red, it drives an action such as changing a priority or an investment.” A metric with no stated consequence is a dashboard entry. Each of the seven above carries one, and the consequences are deliberately modest and executable — reduce an autonomy tier, name a risk, escalate an ownership gap — because an unenforceable consequence is worse than none.

And the eight we deliberately left off. The other eight programme metrics are below the line: they run the work and belong in the monthly CDR review, not the committee pack. Gartner is explicit that both layers are necessary and serve different purposes — but that “only above-the-line metrics should be shared with executives to drive effective business decisions.” Reporting all fifteen upward would reduce the decision value of each one.

WHAT THIS MEANS FOR THE PROGRAMME
Adopt the seven as the standing Security Committee set, with the thresholds above. Review the composition annually — not the thresholds quarterly, which would let the target follow the performance.

The outcome the Security Committee is protecting, stated in the business’s own terms: production output is not interrupted by a cyber event. It is first because it is the only one of the three with a precedent in this industry and a published number attached to it.

EvidenceDetail
TSMC, August 2018A misconfigured software installation propagated a variant of WannaCry across production tool controllers, halting several facilities. TSMC put the revenue impact at approximately 3% of third-quarter revenue, roughly US$255M, plus a gross-margin reduction of about one percentage point.
The cause matters more than the costTSMC stated the infection came from a new software tool installed without virus scanning before connection to the network — not an intrusion campaign. The failure mode was a change-control gap, and that is the same class of gap an unsupervised agent with write authority represents.
the enterprise’s own disclosureThe enterprise operates fabrication facilities across several countries with deep interdependency between sites, and its Form 10-K identifies unauthorised access to facilities or technology infrastructure as a named risk.
The sector’s standards lagSEMI E187 and E188 followed the 2018 incident by three to four years, and neither covers the manufacturing execution system, material-control system or factory host layer.

Which programme metrics have line of sight to it

  • Metric 06 — critical IT, OT and cloud assets under continuous posture assessment, because the OT estate is where continuity is lost.
  • Segmentation maturity under D4 — Gartner names microsegmentation specifically to “limit lateral movement” between converged IT and OT.
  • Metric 01 — validation, because an untested containment path is an assumption.

And the boundary we are explicit about. Response authority inside the production environment is deferred to FY27 by decision. For FY26 this goal is served by assessment, segmentation and validation — not by automated containment. Gartner describes why: converged environments “struggle to support the downtime required for constant patching.”

Process recipes, yield data and design files are the assets that determine whether the enterprise’s technology lead survives. This goal is second because the threat to it is documented in court records, not inferred.

EvidenceDetail
United Microelectronics CorporationPleaded guilty to trade-secret theft and was sentenced to a US$60M fine in respect of the enterprise core product technology. The Deputy Attorney General described the conduct as part of a campaign to acquire American technology.
Fujian JinhuaConvicted at trial in February 2024 on the related charges.
Competitive pressure is currentThe enterprise names ChangXin Memory Technologies as a competitor receiving state support.
The regulatory shock is realThe China Cyberspace Administration decision materially affected the enterprise’s revenue, disclosed in its own filings.
And AI is named as an IP risk by the enterprise itselfThe Form 10-K identifies risks arising from AI use in relation to intellectual property and from new methods of attack.

Why this goal changes the machine-identity metric’s weighting

Metric 04 — credential classes revocable within ten minutes — is weighted so that any class with reach into recipe or design data must be in scope. A revocation catalogue that covers the easy classes and omits the crown-jewel ones would improve the number while leaving the goal unprotected. That is precisely the failure Gartner’s “direct line of sight” test is meant to catch.

One complication worth stating to the committee. Running inference over controlled technical data may itself be an export-control event depending on where the model is hosted and who can reach it. This is a reason the self-hosted AI Lab sits in the defence set — not only for forensic independence, but because hosted guardrails also refuse legitimate defensive analysis.

The third goal is the one CISOs usually leave off the slide, and the one that makes the programme an enabler rather than a tax. Gartner’s own final step in deriving outcome-driven metrics is to express readiness as a risk to business enablement — not as a security score.

The argument in one line

If we cannot say what an agent is permitted to do, the business cannot safely say yes to it. Runtime enforcement of declared scope (metric 05) is therefore an adoption control as much as a security one — it is what allows a high-autonomy use case to be approved at all.

What the evidence showsConsequence for this goal
Enterprise AI initiatives are abandoned at a material rate, most often for governance and trust reasons rather than model performance.Ungoverned adoption is slower in practice, because projects stall at approval or get withdrawn after an incident.
Gartner advises risk-tiering frontier AI use cases and applying layered control planes, with composite AI patterns for the high-risk tier.A tiering scheme is the mechanism that lets low-risk use cases move fast while concentrating scrutiny where it is warranted.
Very few peers have a defined programme for non-human identity governance.Doing this creates a defensible position rather than merely catching up.
THE MEDIUM-TERM CASE IS GENUINELY POSITIVE
Gartner: frontier models “will enable defenders to inspect an unprecedented volume of source code.” The same capability that raises the threat is the one that eventually closes it — and the organisations able to use it will be the ones that already know what their agents are, what they are permitted to do, and how to withdraw their authority.

Why this goal is stated as speed and not as “secure AI adoption”. Because “secure adoption” has no direction. Speed is measurable and the business already tracks it, which is what makes the line of sight from a below-the-line protection level to an above-the-line business outcome real rather than rhetorical.

The diagram is not a presentational device. It is the output of a defined method, and the method is what makes the fifteen programme metrics defensible rather than assembled.

StepApplied here
1 Identify the business processes and the technology stacks that support themProduction and the factory-software layer; the design and yield-analysis estate; the agent and model estate.
2 Determine the business outcomes those processes deliverManufacturing continuity; protection of critical IP and recipe data; speed of the enterprise’s own AI adoption.
3 Identify the technology risks to those outcomesInterruption of tool control; exfiltration or corruption of recipe data; an agent acting outside its declared scope.
4 Derive the technology outcome-driven metricsThe fifteen below-the-line protection levels, three per domain plus three for the foundation.
5 Express readiness as a risk to business enablementThe readiness statement — which is why it is phrased as what the business cannot yet safely do.

What makes something an outcome-driven metric and not just a number

  • It is a continuous outcome of an identifiable investment — fund it and it improves, defund it and it degrades.
  • It sits on a sliding scale, so the committee is choosing a protection level rather than passing or failing.
  • It takes the canonical form “% of X”, which is why every metric on the executive set is written that way.
  • It has direct line of sight to a business outcome — and if it does not, “the technology investment should not be prioritized as it will drive little value.”

And the rule that cut the committee set from fifteen to seven. “You only need five to nine metrics for each target audience”, and “only above-the-line metrics should be shared with executives.” Gartner is also explicit about what to remove: stop reporting activity counts that no investment decision follows from.

The change being asked for

From “how many days do we take to patch?” to “what is our tolerance to be exploited by a known vulnerability?” — set by asset class.

Why the first question cannot be answered usefully any more

The interval between disclosure and working exploit has compressed from days to minutes for some classes of vulnerability, so a days-to-patch target is a statement about our process rather than about our exposure. Gartner is blunter still: we are “not patching our way out of vulnerability exposure”, and mitigation and compensating control must carry the load where patching cannot. Only 36% of organisations even have infrastructure teams actively engaged.

What the reframe enablesBecause
Differentiated targets by asset classA tolerance of near-zero for the factory-software layer and a higher one for a development sandbox are both rational; a single days-to-patch target for both is not.
Decision models rather than severity sortingCISA’s binding directive and the SSVC model already frame remediation as a decision on exploitability and mission impact.
An honest conversation about residual riskBecause the committee is setting a tolerance rather than approving a deadline it will not meet.

What this does not mean. It is not a licence to stop patching. Gartner’s framing of preemptive security is explicit that these approaches are “not substitutes or replacements for cyber hygiene”. The reframe changes what we measure and report, not what we maintain.

The ask, in the words the board should hear

“cybersecurity teams can only guide vulnerability prioritization. IT operations, product teams and business system owners must be held accountable by the board to fix exposures in their own systems.”

Why this is the critical path and not an administrative preference

Roughly 80% of the controls this programme depends on sit outside Cyber Defense & Resilience. Gartner states the consequence directly: “without widespread business engagement most exposure management functions… are unable to function effectively, and warns that preemptive tooling can end up “limit[ed] to detecting, rather than fixing issues, thereby increasing, rather than reducing overhead.”

Dependency areaThe accountable owner being requested
Machine-identity lifecycle, issuance and revocation mechanismIdentity
Segmentation, and binding the agent registry to a runtime enforcement pointPlatform and infrastructure & operations
Repository and build policy, artefact-registry defaultsDeveloper experience
The factory-software layer — the manufacturing execution system (MES), material-control system (MCS) and factory host systems that SEMI E187 and E188 do not coverManufacturing
Standard operating procedures for password and MFA reset, supplier interaction and financial transactionsService desk, procurement, finance and HR
AND THE METRIC THAT MAKES IT VISIBLE
Committee metric 07 — % of exposure findings with an accountable resolver outside CDR. It is on the executive set specifically so that a failure to grant this decision shows up as a number rather than as a complaint.

What is being ratified

Gartner’s four core preemptive domains as the programme’s structure — exposure management, adversary management and threat intelligence, adversary disruption, and posture and policy management — on a foundation of managed services and capabilities, with the fifteen outcome-driven metrics as the definition of success.

Why adopt an external frame rather than write our own

  • It is additive, not a replacement: preemptive security “adds focus beyond prevention, detection, and response”, so prevent–detect–contain–respond remains the delivery method underneath.
  • Much of it is already bought: Gartner notes many preemptive capabilities already exist in current platforms, which is why tranche 1 needs no new money.
  • It gives the committee a defensible external reference for the structure, so debate can be about sequencing and funding rather than about taxonomy.
  • It carries its own anti-hype discipline: focus on “how the tools support these outcomes, not on the hype around AI.”

The alternative framing, if the committee prefers three words to four domains. Gartner’s other formulation of the same idea is predict, prioritise, prevent. The content is unchanged; only the label differs. What should not change is that the outcomes are adopted alongside the structure — a frame without metrics is a diagram.

The specific authority requested

Revocation of machine-identity credentials without prior change approval, for an enumerated Tier A list, with the list changeable only by the Security Committee and every action logged and reviewed.

Why speed is the whole point

Average adversary breakout time is 29 minutes; the fastest observed was 27 seconds, and the median hand-off to a second-stage operator 22 seconds. A containment action gated on a change window is not a containment action.

Why machine identity and nothing else

  • It is the reversible asset class — revoking a credential breaks a workload, not a person’s day, and it can be reissued in minutes.
  • It is the class the reference incident actually abused, and where the adversary held signing authority and could mint credentials faster than they were withdrawn — which is why the catalogue must include authority-level taint, not only credential revocation.
  • No human-account, endpoint-isolation or production-network authority is being requested. Those are not reversible in the same sense and are not in scope for this decision.
GARTNER NAMES THIS AS THE HARDEST ASK ON THE PAGE
“it is extremely rare for organizations to be willing to automatically remediate discovered issues due to the concern over potential disruption… Without a cultural shift, many CISOs will be unable to fully utilize preemptive cybersecurity approaches.” We are asking for that shift once, in the narrowest place where it is defensible.

And the proof obligation we accept in return. The authority should be contingent on rehearsal: a named authority, a written procedure and a measured end-to-end time per credential class, tested quarterly. That is committee metric 04. If the rehearsal lapses, the authority should lapse with it.

The finding

“outdated systems and unused code are no longer just operational problems, but significant cyberthreat liabilities.”

Why frontier models change the calculus specifically

The capability that has genuinely shifted is the speed and scale of finding and weaponising known weakness, not the invention of new attack classes. Legacy and unused code is exactly the surface that rewards cheap, patient, automated review — and unused code is the worst case, because nobody is watching it and nobody will notice it being probed. Gartner’s medium-term observation cuts both ways: models “will enable defenders to inspect an unprecedented volume of source code” — and attackers first.

What changes if this is acceptedMechanism
Decommissioning gets fundedBecause it becomes a risk-reduction line rather than a deferred engineering nicety.
Unused code becomes a tracked exposureDead endpoints, orphaned repositories and stale artefact registries enter the exposure inventory rather than sitting outside it.
Mitigation is legitimised where patching is impossibleConsistent with Gartner’s position that we will not patch our way out of this.
The IT/OT case becomes statableConverged environments “struggle to support the downtime required for constant patching” — which is a technical-debt problem wearing an operations costume.

Why it belongs on a board agenda rather than an engineering backlog. Because the decision that creates technical debt is a funding and prioritisation decision, and it is made above the level of the teams who inherit it. Naming it as material risk is what allows a resolver team to trade a feature for a decommission without needing to win that argument alone.

The proposition

Approve tranche 1 now — it requires no new money, because three of its seven items are governance or process and the rest use capability already licensed. Then gate tranche 2 on what tranche 1 finds, rather than approving a full-programme budget against assumptions.

The two tranche-1 items that may change tranche 2’s scope

ItemWhat it could change
The coverage baseline — which techniques are reachable with the telemetry we already have. Today 16 of 78 by the modality we hold, against 69 of 78 by runtime instrumentation.If the gap is mostly closable by re-pointing existing collection, the tranche-2 telemetry line shrinks. If it is not, the case for runtime instrumentation is made with our own numbers rather than a vendor’s.
The gateway delegation test — whether a delegation identity can be minted at the gateway. Attribution is formally non-identifiable from logs alone, and trace-based grouping recovers only 4-6% of a delegation’s events.If the gateway can carry it, runtime enforcement of declared scope is a build. If it cannot, the fallback is credential-level resolution — weaker, cheaper, and worth knowing before committing to the enforcement design.

Why sequencing disruption before detection is deliberate

Because it is where we are weakest and where the cheap wins are. Deception needs no inventory, no gateway change and no new telemetry pipeline; SOP hardening is process change; revocation authority is a decision. Gartner notes disruption “provides a clear signal of an attack” and feeds other controls.

What the committee is really approving. Not a budget. A method of deciding the budget — spend a quarter establishing the two numbers that determine the largest line items, then fund against measurements rather than estimates. If tranche 1 produces no findings that change tranche 2, that is itself a useful result.

Every other number on the baseline is unflattering. This one is not, and it is worth understanding why it matters more than it looks.

What the enterprise has already disclosed. “Our Board of Directors administers its cybersecurity risk oversight function directly as a whole, as well as through the Security Committee… which oversees monitoring and incident response, risk mitigation, supply chain…”

Why this is the scarce asset

What the committee already hasWhy it usually has to be built
A standing board-level body with an explicit cybersecurity remitMost programmes must first establish a forum, then earn its attention, then win the right to ask it for authority. That is typically a year.
Authority that reaches outside the security functionDecision 02 — naming accountable owners in IT operations, product and business teams — is only actionable because a body with that reach exists. Gartner frames this as a board accountability, not a security one.
Standing agenda time for risk-appetite conversationsDecision 01 is a risk-appetite reset. There is no other forum that can make it.
An existing route for materiality judgementsWhich matters given SEC Regulation S-K Item 106, Item 1C and the Form 8-K Item 1.05 four-business-day determination.

What it means for the maturity assessment

Programme governance is assessed at MIL2 today and targeted at MIL3 — the highest starting point of anything on the baseline. Six of the fifteen assessed capabilities sit at MIL0. The gap is not governance; it is execution capacity and delegated authority, and both are things this committee can grant.

SO THE BRIEFING SHOULD ASK, NOT INFORM
Gartner’s guidance for exactly this board scenario is to temper the fear, uncertainty and doubt (FUD) and use the attention as a decision point “to reset expectations, funding, and accountability.” The scarce resource in the room is decisions, not understanding.

The baseline is assessed against the Cybersecurity Capability Maturity Model (C2M2), published by the U.S. Department of Energy. Its Maturity Indicator Levels (MILs) are cumulative and are assessed per capability, not as a single organisational score.

LevelDefinitionWhat it takes to claim it
MIL0The practice is not performed.Nothing. It is the honest answer for six of our fifteen capabilities.
MIL1Initial practices are performed, but may be ad hoc.Someone does it. It need not be written down or repeatable.
MIL2Practices are documented, and adequately resourced and skilled.A written procedure, named people and enough capacity to run it. This is the FY27 target for most capabilities.
MIL3Practices are guided by policy, periodically reviewed, and measured for effectiveness.Policy, review cadence and an effectiveness measure. Reserved for the two capabilities where the committee is being asked to grant authority: control validation and machine-identity revocation.

Why not the NIST Cybersecurity Framework Tiers

Because the CSF Tiers are not a maturity scale. They describe the rigour of an organisation’s cybersecurity risk-governance and risk-management practices, and NIST’s own Organizational Profile guidance frames them as a characterisation to inform a target profile — not a ladder to climb. Using them as one produces a number the committee will reasonably read as a score, and which cannot be tied to a specific investment.

What CSF 2.0 is doing in this strategy. Two subcategories are load-bearing rather than decorative: GV.RM-02, that risk appetite and tolerance are established and communicated — which is decision 01 — and the GV.RR category on roles, responsibilities and authorities — which is decision 02. We use CSF for governance structure and C2M2 for capability measurement.

AND THE LIMIT OF ANY MATURITY ASSESSMENT
A MIL is a statement about practice, not about protection. A documented, resourced and skilled detection capability can still be blind to a technique nobody wrote an analytic for — which is why the executive set leads with validated coverage rather than with maturity, and why maturity does not appear on the committee’s seven metrics at all.

The word “ownership” usually produces a workshop and a matrix. This ask is narrower and harder: a board instruction naming a single accountable owner for each of five dependency areas, with the obligation to accept or reject exposure findings on a defined clock.

Dependency areaAccountable owner requestedThe first thing they would be asked for
Machine-identity lifecycleIdentityAn enumerated list of credential classes, and a revocation mechanism that works by cohort rather than one credential at a time.
Segmentation and enforcement bindingPlatform / I&OThirty to ninety days of traffic observation before any enforcement, and a decision on where the agent registry binds to a runtime enforcement point.
Repository and build policyDeveloper experienceArtefact-registry anonymous-access defaults reviewed, and build-time provenance for anything reaching production.
The factory-software layerManufacturingAn owner for the manufacturing execution system, material-control system and factory host systems — the layer SEMI E187 and E188 do not cover.
High-risk standard operating proceduresService desk, finance, procurement, HROut-of-band callback to a pre-registered channel on password and MFA reset, supplier interaction, and financial transactions.

Why a named owner and not a shared responsibility

Gartner measured the engagement problem: only 36% of organisations have infrastructure teams actively engaged on remediation, and exposure management functions without business engagement are “unable to function effectively.” It also observes that the obstacles here are predominantly non-technical. The segmentation evidence is the starkest version: of fourteen organisations attempting even limited segmentation, eleven failed — mostly for want of an executive champion and application-owner buy-in.

What accountability means in practice, stated so it can be accepted or refused. Not that the owner fixes everything. That the owner accepts a finding with a date, or rejects it with a reason, within a defined window — and that unowned findings escalate to the Security Committee rather than accumulating in a CDR backlog. Committee metric 07 measures exactly this, which is why it earns a place on a seven-metric set.

THE FAILURE MODE IF THIS IS NOT GRANTED
The programme still builds validation and posture assessment, still generates findings, and manufactures a backlog nobody owns — Gartner’s explicit warning that preemptive offerings can end up “increasing, rather than reducing overhead.” That is the single most likely way this programme creates work instead of protection.

The strategy is built on two of the twelve reports and corroborated by the other ten. It is worth being explicit about which, because it determines what is load-bearing and what is supporting.

RoleReportsWhat depends on it
BaseG00859378 Outcomes of Preemptive Cybersecurity; G00799085 Outcome-Driven Metrics for the Digital EraThe four-domain architecture and the entire outcome mechanism. If either is wrong, the strategy is wrong.
Threat pictureG00852902 AI-Augmented Attacks; G00852689 the 2026-2027 landscapeThe threat tab, including the deepfake prevalence figures and the counterweight numbers.
Board framingG00856541 CISO Board ScenarioDecisions 01, 02 and 05, and the tone of the closing message.
ExecutionG00837909 CTEM roadmap; G00810627 vulnerability exposure; G00853789 infrastructure cybersecurity 2027; G00858028 frontier-AI checklistScoping, resolver teams, the microsegmentation journey, IT/OT governance and risk tiering.
DirectionG00836733 preemptive security; G00845745 Future of the CISO 2030The market trajectory and the longer-horizon risks. Supporting, not load-bearing.
Context onlyG00846817 Hype Cycle for Enterprise ArchitectureRead; not used. Stated so the corpus count is not mistaken for corpus dependency.

How it was read

  • Every report extracted in full and analysed page by page; dossier 27 carries verbatim quotes with page locators for every claim that appears in this document.
  • The two base reports were read end to end before anything was built on them — which is how the five-to-nine metric ceiling and the five-step derivation method came to shape the outcomes tab rather than being retrofitted.
  • Where Gartner’s position is more conservative than the prevailing narrative, Gartner’s position is the one used — including that true AI-powered attacks remain rare and that current offensive use is largely experimental.
  • Where Gartner and a primary source disagree on a number, both are shown and the primary source is preferred. The Verizon report is currently secondary-sourced and flagged as such.

And the discipline that goes with reading a vendor-analyst corpus. Gartner supplies its own caution and we apply it to this document too: focus on “how the tools support these outcomes, not on the hype around AI, and treat preemptive capability as additive to rather than a replacement for cyber hygiene. Several claims in the corpus are strategic planning assumptions rather than measurements; where one is used, it is labelled as such.

WHAT IS NOT FROM GARTNER
The technical depth in this document — MITRE ATLAS and Summiting the Pyramid coverage, the OpenSSF SAF-MCP matrix, the delegation-attribution proof, the deception evidence, the deepfake-detection evaluations and the segmentation failure data — comes from primary technical sources and is carried in dossiers 01 to 28. Gartner sets the frame; the primary research sets the numbers.
WHAT THE NUMBER ASKS, IN PLAIN WORDS
Of everything our discovery signals actually found, what share do we have under control — one record, one named owner, a declared purpose, and a full list of the credentials it holds?
% of discovered AI assets and machine identities that are fully reconciled

1 · What exactly one item is

One object, where an object is either an AI asset or a machine identity, deduplicated on a stated join key.

Definition and why it is drawn here
AI assetCounted as an object in five classes only: a network-reachable model endpoint (URL plus deployment name); a registered agent (registry identifier); a tool or Model Context Protocol server exposed to an agent (server URI plus tool name); a retrieval corpus or index an agent can read; a stored model artefact in a registry.
Machine identityA credential-bearing non-human principal, in ten classes: application registration or service principal; system-assigned managed identity; user-assigned managed identity; workload identity federation subject; platform-issued application programming interface (API) key or static secret; OAuth client-credentials grant; personal access token used by automation; workload certificate or SPIFFE identity; cloud role or service account assumed by a workload; signing identity for code or artefacts.
Deliberately not countedIndividual prompts, chat sessions, and notebook experiments not reachable from a shared endpoint. These are excluded because they are unbounded and would make the count irreproducible — two people counting on the same day would not agree.
The join keyIdentities join on the platform object identifier. Assets join on endpoint URL plus deployment name, or on the registry identifier. Objects that cannot be joined are counted separately and flagged unjoinable, and the unjoinable count is published beside the metric as its own error bar.

2 · The denominator, and where it comes from

The union of all discovery signals, deduplicated. Five signals, in descending order of trustworthiness: S1 resource-tag scan across all cloud subscriptions; S2 six-signal heuristics over application registrations and identities (permission shape, naming pattern, redirect URI, credential type, sign-in pattern, absent owner); S3 gateway and egress telemetry showing traffic to an inference endpoint; S4 a service-management platform record; S5 registry self-declaration. An object is discovered if it appears in at least one signal.

System of record. The the agent registry for assets; the identity platform for identities. The service-management platform holds the reconciliation join.

3 · The test that puts an item in the numerator

ConditionRequirement
R1 — one recordJoined to exactly one record in the system of record. No duplicate, no orphan.
R2 — a named ownerA named individual who is currently employed and has acknowledged ownership within the last twelve months. A team name fails.
R3 — declared scopeDeclared purpose and permitted scope present and non-empty.
R4 — complete credential listEvery credential the object holds is enumerated, with an issue date and an expiry.

4 · How a pass gets wrongly claimed

  • Reconciled against the registry alone — which only proves the object declared itself, not that it exists as declared.
  • An owner field populated with a distribution list, a team, or a person who has left.
  • Partial reconciliation counted as a pass. All four tests must hold; three of four is a fail.
  • Silently dropping unjoinable objects instead of publishing the count.

5 · How it is computed

Computed monthly by the discovery pipeline, not by hand. Each signal is a scheduled query; the union and the dedup are scripted so the number is reproducible from the same inputs. Owned by CDR with Identity as the data provider for S1 and S2.

6 · The arithmetic, worked through

LineValue
Tag scan (S1) returns candidate identities340
Heuristics (S2) flag application registrations as agent-like128
Gateway telemetry (S3) shows distinct inference endpoints44
Registry (S5) holds declared agents61
Union after deduplication — the denominator418
Have a record in the system of record61
Pass all four reconciliation tests — the numerator38
Metric38 / 418 = 9%
Companion — shadow rate (only ever seen by S1–S3)357 / 418 = 85%

What here is real and what is illustrative. The arithmetic above is illustrative, to show how the figure resolves. What is established today is that the registry exists but is self-declared and unreconciled, so the denominator has not yet been enumerated. Publishing the denominator is the first deliverable, not the percentage.

7 · How this number could lie

THE GAMING RISK, AND WHAT CLOSES IT
This number legitimately falls when discovery improves. Adding a sixth signal enlarges the denominator and the percentage drops — which is a better state, not a regression. Report it beside the shadow rate and the median declare-lag (days from first discovery signal to a reconciled record), both of which move in the right direction unambiguously. A rising percentage with a static denominator is the pattern to distrust.

Why self-declaration is the wrong first step, not the wrong idea

The only published discovery methodology for this problem puts developer attestation fourth of four — tag scan, then heuristics, then reconciliation against the configuration management database (CMDB), then attestation of whatever residue remains, with a deadline and an escalation path. Our registry is step one of one. Invert the order and the same registry becomes a trustworthy join key rather than the population itself.

Why this metric is first to fund

Metrics D1.2, D3.2, D4.1 and D4.2 all need a scoped population before they can be computed at all. Gartner recommends pulling automated information technology, operational technology and cloud inventories from the exposure platform rather than maintaining them separately. The three-layer asset data model in the service-management platform is the intended join point.

WHAT THE NUMBER ASKS, IN PLAIN WORDS
Of the adversary techniques we have decided matter, for how many have we run the attack and watched a defence hold? Not “do we own a product that should stop it” — did we try it.
% of priority adversary techniques with at least one defence proven by executing the attack in the last 90 days

1 · What exactly one item is

One technique on the versioned priority technique list. Not one control — controls are not countable objects, because no two people enumerate them the same way.

Definition and why it is drawn here
Why the unit is a technique and not a controlA control count can be made to say anything: one platform is one control or forty, depending on who is asked. A technique is an externally defined, enumerable object with a published identifier, so the denominator can be audited by someone outside the programme.
What sits underneathA register of defence claims, each a tuple of (technique identifier, control instance, expected outcome), where expected outcome is prevent, detect or both. The claim register drives the testing; the technique list drives the metric.
Deliberately not countedTechniques we have documented a decision not to defend against. These leave the denominator only by written decision with a named approver, and the count of excluded techniques is published with the metric.

2 · The denominator, and where it comes from

The priority technique list — named, versioned and published with the metric. Version 1 is built from the techniques MITRE mapped in ATLAS case study AML.CS0068 set within the 78-technique coordination-and-tool-use set used for the coverage baseline, minus documented exclusions. Re-ratified quarterly; the version number travels with every published figure.

System of record. The defence-claim register and the validation run log, held by the validation function. Every pass carries a run identifier so it can be re-examined.

3 · The test that puts an item in the numerator

ConditionRequirement
A prevent claim passes whenThe emulated action is attempted against a representative target and fails, and the failure is attributable to the named control by a log entry showing the denial. “It did not work” without an attributable denial is not a pass.
A detect claim passes whenThe emulated action executes and an alert of the claimed severity appears in the security information and event management (SIEM) platform within the claimed time budget, and the alert names the technique. A generic anomaly alert is not a pass.
EnvironmentProduction, or a production-equivalent replica with the same configuration. A pass achieved only in a lab is recorded as lab pass and does not enter the numerator.
FreshnessA pass expires after 90 days, because configuration drifts. The metric is therefore always a statement about the present.

4 · How a pass gets wrongly claimed

  • Counting a technique as covered because an analytic exists. Existence is not evidence. That is metric D2.3, which is a different and easier test.
  • A lab pass promoted to a production pass without re-running it in production.
  • An alert that fires but cannot be attributed to the emulated action — common when the test runs during normal change activity.
  • An alert outside the claimed time budget recorded as a pass because it eventually arrived.
  • Shrinking the technique list to raise the percentage, without publishing the exclusions.

5 · How it is computed

Run continuously by the validation and purple-team function; the figure is recomputed at each publication from the run log rather than maintained as a spreadsheet. Evidence for each pass is the run identifier, the target, the timestamp and the attributed log line.

6 · The arithmetic, worked through

LineValue
Priority technique list v122 techniques
Defence claims made against them31 claims
Techniques with no claim at all7
Claims that passed in production within 90 days12
Techniques with at least one passing claim — the numerator9
Metric9 / 22 = 41%
Companion — claim-level pass rate12 / 31 = 39%
Companion — techniques with no claim7 / 22 = 32%

What here is real and what is illustrative. The arithmetic is illustrative. The established position today is zero for the autonomous technique block: no validation run has been completed against it, so the honest current value is 0%.

7 · How this number could lie

THE GAMING RISK, AND WHAT CLOSES IT
Two ways. First, claim few defences — a 100% claim-level pass rate means nothing if seven techniques carry no claim, which is exactly why the metric counts techniques rather than claims. Second, shrink the list. Both are closed by publishing the list version, the exclusion count and the no-claim count alongside every figure.

Why this is the single most important number on the page

It is the direct measure of Gartner’s stated benefit for this domain: visibility into “where and why they fail prior to an actual attack.” It is also the only metric in the set that cannot be improved by buying something. It moves when a defence is executed against and holds.

And the reporting discipline that must travel with it

Never publish it as a bare percentage. The coverage framework underneath states its own limits twice and honestly: every technique has an analytic, and every mapping is partial. If any technique ever shows complete coverage, the measurement is wrong. Score detections on robustness, precision and implementation coverage rather than on techniques touched.

WHAT THE NUMBER ASKS, IN PLAIN WORDS
For each store of data we would most hate to lose, do we know who can reach it, who can change it, where its credentials could carry an intruder, and what would tell us?
% of crown-jewel data stores with a current, machine-generated blast-radius map

1 · What exactly one item is

One data store on the crown-jewel register. A store is a named, addressable repository of data — not a system, not a team, not an application.

Definition and why it is drawn here
The six store classesSource-code repository; process-recipe store; yield and test-data store; design-file store; model-weights store; signing-key store.
Deliberately not countedCopies and caches that are not separately addressable. Where a copy is separately addressable and reachable by a different principal set, it is a separate store — because its blast radius is different.

2 · The denominator, and where it comes from

The crown-jewel register: a business-owned, Security-Committee-ratified, versioned list. Inclusion test — a store is crown-jewel if loss, corruption or disclosure would (a) halt or degrade production output, (b) disclose process-recipe or design intellectual property, or (c) allow trusted code or artefacts to be published in the enterprise’s name.

System of record. The register itself, plus the identity graph and exposure platform that generate the maps.

3 · The test that puts an item in the numerator

ConditionRequirement
B1 — effective readersThe resolved principal list with read access, expanded through group and role nesting, including machine identities. The access-control list as written is not the effective list and does not pass.
B2 — effective writers and publishersThe narrower set holding write, merge, tag, sign or deploy rights.
B3 — reachable egressThe destinations and downstream systems a workload holding the store’s credentials can reach.
B4 — the detecting telemetryThe named log source and field that would show a bulk read or an unexpected write.
Currency and provenanceAll four generated from live platform data within the last 90 days. A hand-maintained list fails regardless of accuracy, because it cannot be re-derived.

4 · How a pass gets wrongly claimed

  • Using the access-control list as written rather than effective permissions. This is the usual failure and it understates the radius by an order of magnitude.
  • Mapping human principals only and omitting machine identities — the second usual failure, and the one that matters most here.
  • A map with no egress element, which answers who can read it but not where it can go.
  • A map with no named detection source, which answers the exposure question but leaves no way to notice abuse.

5 · How it is computed

Generated monthly by query against the identity graph and exposure platform. CDR owns the generation; the business owns the register and its ratification date.

6 · The arithmetic, worked through

LineValue
Crown-jewel register v164 stores
Have an effective read list41
Also have the write and publish list22
Also have egress mapped11
Also have a named detection source — the numerator9
Metric9 / 64 = 14%

What here is real and what is illustrative. Illustrative. The register does not yet exist in ratified form, so neither the numerator nor the denominator can be stated today. Establishing the register is a business task, not a CDR task.

7 · How this number could lie

THE GAMING RISK, AND WHAT CLOSES IT
A short register makes this trivially easy. Publish the register size and its last ratification date with every figure. And note the corollary: if the business will not name its crown jewels, that refusal is the finding, and it belongs in front of the Security Committee rather than inside a metric.

Why this is the scoping metric for the whole programme

Gartner is explicit that exposure scoping should follow potential business impact rather than threat severity alone. The crown-jewel register is that scope, and four other metrics inherit it: deception placement (D3.1) is scoped to crown-jewel environments, revocation classes (D3.2) are weighted by crown-jewel reach, posture assessment (D4.2) defines critical by crown-jewel path, and the validation target set is chosen for blast radius. Get this register wrong and four other metrics measure the wrong things accurately.

WHAT THE NUMBER ASKS, IN PLAIN WORDS
Of the procedures where a convincing phone call could move money or reset a credential, how many have we attacked ourselves and failed to break?
% of named high-risk procedures that pass a live social-engineering test with all four verification controls in place

1 · What exactly one item is

One named procedure. A procedure is countable because it has a written trigger, a sequence of steps and a named execution owner. Not a department, not a policy, not a training course.

Definition and why it is drawn here
The inclusion testA procedure is in scope if an instruction received over a voice, video or messaging channel can cause: (a) a credential or multi-factor authentication (MFA) factor to change, (b) money to move, (c) a supplier or bank detail to change, or (d) data or access to be granted.
Deliberately not countedAwareness training completion. It is a different thing measured on a different scale, and counting it here would let a training push move a control metric.

2 · The denominator, and where it comes from

The high-risk procedure register. Gartner names the areas to start from: password reset, partner and supplier interaction, and financial transactions. Version 1 is expected to hold on the order of a dozen procedures — help-desk password reset, help-desk MFA reset, privileged access grant, new supplier onboarding, supplier bank-detail change, payment release above threshold, purchase-order amendment, contract signature, employee data change, physical access grant.

System of record. The procedure register, owned jointly by the service desk, finance, procurement and human resources. The exercise results are held by CDR.

3 · The test that puts an item in the numerator

ConditionRequirement
H1 — out-of-band callbackVerification is completed by calling back on a channel drawn from the system of record, never from the requester.
H2 — a secret off-channelA shared secret or passphrase that never traverses the channel being verified.
H3 — dual authorisationA second authoriser above a stated value or privilege threshold. The threshold must be a number, not a judgement.
H4 — stated authority to refuseThe procedure states, in writing, that staff may delay or refuse a request pending verification and that no adverse consequence follows from doing so, including for a request that appears to come from a senior executive.
And the deciding evidenceA social-engineering exercise against that specific procedure in the last 180 days that did not succeed. Written-only is not hardened. A procedure bypassed during an exercise reverts to fail immediately.

4 · How a pass gets wrongly claimed

  • The callback number taken from the caller. This is the single most documented failure mode and it defeats the control entirely.
  • The passphrase spoken on the same call it is meant to verify.
  • A threshold described as “significant” or “unusual” rather than stated as a number.
  • Refusal authority that exists in the security policy but not in the procedure the agent is reading at the time.
  • Counting the procedure because the text was updated, without an exercise. Manufactured urgency causes staff to bypass procedures they know — only the exercise tests that.

5 · How it is computed

Procedure owners attest the text; CDR runs the exercises on a rolling schedule so every procedure is tested within 180 days. The exercise result, not the attestation, sets the value.

6 · The arithmetic, worked through

LineValue
High-risk procedure register v111 procedures
Contain all four controls in the written text0
Also passed a live exercise in the last 180 days — the numerator0
Metric0 / 11 = 0%

What here is real and what is illustrative. Here the zero is real. This capability is assessed at MIL0 — the practice is not performed — and it is the largest single gap against the most prevalent actual threat. The register size of eleven is illustrative; the numerator is not.

7 · How this number could lie

THE GAMING RISK, AND WHAT CLOSES IT
The obvious move is to count documents. The exercise requirement closes it. The second move is to exercise only the easy procedures — closed by requiring every registered procedure to be tested within the 180-day window, and by publishing the count of procedures never yet exercised.

Why this is on the committee set at all

Because it is the most prevalent real AI-augmented attack and the cheapest to mitigate. 41% of organisations have experienced deepfake-enabled social engineering on an audio call and 35% on video; deepfakes account for one in five biometric fraud attempts. And the investment is process redesign and training — no technology purchase.

IncidentWhy it is the relevant precedent
Arup, February 2024 — a confirmed US$25.6M loss after an employee joined a video call on which every other participant was synthetic.Real-time video deepfakes are operationally viable against a competent finance function.
MGM Resorts, September 2023 — a help-desk MFA reset, roughly ten minutes to compromise, more than $100M in impact.No synthetic media was needed. The process weakness is exploitable with a phone call; AI only makes it cheaper at scale. This is the more important case for us, because it is the one our controls must stop first.

Why detection tooling is not the control

  • Commercial detectors lose roughly 45 to 50% of their area under the curve (AUC) on realistic in-the-wild content compared with laboratory benchmarks.
  • Liveness certification structurally excludes injection attacks — the technique actually used — so a certified product can be fully exposed to it.
  • Content provenance is destroyed by any screenshot or re-encode, and the absence of a credential is not evidence of fakery.
  • Untrained people identify deepfakes in roughly 0.1% of trials.
SO THE CONTROL IS PROCEDURAL, AND THE HARD PART IS AUTHORITY
The second recurring failure is not technical: manufactured urgency causes staff to bypass a procedure they know perfectly well. That is a matter of stated authority, not training. Staff must be explicitly permitted to refuse and delay a senior request without career risk, which is why H4 is a pass condition and not a nice-to-have.
WHAT THE NUMBER ASKS, IN PLAIN WORDS
Of the attacks we have decided are worth worrying about, how many have turned into work a named person owns rather than a paragraph in a report?
% of priority threat scenarios converted into an implementable requirement with a named individual owner

1 · What exactly one item is

One scenario on the scenario register. A scenario is a named adversary objective plus the sequence of techniques used to reach it against a named the enterprise asset. It is countable because it is a register entry with an identifier.

Definition and why it is drawn here
What is not a scenarioA threat actor name. A technique on its own. A news article. These become scenarios only when written against a named asset with an objective, which is what makes them testable.
Register size is capped deliberatelyA register of two hundred scenarios is a backlog, not a plan. Version 1 is capped so that conversion is achievable within a quarter, and the cap is published.

2 · The denominator, and where it comes from

The scenario register, derived from the programme’s priority intelligence requirements (PIRs) and re-ratified quarterly.

System of record. The scenario register and the detection-requirement tracker, joined on scenario identifier.

3 · The test that puts an item in the numerator

ConditionRequirement
Five attributes, all requiredThe scenario has produced at least one requirement that is: (1) written in implementable terms — data source, logic intent, expected precision; (2) assigned to a named individual; (3) carries a due date; (4) carries a state of open, built, validated or retired; (5) is traceable back to the scenario identifier and forward to a built analytic or an explicit decision.
An explicit decline is a passA documented “we will not detect this, here is why, and here is the compensating control” counts as converted. This is deliberate — without it the metric punishes honest triage and nobody will record a decline.

4 · How a pass gets wrongly claimed

  • An owner recorded as a team or a function. Teams do not convert scenarios; people do.
  • A requirement written as an intention (“improve coverage of credential abuse”) rather than as something an engineer can build.
  • A scenario marked converted with no forward traceability, so nobody can tell whether the analytic was ever built.
  • Declines recorded without a compensating control, which turns the escape hatch into a way of clearing the register.

5 · How it is computed

Computed from the requirement tracker at each quarterly scenario review, joined on scenario identifier. Owned by threat intelligence, with detection engineering as the receiving function.

6 · The arithmetic, worked through

LineValue
Scenario register v118 scenarios
Have a requirement with all five attributes6
— of which built or validated4
— of which declined with a named compensating control2
Metric6 / 18 = 33%

What here is real and what is illustrative. Illustrative. Today there is no scenario-to-requirement conversion process, so this capability sits at MIL0 and the register does not exist.

7 · How this number could lie

THE GAMING RISK, AND WHAT CLOSES IT
Convert the easy scenarios and leave the hard ones perpetually open. The counter is to publish age of the oldest unconverted scenario beside the percentage — a healthy register has a low maximum age, not just a high conversion rate.

Why this metric exists rather than a threat-intelligence volume measure

Because intelligence that does not change a control is overhead. Gartner puts it plainly: adversary management “is of most value when combined with one of the other pillars to provide actionable enforcement.” This metric is the join between D2 and everything else — it measures whether the intelligence function produces work the rest of the programme can act on.

And it is the mechanism Gartner recommends for the roadmap itself. Scenario simulations should feed the capability roadmap rather than sit beside it, which is why the conversion record is required to be traceable forward to a built analytic or a written decline. The traceability is the deliverable; the percentage is just its summary.

WHAT THE NUMBER ASKS, IN PLAIN WORDS
For how many of those techniques would the evidence already be sitting in our logs if it happened tonight — with the fields actually populated, not merely permitted by the schema?
% of priority techniques with a named log source that is collected today and carries the required fields

1 · What exactly one item is

One technique — the same versioned priority technique list used by D1.2. Same denominator, different test: D2.3 asks whether the evidence would exist; D1.2 asks whether we proved we would catch it.

Definition and why it is drawn here
Why the two metrics share a denominatorSo that the pair can be read as a gap. If D2.3 is high and D1.2 is low, we have the data and lack the analytics. If both are low, we have a telemetry problem first. Different denominators would make that comparison meaningless.

2 · The denominator, and where it comes from

The priority technique list, version-matched to D1.2. The published coverage baseline uses the 78-technique coordination-and-tool-use set.

System of record. The detection-engineering coverage matrix, with field population verified by query against the SIEM rather than asserted.

3 · The test that puts an item in the numerator

ConditionRequirement
T1 — named and collectedA specific log source and field set is named, and that source is actually being ingested today and retained for at least the stated window. A source that could be enabled does not pass.
T2 — required fields populatedThe minimum fields the technique requires are present and populated in real events, verified by sampling. A schema that permits a field is not the same as a field that carries data.
The metric is T1 and T2 togetherDeliberately the harder of the two available bars.
What is explicitly out of scope hereWhether an analytic exists (D2.3 does not ask) and whether it was proven to work (D1.2).

4 · How a pass gets wrongly claimed

  • Counting a source because the vendor documents it. Verify by sampling events; documentation is not data.
  • Counting a schema field that is defined but empty in practice — the most common false pass in agent telemetry today.
  • Counting a technique as covered by a generic source that would contain the event in principle but has no field distinguishing it.
  • Reporting a single coverage percentage with no robustness dimension, which implies techniques are equally covered when they are not.

5 · How it is computed

Recomputed monthly from the coverage matrix, with the field-population check run as a scheduled query. Owned by detection engineering.

6 · The arithmetic, worked through

LineValue
Priority set — coordination and tool-use techniques78
Pass T1 and T2 by the telemetry modality collected today — the numerator16
Metric16 / 78 = 21%
Would pass under runtime instrumentation, not deployed69 / 78 = 88%

What here is real and what is illustrative. These figures are real, not illustrative — they come from the generated coverage matrix. The 16-versus-69 gap is the central technical fact of the baseline, and establishing whether it is closable with collection we already own is a tranche-1 deliverable.

7 · How this number could lie

THE GAMING RISK, AND WHAT CLOSES IT
Report a single number and imply uniform coverage. The framework the matrix is built on states the limit itself: coverage is “0 direct, 78 partial”. So a technique showing complete coverage indicates a measurement error, not an achievement, and the figure must be published with its robustness dimension.

What the 16-versus-69 gap actually means

It is not a gap in analytics. It is a gap in where the telemetry is taken from: gateway-level observation sees a fraction of what runtime instrumentation sees, because the interesting behaviour happens between an agent and its tools rather than at the network boundary. That makes the gap a collection-architecture decision with a cost attached — which is why decision 06 gates the tranche-2 telemetry line on establishing this number with our own estate rather than a reference one.

One caveat that changes how the number should be read. Some of the most valuable detections need no broad telemetry at all. “An agent invoked a tool it never declared” requires only the registry joined to tool-call records, and needs no behavioural baseline. Of the eleven published hunting queries for agent activity, none performs that join. So a low coverage percentage does not mean nothing useful can be built today.

WHAT THE NUMBER ASKS, IN PLAIN WORDS
In the places an intruder would go for our most valuable data, how often would they touch something whose only purpose is to tell us they are there?
% of crown-jewel environments with at least one live, monitored, tested decoy on the intruder’s path

1 · What exactly one item is

One environment — a network or platform boundary containing at least one store from the crown-jewel register. Environments are countable because the boundary is a configuration object.

Definition and why it is drawn here
Why environments and not decoysCounting decoys rewards volume. Counting environments asks the question that matters: is there anywhere an intruder can reach a crown jewel without touching something that tells us?
Density is tracked separatelyOne qualifying decoy per environment is the pass bar here. Placement quality and density are tracked as a companion, because a single token in a large environment is weak but is not zero.

2 · The denominator, and where it comes from

The set of environments derived from the crown-jewel register (D1.3). It inherits that register, so it inherits its ratification date too.

System of record. The deception deployment inventory, joined to the alert-routing configuration.

3 · The test that puts an item in the numerator

ConditionRequirement
C1 — on the pathReachable by an intruder who has reached that boundary, and placed where enumeration would encounter it — not in an unused subnet.
C2 — monitored, not merely loggedTouching it raises an alert routed to a monitored queue with a defined response procedure. A log entry nobody is watching fails.
C3 — credibleNaming, metadata and age consistent with real assets in the same environment, and not identifiable by the known public fingerprinting techniques.
C4 — tested end to endSomeone touched it and the alert arrived at the queue, within the last 180 days.

4 · How a pass gets wrongly claimed

  • Decoy deployed, alert unrouted. This is the most common failure and it produces a metric that looks healthy and detects nothing.
  • A decoy that is fingerprintable, which turns it into a signal to the adversary that they are being watched.
  • A stale decoy whose metadata no longer matches the environment around it.
  • A decoy placed off the path, where it is safe but never encountered.

5 · How it is computed

Computed from the deployment inventory monthly; the C4 test is scheduled so every decoy is exercised within 180 days. Owned by CDR.

6 · The arithmetic, worked through

LineValue
Crown-jewel environments12
With any decoy deployed0
Passing all four conditions — the numerator0
Metric0 / 12 = 0%

What here is real and what is illustrative. The zero is real — nothing is placed today. The denominator of twelve is illustrative and depends on the crown-jewel register.

7 · How this number could lie

THE GAMING RISK, AND WHAT CLOSES IT
Deploy widely and route nothing. C2 and C4 exist specifically to close that, and they are the two conditions most likely to be quietly dropped under delivery pressure.

Why this is in the first thirty days rather than FY27

Because it needs no inventory, no gateway change and no new telemetry pipeline — and because disruption is where we are weakest. Gartner’s stated benefit for the domain: it “may slow down an attack as an attacker spends time… on a target of no value” and it provides “a clear signal of an attack.”

And why the evidence supports it more strongly than most controls

  • In a controlled study of language-model agents, bait was taken at roughly 78% against a human baseline of about 37% — agents are more susceptible to deception than people, not less.
  • The same work found a recognition-action gap of 73.4%: agents frequently identified a lure as suspicious and interacted with it anyway.
  • Attention diversion was statistically absent — decoys did not distract agents from real objectives, so the control adds signal without adding risk.
  • National-level trials across 121 organisations and 14 vendors give the operational counterpoint: the control fails on routing and response, not on placement.
WHAT THE NUMBER ASKS, IN PLAIN WORDS
For each kind of non-human credential in the estate, could we turn all of them off at once, and have the systems that accept them start refusing, inside ten minutes?
% of machine-identity credential classes revocable cohort-wide in under ten minutes, measured at the resource

1 · What exactly one item is

One credential class, where a class is the tuple (credential type × issuing platform × revocation mechanism). For example: “application-registration client secret in the corporate tenant”, “user-assigned managed identity”, “Kubernetes service-account token in cluster group A”, “artefact-signing key”.

Definition and why it is drawn here
Why classes and not credentialsBecause revocation capability is a property of the mechanism, not of the individual credential. Counting credentials would make the metric move with the size of the estate rather than with our ability to act — it would fall as the estate grew even if capability improved.
Enumerated once, then versionedThe class list is established once and re-ratified quarterly. It is small — on the order of a dozen to twenty entries — which is what makes the metric readable.
Scope ruleAny class with reach into a crown-jewel store must be in scope. Classes leave scope only by Security Committee decision.

2 · The denominator, and where it comes from

The enumerated credential-class list. It is published, because a percentage over an unpublished class list is not auditable.

System of record. The identity platform for issuance; the revocation rehearsal log for the measured times.

3 · The test that puts an item in the numerator

ConditionRequirement
V1 — a named authorityA role that can authorise revocation for that class without a change-advisory board.
V2 — a written procedureCurrent within twelve months.
V3 — cohort capabilityThe mechanism can revoke the whole class, or a defined subset, in one action — not credential by credential.
V4 — a measured time under ten minutesFrom decision to the resource refusing the credential. Measured, not estimated. The measurement point is the resource, not the directory.
V5 — rehearsed within 90 daysWith the measured time recorded against the run.

4 · How a pass gets wrongly claimed

  • Measuring at the directory rather than at the resource. Revoking a secret does not invalidate access tokens already issued until they expire — so a class can look revocable in the directory and remain fully usable at the resource for the token lifetime. Where the platform cannot close that window, the class fails until a continuous-evaluation mechanism is in place.
  • A procedure that exists but has never been executed, so the time is an estimate.
  • Credential-by-credential revocation counted as cohort capability. Under time pressure the difference is the whole control.
  • An authority named in a document but requiring a change window in practice.

5 · How it is computed

Measured at each rehearsal and recorded per class with the run identifier. Owned by Identity as the mechanism provider, with CDR as the authority holder.

6 · The arithmetic, worked through

LineValue
Enumerated credential classes14
With a named authority and a written procedure0
With a measured end-to-end time under ten minutes — the numerator0
Metric0 / 14 = 0%

What here is real and what is illustrative. The zero is real: nothing is defined today — no authority, no target time, no rehearsal. The class count of fourteen is illustrative until the enumeration is done, and that enumeration is the first deliverable.

7 · How this number could lie

THE GAMING RISK, AND WHAT CLOSES IT
Enumerate few classes, or descope the difficult ones. Closed by publishing the class list, by the rule that crown-jewel-reaching classes cannot be omitted, and by requiring Security Committee sign-off for any descoping.

Why ten minutes, and why machine identity specifically

Average adversary breakout time is 29 minutes, with a fastest observed of 27 seconds and a median hand-off to a second-stage operator of 22 seconds. A revocation that takes a change window is not a control. And machine identities are the reversible asset class — revoking one breaks a workload rather than a person’s day, and it can be reissued in minutes. That is what makes pre-authorised action defensible here and nowhere else yet.

Revocation at the credential level is not sufficient

If the adversary holds signing material they can mint valid credentials faster than we withdraw them — which is what happened in the reference case. So the catalogue must include authority-level actions: issue-time invalidation, lease revocation by prefix, signing-authority taint, and continuous session signals.

GARTNER NAMES THIS AS THE HARDEST ASK IN THE PROGRAMME
“it is extremely rare for organizations to be willing to automatically remediate discovered issues due to the concern over potential disruption… Without a cultural shift, many CISOs will be unable to fully utilize preemptive cybersecurity approaches.” We are asking for that shift once, in the narrowest place where it is defensible, with the scope list changeable only by the Security Committee. And the authority should lapse if the rehearsal lapses — which is what condition V5 is for.
WHAT THE NUMBER ASKS, IN PLAIN WORDS
When something real happens, how often does the earliest trace of it turn out to have come from a trap we set, rather than from a customer, a supplier, or luck?
% of confirmed incidents whose first signal, in the post-incident timeline, came from a disruption control

1 · What exactly one item is

One confirmed incident — an event that passed triage into the incident process with a severity assigned. Not an alert.

Definition and why it is drawn here
First signal, defined preciselyThe earliest alert or observation by timestamp that the post-incident review assesses as relating to the incident, whether or not anyone acted on it at the time. This is deliberate: the metric measures what the estate produced, not what the analyst noticed.
The six source categoriesEvery confirmed incident’s first signal is classified as exactly one of: disruption control (a decoy touched, a honeytoken used, a revocation-triggered failure observed); threat hunt; detection analytic; third-party or external notification; user report; discovered during unrelated work. This metric is the first category.

2 · The denominator, and where it comes from

Confirmed incidents in the rolling four-quarter window. The window is four quarters because a single quarter produces a number too small to interpret.

System of record. The post-incident review record. The classification is made at review time and is not revisited afterwards.

3 · The test that puts an item in the numerator

ConditionRequirement
The pass conditionThe post-incident timeline names the first signal, and its source category is disruption control.
Reporting formAlways as numerator and denominator with absolute counts shown, never as a bare percentage. With fewer than about twenty incidents in the window, a percentage alone is misleading.

4 · How a pass gets wrongly claimed

  • Classifying by which alert triggered the response rather than which signal came first. The two are frequently different, and the difference is itself worth reporting.
  • Counting alerts instead of incidents, which makes the denominator a function of tuning.
  • Publishing a percentage on a denominator of three.

5 · How it is computed

Assembled at each post-incident review; the metric is recomputed quarterly over the trailing four quarters. Owned by the incident-response function.

6 · The arithmetic, worked through

LineValue
Confirmed incidents, rolling four quarters23
First signal was a detection analytic11
First signal was a user report6
First signal was third-party notification4
First signal was a threat hunt2
First signal was a disruption control — the numerator0
Metric0 / 23 = 0%

What here is real and what is illustrative. Illustrative distribution. The zero in the disruption row is structurally certain, because no disruption controls are deployed — but the other rows are not the enterprise figures and the classification has not yet been applied retrospectively. Doing that retrospective classification is cheap and would give a real baseline within weeks.

7 · How this number could lie

THE GAMING RISK, AND WHAT CLOSES IT
It cannot easily be gamed, which is its strength. It is also the only one of the fifteen that is a lagging measure — it cannot be improved directly, only by deploying the controls in D3.1 and D3.2 and waiting. Treat a low value as a statement about deployment, not about the incident-response team.

Why carry a metric that cannot be directly influenced

Because it is the only measure that tests whether deception and hunting earn their place rather than merely being deployed. D3.1 counts placement; this counts payoff. Without it, a fully green D3.1 could coexist with a deception capability that has never once been the thing that told us.

And it is the number that would have changed the reference incident. In the reference case the defender’s own account identifies the failure as one of escalation, not detection — signals existed and did not become an incident quickly enough. A first-signal classification applied consistently is how that pattern becomes visible before the post-mortem rather than during it.

WHAT THE NUMBER ASKS, IN PLAIN WORDS
For how many of our agents would an attempt to use a tool they never declared simply be refused — right now, in production, not according to a document?
% of production agents whose registry-declared scope is refused at runtime when exceeded

1 · What exactly one item is

One production agent: an agent with a registry identifier that is either reachable by someone other than its author, or runs on a schedule.

Definition and why it is drawn here
Deliberately not countedExperiments reachable only by their own author. This exclusion is what makes the count stable — without it the denominator tracks developer activity rather than production exposure.
Declared scope, definedThe registry record’s enumerated permitted set: tools, data sources, egress destinations, and a maximum autonomy tier. All four must be enumerated for the record to be scoreable.

2 · The denominator, and where it comes from

Registered agents that meet the production test. Published alongside the total registered count, so the exclusion is visible.

System of record. The the agent registry, joined to the enforcement point's policy decision log.

3 · The test that puts an item in the numerator

ConditionRequirement
E1 — an enforcement point that can refuseCalls traverse a gateway, proxy or sidecar that is able to deny a call, not merely observe it.
E2 — delegation identity resolvedThe enforcement point knows which registered agent, acting for which principal, is making the call — not merely which credential was presented.
E3 — registry evaluated at call timeThe declared set is read from the registry at the moment of the call, so a scope change takes effect without a redeploy. A scope compiled into the agent at build time fails.
E4 — refusals logged and alertedA denial is recorded with the technique-relevant fields and raises an alert.
The deciding evidenceAn out-of-scope tool call is attempted in production and the refusal plus the log entry are observed. Configuration evidence alone does not pass.

4 · How a pass gets wrongly claimed

  • An enforcement point that logs but cannot deny. Observation is not enforcement and the distinction is the entire metric.
  • Scope compiled in at build time, which cannot be changed under incident conditions.
  • Credential-level attribution recorded as delegation identity. Attribution is formally non-identifiable from logs alone, and trace-based grouping recovers only 4 to 6% of a delegation’s events — so if the gateway cannot mint a delegation identity, E2 is partial and the agent fails.
  • Counting an agent because an AI security posture management (AI-SPM) tool has inventoried it. That is a static discipline and cannot deliver this.

5 · How it is computed

Computed from the enforcement point's policy configuration joined to the registry, with the E-condition evidence held per agent. Owned jointly by CDR and AI Enablement — this is a build, not a procurement.

6 · The arithmetic, worked through

LineValue
Registered agents61
Meet the production test — the denominator34
Traverse an enforcement point that can deny (E1)0
Pass all of E1 to E4 — the numerator0
Metric0 / 34 = 0%

What here is real and what is illustrative. The zero is real: nothing binds the registry to runtime enforcement today. The agent counts are illustrative pending the D1.1 enumeration.

7 · How this number could lie

THE GAMING RISK, AND WHAT CLOSES IT
Count agents that pass through a gateway without testing whether it would refuse anything. The production attempt test closes it, and it is cheap to run.

The finding that makes this a build rather than a purchase

AI security posture management cannot deliver this. It is a static configuration and inventory discipline. As one vendor states the distinction: “static AI-SPM tells you what an agent can do; runtime-informed AI-SPM tells you what it actually does.” Gartner’s AI trust, risk and security management framing is reported to acknowledge that it “doesn’t adequately cover the governance of the autonomous actions these models can take.”

SO WRITE IT AS AN EXPLICIT, SEPARATELY-TESTED REQUIREMENT
Do not assume it arrives with a posture-management purchase. The category is consolidating into cloud-native application protection platforms, so AI-SPM will likely appear as a module of something the enterprise already licenses — useful for inventory and misconfiguration, and silent on runtime scope.

The detection this unlocks, which nobody ships

Once E2 and E3 hold, declared scope becomes comparable as well as enforceable — which gives the highest-value agent detection available: an agent invoked a tool it never declared. It needs no behavioural baseline and no anomaly model. Eleven published hunting queries exist for agent activity and none of them performs that join.

One dependency worth testing in week one. Whether the gateway can carry a delegation identity at all — the proposed mechanism is a World Wide Web Consortium (W3C) baggage header carried in the Model Context Protocol request metadata. If it cannot, the fallback is credential-level resolution: weaker, workable, and far better known before the enforcement design is committed than after.

WHAT THE NUMBER ASKS, IN PLAIN WORDS
Of the assets whose compromise would stop production or expose recipe data, how many are checked against a written standard every week, with somebody owning the drift when it appears?
% of critical assets assessed against a named standard at least weekly, with drift raised as an owned finding

1 · What exactly one item is

One asset, at the granularity the assessing tool reports — host, cluster, cloud resource, or tool controller. The granularity is published, because it determines the number and is the easiest thing to shift quietly.

Definition and why it is drawn here
Critical, definedOn the crown-jewel path — holds, processes, or has write authority into a store on the crown-jewel register — or is a production tool controller or factory-software component.
Why factory software is named explicitlySEMI E188 explicitly excludes the manufacturing execution system, the material-control system and factory host systems; E187 addresses supplier-provided equipment. Neither covers the factory-software layer that holds write authority into the tools, and that is the layer reachable from information technology. It is a gap in the industry standards, not in the enterprise’s implementation, so this metric names it rather than inheriting the omission.

2 · The denominator, and where it comes from

Critical assets from the asset register, across information technology, operational technology and cloud — reported with the three sub-populations visible, because their coverage differs greatly.

System of record. The asset register, with the assessing platforms as the evidence source.

3 · The test that puts an item in the numerator

ConditionRequirement
P1 — a named written standardA benchmark or internal baseline, with its version recorded. “Hardened” without a named standard fails.
P2 — at least weekly, unattendedAssessment runs at least every seven days without human initiation.
P3 — drift becomes an owned findingA deviation produces a finding with an owner and a date — not merely a changed dashboard state.
P4 — absence detectionThe platform knows what it is not seeing: an asset present in the register but not reporting becomes a finding. Without P4, unassessed assets are indistinguishable from passing ones.

4 · How a pass gets wrongly claimed

  • Omitting P4. This is the near-universal omission and it is the difference between a coverage figure and a marketing figure.
  • Counting a point-in-time scan as continuous assessment.
  • A dashboard that shows drift without creating an owned finding, so nothing follows from it.
  • Reporting one blended percentage across information technology, operational technology and cloud, which hides the operational-technology position entirely.

5 · How it is computed

Computed weekly from the assessing platforms joined to the asset register, with the three sub-populations reported separately. Owned by CDR with platform and infrastructure operations as the remediating function.

6 · The arithmetic, worked through

LineValue
Critical assets on the register2,140
— information technology and cloud1,870
— operational technology and factory software270
Under a seven-day unattended assessment with drift findings and absence detection — the numerator640
Metric640 / 2,140 = 30%
— operational technology sub-populationapproximately 0%

What here is real and what is illustrative. Illustrative. Cloud posture is partial today, AI-SPM is absent, and the operational-technology estate is largely unassessed — but the register has not been consolidated, so no numerator can be stated yet.

7 · How this number could lie

THE GAMING RISK, AND WHAT CLOSES IT
Two ways, both about the denominator. Define critical narrowly — closed by inheriting the crown-jewel register rather than setting scope inside the metric. And report a blended figure that lets strong cloud coverage mask an unassessed production environment. Closed by publishing the three sub-populations always.

Why assessment and response are deliberately separated for operational technology

Response authority inside the production environment remains deferred to FY27 by decision. Assessment is achievable where response authority is not, so this metric includes operational technology while the containment metrics exclude it. Gartner describes the underlying constraint precisely: “Fragmented IT/OT convergence creates severe risks, as current systems lack unified governance and struggle to support the downtime required for constant patching.”

The segmentation journey this metric sits inside

FindingConsequence for sequencing
Gartner recommends “the journey from macrosegmentation to network security microsegmentation to limit lateral movement.”FY27, not tranche 1.
Discovery first: 30 to 90 days of traffic observation before enforcement; a pilot in 8 to 12 weeks; a first enterprise segment in 3 to 6 months.Enforcing before dependency mapping is the usual cause of outage-driven rollback.
Of fourteen organisations attempting even limited segmentation, eleven failed — predominantly for want of an executive champion and application-owner buy-in.Secure the sponsor and the application owners before the tooling. The failure mode is organisational.
For legacy process equipment, network-enforced segmentation is the only viable primary control — agents cannot run and virtual local area network (VLAN) tagging is often unsupported.The production environment path is network-level and sits behind the FY27 decision.
WHAT THE NUMBER ASKS, IN PLAIN WORDS
Of the workloads that read content written by somebody else, how many cannot reach anywhere we have not explicitly allowed — proven by trying it?
% of untrusted-input workloads with egress denied by default and the denial proven by test

1 · What exactly one item is

One workload — a deployable unit with its own network identity: a container workload, a function, or a virtual machine role. The granularity is published.

Definition and why it is drawn here
Untrusted-input, defined by one testDoes the workload process content it did not author and cannot fully validate? Concretely: inference over user or external content; document and email ingestion; web retrieval; repository-content processing; and Model Context Protocol tool servers reachable by agents. The question to ask is whether the content it reads could contain instructions.
Deliberately not countedWorkloads whose only inputs are internally generated and schema-validated. Including them would dilute the denominator with the easy cases.

2 · The denominator, and where it comes from

Workloads meeting the untrusted-input test, enumerated from the platform inventory rather than declared.

System of record. The platform inventory joined to the network-policy and egress-proxy configuration.

3 · The test that puts an item in the numerator

ConditionRequirement
G1 — explicit allow-listA named list of permitted destinations; everything else refused. A wildcard, or a broad content-delivery range that reaches anywhere, fails.
G2 — enforced where the workload cannot reconfigure itNetwork policy with an enforcing data plane, or an egress proxy the workload must traverse.
G3 — denials loggedWith source workload identity and destination, so a blocked attempt is a signal and not just a failure.
G4 — proven by testAn attempt to a non-allowed destination, made from inside the workload, is refused — demonstrated within the last 90 days.

4 · How a pass gets wrongly claimed

  • A policy authored with no enforcing plane. Kubernetes NetworkPolicy silently does nothing without a container network interface (CNI) plugin that enforces it — this is a documented trap and it produces a confident false pass.
  • An allow-list containing a wildcard or a broad provider range, which permits egress to anywhere behind that provider.
  • Domain name system (DNS) resolution permitted to arbitrary resolvers, which leaves an exfiltration path open regardless of the allow-list.
  • A proxy that can be bypassed by addressing a destination directly by internet protocol address.

5 · How it is computed

Computed monthly from configuration, with the G4 test scheduled per workload class rather than per workload. Owned by platform engineering with CDR defining the untrusted-input test.

6 · The arithmetic, worked through

LineValue
Workloads meeting the untrusted-input test87
Have an explicit allow-list (G1)12
Also enforced at a point the workload cannot change (G2)6
Also logged and proven by test (G3, G4) — the numerator4
Metric4 / 87 = 5%

What here is real and what is illustrative. Illustrative. The untrusted-input population has not been enumerated, which is the first task here.

7 · How this number could lie

THE GAMING RISK, AND WHAT CLOSES IT
Author policies and never test them — which is exactly what the Kubernetes trap makes easy, because the configuration looks correct. G4 is the only condition that catches it, and it is the one most likely to be dropped.

Why this is framed as isolation rather than input filtering

Because prompt filtering is a probabilistic control against an adversary who can iterate, while egress control is a deterministic one. Gartner’s related guidance points at securing agent actions rather than agent prompts. The practical consequence: we do not need to detect a malicious instruction if the workload that receives it cannot reach anywhere useful.

And it is the control that pairs with D4.1. D4.1 constrains what an agent may call; D4.3 constrains where a workload may reach. Together they close the two paths that matter, and either alone leaves the other open. D4.3 is the cheaper of the two and does not depend on the registry, which is why it can proceed while the enforcement-point question is still being settled.

WHAT THE NUMBER ASKS, IN PLAIN WORDS
Of the exposures we raise, how many end up with a named person outside our own team who has said either “yes, by this date” or “no, and here is what covers it instead”?
% of deduplicated exposure findings with a named individual outside CDR who has accepted or rejected them on the clock

1 · What exactly one item is

One deduplicated finding. The deduplication rule is one finding per (weakness × affected asset group × owner)not one per scanned instance.

Definition and why it is drawn here
Why the deduplication rule has to be statedBecause per-instance counting inflates numerator and denominator together and makes the percentage meaningless: a single misconfiguration across four hundred hosts would dominate the figure. Stating the rule is what makes the metric comparable quarter to quarter.
Deliberately not countedInformational findings with no remediation path, and findings whose only owner is CDR itself — the latter are reported separately, because a programme that owns its own findings is measuring nothing.

2 · The denominator, and where it comes from

All deduplicated findings raised in the reporting period.

System of record. The exposure-management workflow system. Email does not count as a record — if the acceptance is not in the workflow, the finding is unowned.

3 · The test that puts an item in the numerator

ConditionRequirement
A1 — a named individual outside CDRRecorded as accountable. A team, a function or a distribution list fails.
A2 — accepted or rejected on the clockThe owner has either accepted with a remediation date, or rejected with a stated reason and a named compensating control — within the acceptance window. Ten working days is the proposed window; the number is a Security Committee decision, not a CDR one.
A3 — in the workflow systemNot in a mailbox, a spreadsheet or a meeting minute.
A rejection is a passThe metric measures whether ownership was resolved, not whether everything gets fixed. This is essential: without it the metric punishes honest risk acceptance and nobody will record a decision at all.

4 · How a pass gets wrongly claimed

  • Assigning to a team. Teams do not accept risk; named people do, and that is the whole point of the board ask.
  • Counting per scanned instance, which makes the figure move with scanner configuration.
  • Treating an unanswered finding as accepted by default. Silence is the failure state this metric exists to make visible.
  • A rejection with no compensating control named, which is a decision to do nothing recorded as a decision.

5 · How it is computed

Computed from the workflow system at the end of each reporting period. Owned by CDR as the raiser; the value is determined entirely by behaviour outside CDR, which is deliberate.

6 · The arithmetic, worked through

LineValue
Deduplicated findings raised in the quarter1,340
With a named individual outside CDR384
— of which accepted with a date, within the window151
— of which rejected with a reason and a compensating control59
Resolved ownership on the clock — the numerator210
Metric210 / 1,340 = 16%
Companion — escalated to the Security Committee unowned174

What here is real and what is illustrative. Illustrative. What is established is that there is no systematic assignment today, and that roughly 20% of the relevant controls sit with CDR — so about four fifths of the remediation capability is outside the team raising the findings.

7 · How this number could lie

THE GAMING RISK, AND WHAT CLOSES IT
Raise fewer findings, or raise only the ones with obvious owners. The counter is the unowned escalation count, published beside the percentage — that number rises when findings are parked and cannot be improved by selective raising.

Why this is the programme’s critical path

Gartner is unusually direct: “cybersecurity teams can only guide vulnerability prioritization. IT operations, product teams and business system owners must be held accountable by the board to fix exposures in their own systems.” And separately: “Without widespread business engagement most exposure management functions… are unable to function effectively.” Only 36% of organisations have infrastructure teams actively engaged on remediation. Gartner also observes that the obstacles here are predominantly non-technical.

THE SEQUENCING RISK IF THIS LAGS THE TOOLING
Gartner warns preemptive offerings may be “limit[ed] to detecting, rather than fixing issues, thereby increasing, rather than reducing overhead.” If validation and posture assessment stand up before accountable resolvers exist, the programme manufactures a findings backlog nobody owns. That is the single most likely way this creates work rather than protection — which is why decision 02 is sequenced before any tooling.

And what the ask actually is. Not goodwill, and not a responsibility workshop. A board instruction naming accountable owners for five dependency areas: Identity for the machine-identity lifecycle; platform and infrastructure operations for segmentation and enforcement binding; developer experience for repository and build policy; manufacturing for the factory-software layer; and the service desk, finance, procurement and human resources for the high-risk procedures. Of the six decisions requested, this and the appetite reframe are the only two that cost nothing — and between them they determine whether the other four are deliverable.

WHAT THE NUMBER ASKS, IN PLAIN WORDS
If we had to analyse attacker code and attacker prompts tomorrow, could we — on our own infrastructure, without a hosted model refusing the work, and in a way that would still stand up months later?
Forensic analysis capability that does not depend on a hosted model’s guardrails — reported as met, partly met or not met

1 · What exactly one item is

One capability. There is no population to divide by, so there is no percentage.

Definition and why it is drawn here
Why this one breaks the “% of” form, deliberatelyGartner’s canonical outcome-driven metric form is “% of X”, and for fourteen of the fifteen we hold to it. Here the population is a single capability. Inventing a denominator to preserve the form would be dishonest, so this is reported as met, partly met or not met, with the condition list shown and the quarterly test result as the supporting number. That gives the committee a trend without a production sitericated percentage.

2 · The denominator, and where it comes from

Not applicable. Reported as a state with four named conditions.

System of record. The capability's own documentation, plus the quarterly refusal-test log.

3 · The test that puts an item in the numerator

ConditionRequirement
F2a — local weightsA self-hosted inference capability on infrastructure the security team controls, with model weights held locally.
F2b — it will do the workIt processes incident artefacts — malicious code, adversary prompts, exfiltrated content — without refusal. Verified quarterly against a standing test set, not assumed.
F2c — evidence handlingArtefact handling aligns with digital-evidence requirements and the environment is documented for later admissibility.
F2d — export-control determinationA written assessment of inference over controlled technical data, because where the model runs and who can reach it changes the answer.
Reporting formMet only when all four hold. Partly met lists which conditions fail. Not met is the position today.

4 · How a pass gets wrongly claimed

  • Treating a hosted model with a commercial agreement as equivalent. The agreement governs data use; it does not govern refusal behaviour.
  • Standing the capability up and never running the refusal test, which leaves F2b an assumption at exactly the moment it matters.
  • Omitting the export-control determination, which is the condition most likely to stop the capability being usable on the artefacts that matter most.

5 · How it is computed

Reviewed quarterly. The refusal test is the only part that produces a number, and that number is the supporting measure.

6 · The arithmetic, worked through

LineValue
F2a — local weightsNot met
F2b — quarterly refusal test passingNot met
F2c — evidence-handling alignmentNot met
F2d — written export-control determinationNot met
StateNot met

What here is real and what is illustrative. This is the real position, not an illustration. The capability does not exist today.

7 · How this number could lie

THE GAMING RISK, AND WHAT CLOSES IT
Report “in progress” indefinitely. The four conditions are binary and individually checkable, which is what the three-state reporting form is for — partly met must name which conditions fail.

Why this belongs in a strategy document at all

Because it cannot be procured mid-incident, and because the failure is documented rather than hypothetical. Hosted models refused legitimate blue-team work at 2.72× the neutral rate across 2,390 real security tasks, and could not be worked around by rephrasing. In the reference incident the defender hit exactly this.

And one requirement that is easy to miss. Every evidence-touching inference must record the model hash, engine version and decode parameters, or the analysis cannot be reproduced later. That is a logging requirement on day one, not a refinement — retrofitting it means the early analysis is the analysis that cannot be defended.

WHAT THE NUMBER ASKS, IN PLAIN WORDS
Of our written response procedures, how many state how many at once they can handle and what happens when that is exceeded — and have been exercised in the last year?
% of named response workflows with a stated concurrency limit, overflow behaviour and an exercise in the last 12 months

1 · What exactly one item is

One response workflow: a named, triggerable procedure with an entry condition and a defined end state. Countable because it has an identifier in the runbook set.

Definition and why it is drawn here
Deliberately not countedGuidance documents with no trigger and no end state. They may be useful, but they cannot be exercised and they cannot be delegated, which is what this metric is about.
Why capacity rather than existenceBecause “we have a runbook” says nothing about whether it survives twenty concurrent invocations — and volume is precisely what changes when the adversary operates at machine speed.

2 · The denominator, and where it comes from

The named workflow set. Published with the count, because a small well-specified set is a better position than a large vague one and the metric should not disguise which we have.

System of record. The runbook repository, with the exercise log as the evidence for the last condition.

3 · The test that puts an item in the numerator

ConditionRequirement
W1 — a stated concurrency limitHow many simultaneous executions the workflow can sustain, as a number.
W2 — overflow behaviourWhat happens when the limit is exceeded: queue, degrade, or escalate. “Undefined” is the answer this metric exists to eliminate.
W3 — a lifecycleA named owner, a review date, and a retirement or supersession state.
W4 — exercised within 12 monthsWith the result recorded.

4 · How a pass gets wrongly claimed

  • A capacity described qualitatively (“scales as needed”) rather than as a number.
  • An overflow behaviour that is actually the absence of one — work silently queuing with no alert is not a defined behaviour.
  • An exercise that tested the happy path at concurrency one, which is the usual form and tests nothing this metric is asking about.

5 · How it is computed

Computed from the runbook repository and the exercise log, reviewed quarterly. Owned by the incident-response function.

6 · The arithmetic, worked through

LineValue
Named response workflows46
State a concurrency limit and overflow behaviour (W1, W2)11
Also have a lifecycle and an exercise within 12 months — the numerator8
Metric8 / 46 = 17%

What here is real and what is illustrative. Illustrative. Detection-engineering and response capacity is not established today, so the workflow set has not been enumerated with capacity attributes.

7 · How this number could lie

THE GAMING RISK, AND WHAT CLOSES IT
Define the set narrowly so that every workflow in it passes. Publish the set size and its change history — a denominator that shrinks between reporting periods is the signal to look for.

Why measure the baseline before automating anything

Gartner projects that AI agents will autonomously manage 25% of incident-response workflows for data-security events by 2028. That is a strategic planning assumption rather than a measurement, and it should be read as one. But either way the implication holds: a workflow whose capacity and end state are undefined cannot be delegated, because there is no specification for the thing taking it over to meet.

And the honest reason this is a foundation metric rather than a response one. Gartner rests the four domains on managed services and capabilities, noting most organisations need partner support “unless they have large and high-level security teams.” A workflow set with stated capacity is also the artefact that makes a partner conversation possible — it is what a service provider would be asked to meet.

Why these three and not a longer list

Because a longer list would let anything in. Each gate names a specific reason the existing security programme cannot be expected to cover the item: the asset is new, the control has broken, or the required speed exceeds human execution. Anything that does not fit one of those three is, by definition, work the existing programme is already accountable for.

The argument againstThe answer
“Gate 2 will be used to smuggle everything in — any control can be said to be under AI pressure.”Fair, and it is the weakest gate. So G2 requires a named mechanism — speed, scale or fidelity — and evidence that the control previously held. “AI makes phishing worse” does not pass; “voice verification no longer works, here is the prevalence figure” does.
“Parking the posture-management and patching work leaves real risk unowned.”It leaves it owned elsewhere, which is where it already was. The failure mode we are avoiding is a programme that inherits every unsolved problem in the estate because it is the newest thing with budget.
“The dependency bucket is where the plan will actually fail.”Almost certainly true, and that is why it exists as a named bucket with dates rather than as an assumption. Gartner’s finding is that these obstacles are predominantly non-technical, and that exposure functions without business engagement are “unable to function effectively”.

One consequence worth accepting openly. Applying these gates shrinks the programme. Eleven core items is a smaller thing than the earlier draft described, and it will read as less ambitious. It is also the only version that can be delivered and measured, and the only version that answers the question actually asked.