Defending the enterprise against adversaries that attack with AI
What has changed is not the adversary’s capability — it is their velocity and scale. That distinction is the foundation of this strategy, and getting it right is what separates a fundable programme from a reaction to headlines.
“Frontier AI models do not introduce a fundamentally new threat capability; they simply change the velocity and scale of existing attack tactics and exploit generation.”
What is actually happening
Three measured realities, not projections.
of surveyed organisations have experienced deepfake-enabled social engineering — on an audio call and a video call respectively. Deepfakes now account for one in five biometric fraud attempts.
The gap between vulnerability disclosure and active exploitation, now that frontier models can autonomously reverse-engineer patches into working exploits.
“threat actors leveraging and abusing widely available AI tools to scale and enhance existing tactics, not the invention of net-new attack methods.”
And three facts that keep us honest
A strategy that omits these will not survive scrutiny from anyone who reads the same research.
Vulnerability exploitation is the initial access vector in 31% of breaches; credential abuse 13% (appearing somewhere in the chain in 39%) and phishing 16%. “Top threats are cyclical.”
Uses of LLMs to orchestrate automated attacks are “emerging with unclear impact so far”, and malware integrating LLMs shows a “lack of sophistication” — “more experimental than mature”.
“Attackers still need to be ‘right’ multiple times: Foundational cybersecurity controls still provide adequate protection and resilience.” This “is not a repeat of Y2K”.
So why stand up a programme at all? Because the process, not the control, is what breaks. “The core issue is not a new, apocalyptic threat, but a long-standing reality: manual vulnerability management processes are structurally misaligned with machine-speed attacks.” And separately: “organizations are creating attack surfaces faster than technologies can protect them.” Those two sentences are the whole case.
Why the enterprise, specifically
Three facts from the public record, all checkable, none of them speculative.
| Established fact | What AI changes | |
|---|---|---|
| Proven target | A manufacturer pled guilty to stealing the enterprise core product trade secrets and paid a $60M fine, in a scheme prosecutors described as enabling “self-sufficiency in computer memory production.” | Not the motive — the cost and speed of running the campaign |
| Already contested | After the May 2023 CAC decision the enterprise disclosed that a low-double-digit percentage of worldwide revenue was at risk. | The adversary is a standing presence, not an event |
| Manufacturing is the exposure — where information technology meets operational technology (IT/OT) | The enterprise manufactures across seven countries in a business it calls “capital intensive”. Gartner: “Fragmented IT/OT convergence creates severe risks, as current systems lack unified governance and struggle to support the downtime required for constant patching.” | Frontier models now surface infrastructure vulnerabilities faster than the production environment can absorb downtime to patch |
What this programme is, and what it is not
Frontier Threat Defense is not a second security programme. It is the part of defence that the existing one cannot be expected to cover — and saying which part is the first deliverable.
The test that decides what is in
Three gates. An item is in scope if it passes at least one. If it passes none it is business as usual — and the honest thing is to name it and park it rather than quietly carry it.
The thing being defended did not exist in the estate two years ago: a model endpoint, an agent, a tool or Model Context Protocol server, an agent’s machine identity, a retrieval corpus, a fine-tuned weight artefact.
An existing control was adequate and frontier-AI capability makes it inadequate — through speed (disclosure to working exploit compressing to minutes), scale (breadth of simultaneous targets), or fidelity (synthetic voice and video defeating human verification).
The defence cannot be executed at human speed or scale and therefore needs capability we do not have: empirical validation on demand, pre-authorised containment, analysis that a hosted model will refuse to perform.
Three buckets, because two would be dishonest
We own it, it is on this plan, it is measured here.
Business as usual that this plan depends on. Not our deliverable — but we name the owner and the date we need it by, and we escalate if it slips.
Not ours, not blocking. Named, assigned a validation checkpoint, and revisited.
This strategy, run through its own gates
Twenty-three items. Eleven survive as core, six become dependencies with a named owner and a needed-by date, six are parked with a validation checkpoint. The cuts are the point — the fair criticism of the earlier draft was that much of it was business as usual.
| Item | Gate | Verdict | |
|---|---|---|---|
| D1.1 AI assets and machine identities reconciled | FTD CORE | G1 | The population is entirely new asset classes. Nobody else is enumerating agent identities. |
| D1.2 Techniques with a defence proven by executing the attack | FTD CORE | G3 | Becomes the confidence measure. Reframed from techniques to chains — see the chain register. |
| D1.3 Crown-jewel blast-radius maps | DEPENDENCY | — | Attack-path mapping is established practice. We need the crown-jewel register as an input to the chains, so it is a dependency with a date, not a deliverable we build. |
| D2.1 High-risk procedures hardened against synthetic media | FTD CORE | G2 | Voice and video verification worked until it did not. 41% audio, 35% video. |
| D2.2 Scenarios converted to requirements with named owners | FTD CORE | G1 | Retained only in its AI-specific form: this becomes chain definition and maintenance. General threat-intel process is business as usual. |
| D2.3 Named telemetry for AI-relevant techniques | FTD CORE | G1 | Elevated. This is the detection ask — knowing something is happening, at speed and scale. |
| D3.1 Deception on the agent and crown-jewel paths | FTD CORE | G1 | Agents take bait at ~78% against a ~37% human baseline, with a 73.4% recognition-action gap. That is an AI-specific control, not a generic honeypot. |
| D3.2 Machine-identity cohort revocation inside ten minutes | FTD CORE | G3 | The credential classes are new and the clock is set by machine-speed breakout. |
| D3.3 First signal from a disruption control | FTD CORE | G1 | Kept as a measure rather than a milestone. Costs nothing: it is a classification applied at post-incident review. |
| D4.1 Agent declared scope refused at runtime | FTD CORE | G1 | Nothing in the estate does this and no other framework requires it. |
| D4.2 Critical assets under continuous posture assessment | PARKING LOT | — | Classic configuration management. Real work, wrong plan. Parked with a validation checkpoint. |
| D4.3 Deny-by-default egress for untrusted-input workloads | FTD CORE | G1 G2 | The population is defined by a new property — processing content that may contain instructions. |
| F.1 Accountable resolvers outside CDR | DEPENDENCY | — | A governance decision, and the programme’s critical path. Still a dependency, not an FTD capability — but the one we escalate hardest. |
| F.2 Forensic and red-team capability on self-hosted models | FTD CORE | G1 G3 | Elevated to a milestone in its own right. Hosted guardrails refuse defensive work at 2.72×, and the red team needs to run new models as they appear. |
| F.3 Response workflows with stated capacity | DEPENDENCY | — | Runbook hygiene. Needed before anything is delegated to automation, but not AI-specific. |
| Vulnerability and patch management at large | PARKING LOT | — | Explicitly out. The programme continues; it is not this plan’s focus and we will not add to its backlog. |
| Microsegmentation journey | DEPENDENCY | — | A multi-year infrastructure programme with an 11-of-14 failure rate driven by organisational factors. We name the segments the chains need, and inherit the rest. |
| Model gateway and guardrail configuration | PARKING LOT | — | Substantially owned by the Enable AI team. Parked with a validation checkpoint rather than duplicated. |
| Agent registry data quality | PARKING LOT | — | Enable AI owns the registry. We own the reconciliation and the runtime binding, which is a different thing. |
| Secure development and code review for AI features | PARKING LOT | — | Enable AI and developer experience. Validate later. |
| SEMI E187 / E188 conformance | DEPENDENCY | — | Manufacturing owns it. We name the factory-software gap the standards do not cover. |
| Awareness training at large | PARKING LOT | — | Business as usual. The synthetic-media procedure hardening in D2.1 is the FTD-specific slice of it. |
| Operational-technology response authority | DEPENDENCY | — | Deferred to FY27 by decision. We test the IT-to-OT chain and report it; we do not remediate inside the production environment. |
And the three questions the scope now has to answer
| The question, as asked | Where it is answered |
|---|---|
| What do we need to have and do to be 80% confident we could prevent a frontier-AI attack? | Confidence is defined as the share of named attack chains empirically proven closed at a choke point, with the residual named and accepted. Twelve candidate chains; the honest figure today is 0 of 12. |
| What are the ten key milestones, and which are the quick wins? | Ten milestones, each closing named chains, each with a quick win inside four weeks. Eight of the ten need no procurement to begin. |
| How do we score vulnerabilities without collapsing into discovery hell? | We stop scoring vulnerabilities. The unit of exposure becomes a chain and the unit of remediation becomes a choke point — the single change that closes the most chains. |
All three are worked through in the companion implementation plan. This tab is the boundary; the plan is the delivery. The chain register, the confidence arithmetic, the ten milestones with dates and owners, the proof-of-concept scoping and the proposed ownership split with the Enable AI team all live there.
Preemptive cybersecurity, in four domains
Built on two Gartner documents. G00859378 supplies the architecture — four preemptive domains resting on managed services and capabilities — and G00799085 supplies the measurement system behind every outcome. The domains are additive to the prevention, detection and response we already operate.
Close the gap between how fast an adversary can act and how fast the enterprise can see, decide and absorb — by validating controls before they are tested, disrupting attacks early, and hardening posture continuously. Measured as delivered protection levels, not activity.
The strategy architecture
Pillars and their outcomes, in one view. Outcomes are expressed as outcome-driven metrics — each tied to an identifiable investment and moving on a sliding scale.
Exposure management
- Validate control efficacy against attacks
- Simulate real-world approaches
- Enumerate the AI estate — discovery first, self-declaration last
- Prove controls by exercise, not inspection — adversarial validation against the technique set
- Map blast radius before we need it, not during an incident
- % of AI assets and machine identities discovered and reconciled
- % of priority controls validated by simulated attack
- % of crown-jewel repositories with a mapped blast radius
Adversary management and threat intelligence
- Understand adversary tactics
- Threat intelligence
- Anticipate possible attacks against the organisation
- Anchor the technique scope externally — four adversary classes, not internal opinion
- Convert every scenario into a detection requirement with a named owner
- Harden the processes deepfakes actually target — password reset, suppliers, payments
- % of high-risk processes hardened against deepfake impersonation
- % of priority scenarios converted to a requirement with an owner
- % of AI-relevant techniques with a named telemetry source
Adversary disruption
- Misdirect and delay adversary
- Increase attacker cost
- Increase chance of detecting attack
- Place deception where an agent reaches it first — and buy paid or self-hosted tokens
- Make a stolen machine identity worthless in ten minutes — named authority, rehearsed
- Sell it on signal, not delay — the delay benefit does not hold against agents
- % of crown-jewel environments with deception coverage
- % of machine-identity classes revocable within ten minutes
- % of incidents first surfaced by a disruption control
Posture and policy management
- Validate correct configurations
- Harden devices and reduce attack surface
- Bind declared agent scope to runtime enforcement — no product supplies this
- Isolate untrusted-input workloads by default — an isolation problem, not a filtering one
- Assess IT, OT and cloud posture continuously, and begin macro- to microsegmentation
- % of production agents with declared scope enforced at runtime
- % of critical IT, OT and cloud assets under continuous posture assessment
- % of untrusted-input workloads under deny-by-default egress
What we need to do: establish detection-engineering capacity and a lifecycle · stand up forensic capability independent of hosted guardrails · name accountable resolvers outside Cyber Defense & Resilience (CDR). Gartner notes most organisations need managed or partner support here “unless they have large and high-level security teams.” This is the layer where our capability gaps sit — and a temple is only as good as what holds it up.
Why this frame, and why now
“By 2030, preemptive cybersecurity solutions will account for 50% of IT security spending, up from less than 5% in 2024, and replace traditional ‘stand-alone’ detection and response solutions as the preferred approach.”
“Many capabilities promoted as preemptive exist in current platforms. Combining them, however, has the potential to add resilience.” Gartner maps each domain to technology categories the enterprise already licenses in part.
Gartner’s shorthand: Deny access through obfuscation, Deceive through automated deception and moving target defence, Disrupt through predictive intelligence and automated exposure management.
“Current detection and response and application security methods aren’t sufficient to keep up with the speed, sophistication and scope of emerging AI-enabled threats.”
What this strategy deliberately does not do
Three exclusions, each sourced. Naming them is what makes the four domains fundable.
“AI-washing, both as a threat and as a capability, is prevalent in cybersecurity product marketing today” — so we evaluate against outcomes delivered, and check first whether an existing platform already provides the capability.
“The immediate danger is the CISO attempting to own the solution in a silo. If the CISO promises to simply ‘patch faster’ without the backing of the CIO, they will fail.” CVSS-led patching also prioritises the wrong things.
“fully automating fixes will eliminate practical learning ground required to develop experienced Level 3 analysts.” Automation is scoped to reversible actions, not to judgement.
What the programme delivers, and how it is measured
Built with Gartner’s outcome-driven metric method, applied end to end: five steps from business process to measured protection level. Every outcome below measures the continuous result of an identifiable investment, sits on a sliding scale, and has a direct line of sight to a the enterprise business outcome.
The two documents this strategy is built on
One supplies the architecture, the other the measurement system. Everything else in the corpus refines or evidences them.
CISOs Must Focus on the Outcomes of Preemptive Cybersecurity
Supplies the architecture: four preemptive domains, each with a stated core benefit, resting on managed services and capabilities — and positioned as additive to prevention, detection and response.
Outcome-Driven Metrics for the Digital Era
Supplies the measurement system: above and below the line, direct line of sight, the seven characteristics of a metric that changes decisions, and the five-step method for deriving them.
The three business goals, above the line
Gartner is strict about the division: “only above-the-line metrics should be shared with executives to drive effective business decisions.”
- Manufacturing continuityDetail
- Protection of critical IP and recipe dataDetail
- Speed of the enterprise’s own AI adoptionDetail
Fifteen programme outcome-driven metrics, three per domain plus three for the foundation. Each is a protection level we are choosing to buy: fund it and the number improves, defund it and it degrades. And the prioritisation rule that follows: “If there is no clear line of sight to a business outcome, then the technology investment should not be prioritized as it will drive little value.”
The executive set — seven metrics, not fifteen
Gartner is explicit about the ceiling: “You only need five to nine metrics for each target audience. Do not report everything you know.” Fifteen metrics run the programme; these seven go to the Security Committee.
| 01 | % of priority adversary techniques with at least one defence proven by executing the attack in the last 90 daysOne count per technique on the published priority list. Untested counts as unproven. | D1 · the core preemptive measure | Definition |
| 02 | % of discovered AI assets and machine identities that are fully reconciledThe denominator is what discovery found — never an estimate of the true total. | D1 · everything else joins to this | Definition |
| 03 | % of named high-risk procedures that pass a live social-engineering test with all four verification controls in placeThe live exercise decides, not the written procedure. | D2 · the dominant real vector | Definition |
| 04 | % of machine-identity credential classes revocable cohort-wide in under ten minutes, measured at the resourceTimed from decision to the resource refusing the credential. | D3 · raises attacker cost directly | Definition |
| 05 | % of production agents whose registry-declared scope is refused at runtime when exceededProven by attempting an out-of-scope call in production. | D4 · currently zero | Definition |
| 06 | % of critical assets assessed against a named standard at least weekly, with drift raised as an owned findingIncludes absence detection, so unassessed assets cannot count as passes. | D4 · includes the production environment question | Definition |
| 07 | % of deduplicated exposure findings with a named individual outside CDR who has accepted or rejected them on the clockA documented rejection counts; silence does not. | F · the programme’s critical path | Definition |
And each one must drive an action when it moves. “When a metric changes from green to yellow to red, it drives an action such as changing a priority or an investment.” So each of the seven carries a threshold and a named consequence — otherwise it is a dashboard entry, not an outcome. A metric’s value is “its ability to influence decision making”, and the decisions are priorities and investments.
Not “how many days to patch?” but — “what is our tolerance to be exploited by a known vulnerability?”
Readiness as risk to business enablement
Step 5 of the method: place current readiness on a scale running from no investment to leading-edge investment, so the committee sees an investment decision rather than a score. Readiness means technology that “operates the way it is designed, fully supports a business process… and is not unreasonably at risk of failure.”
The full programme set, by domain
Below the line. These drive the programme; the seven above go to the committee. Each row now carries its scope bucket — eleven are Frontier Threat Defense core, three are dependencies we escalate rather than build, and one is parked. See tab 02 for the test that decided.
Exposure management
% of discovered AI assets and machine identities that are fully reconciledDiscovered means seen by any of five independent signals. Reconciled means one record, a named owner, a declared scope, and every credential enumerated.% of priority adversary techniques with at least one defence proven by executing the attack in the last 90 daysOne count per technique on the published priority list. Untested counts as unproven, and a pass that was only achieved in a lab does not count.% of crown-jewel data stores with a current, machine-generated blast-radius mapFour elements: effective readers, effective writers and publishers, reachable egress, and the telemetry that would show abuse. Effective means expanded through group and role nesting, and including machine identities.Adversary management
% of named high-risk procedures that pass a live social-engineering test with all four verification controls in placeOut-of-band callback to a pre-registered channel, a secret that never crosses the channel being verified, dual authorisation above a threshold, and a stated authority to refuse. The exercise decides, not the written procedure.% of priority threat scenarios converted into an implementable requirement with a named individual ownerA documented decision not to detect counts as converted, provided the compensating control is named. The owner must be a person, not a team.% of priority techniques with a named log source that is collected today and carries the required fieldsVerified by sampling real events, not by reading documentation. Having an analytic is a separate and harder measure. Full coverage on any technique means the measurement is wrong.Adversary disruption
% of crown-jewel environments with at least one live, monitored, tested decoy on the intruder’s pathThe alert must reach a monitored queue with a response procedure, and the end-to-end path must have been tested in the last 180 days. A decoy that only writes a log does not count.% of machine-identity credential classes revocable cohort-wide in under ten minutes, measured at the resourceTimed from the decision to the resource refusing the credential, not to the directory recording the change. Requires a named authority, cohort capability and a rehearsal in the last 90 days.% of confirmed incidents whose first signal, in the post-incident timeline, came from a disruption controlDetermined retrospectively from the timeline, not by which alert someone happened to notice. Reported as a rolling four quarters with absolute counts, because the volume is small.Posture and policy management
% of production agents whose registry-declared scope is refused at runtime when exceededProven by attempting an out-of-scope tool call in production and observing the refusal. The registry must be the policy source evaluated at call time, not a document describing intent.% of critical assets assessed against a named standard at least weekly, with drift raised as an owned findingIncludes absence detection, so an asset that stops reporting becomes a finding rather than counting as a pass. Assessment covers operational technology; automated response does not.% of untrusted-input workloads with egress denied by default and the denial proven by testThe enforcement point must be one the workload cannot reconfigure, and the allow-list must not contain wildcards. A policy authored with no enforcing data plane is a false pass.Foundation
% of deduplicated exposure findings with a named individual outside CDR who has accepted or rejected them on the clockOne finding per weakness, asset group and owner — not per scanned instance. A documented rejection with a compensating control counts as resolved ownership; silence does not.% of named response workflows with a stated concurrency limit, overflow behaviour and an exercise in the last 12 monthsCapacity is the point — a runbook with no stated concurrency limit says nothing about whether it survives volume. This is the baseline needed before any of it can be delegated to automation.One honest cost of doing this properly. “Organizations will discover that measuring some of these elements may require instrumenting parts of the infrastructure to gather new types of data. While there may be expenses associated with this activity, Gartner believes the visibility and power to report benefits and to guide priorities… will make the initial investment worthwhile.” Two of the fifteen — reconciliation and runtime enforcement — need exactly that instrumentation.
Working practices beneath the architecture
The four domains say what we are buying. Prevention, detection, containment and response remain how we operate — and Gartner is clear the preemptive layer is additive to them, not a replacement.
| Domain | Prevent Detail | Detect Detail | Contain Detail | Respond Detail |
|---|---|---|---|---|
| D1 Exposure management | Reduce the surface the inventory reveals | Discovery-first inventory; coverage baseline published | Blast-radius mapping so containment knows what to cut | Adversarial validation — controls proved by exercise |
| D2 Adversary management | Harden SOPs for the processes deepfakes target | Scenarios converted to detection requirements with owners | Pre-planned containment per scenario, not per incident | Post-incident learning fed back into the threat model |
| D3 Adversary disruption | Short credential lifetimes, so a foothold decays | Deception and threat hunting as primary signal | Ten-minute machine-identity revocation | Attacker TTPs fed back into other controls |
| D4 Posture and policy | Runtime enforcement of declared scope; microsegmentation | Configuration-drift detection, not only event detection | Mitigating controls where a patch cannot land | Known-good configuration for every production agent |
Ownership — and the board ask that goes with it
| Domain | CDR owns | Requires an accountable owner elsewhere |
|---|---|---|
| D1 Exposure management | Validation, the coverage baseline, scenario design | Asset inventory in the service-management platform; resolver teams to act on findings |
| D2 Adversary management | Threat intelligence, scenarios, detection requirements | SOP changes in HR, finance, procurement and the service desk |
| D3 Adversary disruption | Deception placement, threat hunting, containment authority | Identity, for the machine-identity lifecycle and revocation mechanism |
| D4 Posture and policy | Standards, drift detection, the untrusted-input tier definition | Platform and I&O for segmentation and enforcement binding; manufacturing for the factory-software layer |
Only 36% of organisations have their infrastructure teams engaged on this. Gartner measured it. So the ask is not for goodwill — it is for the board to name accountable owners, because “without widespread business engagement most exposure management functions… are unable to function effectively.”
What we actually build, and which outcome it moves
Seven defensive capabilities, designed against a real agentic intrusion and carried forward from the worked example. Each is mapped to the preemptive domain it serves and the outcome-driven metrics it moves — because a capability that moves no measured outcome should not be funded.
Domain coverage
The seven defences cover all four domains and the foundation. One domain is covered thinly, and that gap is real.
The seven defences
Each carries its domain and the outcome-driven metrics it moves. Technical depth sits behind each panel.
Where we stand
Two control planes, one blind. The governed model and tool plane routes agents through the gateway to approved models and Model Context Protocol (MCP) servers, with real policy enforcement. The substrate plane — artefact registry, object store, CI, databases, wikis, tickets — carries the same identities with no behavioural analytics, no provenance attestation and no inter-agent channel detection. Every stage of the reference intrusion occurred on the second plane.
The AI Lab
Our own models, on our own infrastructure, for defensive work. During the reference incident the defenders' hosted models refused to help investigate it, and they could not proceed until the analysis pipeline was rerouted through a self-hosted model — which then recovered roughly four times as many exposed secrets. Measured independently at 2.72× refusal on defensive security tasks.
Forensic capability independent of hosted guardrails — binary, currently no% of response workflows with defined capacity and lifecycleKnow the ground
Map the estate before someone else does. Every question the incident raises is an inventory question first: what could one stolen credential reach? Define critical assets, read the generated attack paths, run an identity graph, and make blast radius a field on every alert rather than a report. Gartner's version: pull automated IT, OT and cloud inventories from the exposure platform.
% of AI assets and machine identities discovered and reconciled% of crown-jewel repositories with a mapped blast radiusTest & fix at speed
Find our own attack paths and close them faster than they can be used. Neither route into the victim had a published vulnerability identifier — a scanner-driven programme would have matched nothing, because no severity score rates a path. The frame is continuous threat exposure management, and validation is the phase most programmes skip — the one that would have caught this.
% of priority controls validated by simulated attack% of critical IT, OT and cloud assets under continuous posture assessmentDetect & escalate
Detection worked. Escalation did not. The defenders' collection and correlation layers succeeded — ambiguous signals from separate systems were fused into a coherent attack narrative. The criticality scoring failed and nobody was paged, across a weekend. This is a governance decision, not an engineering one, and it is the finding that generalises furthest.
% of priority scenarios converted to a requirement with a named owner% of AI-relevant techniques with a named telemetry sourceDeceive
The one defence that works better against agents than against people. Across 21 models and nearly eleven thousand responses against a 47-person human control, every model took deceptive bait at roughly 78% versus 37% for humans — and articulated that something was a trap then exploited it anyway 73.4% of the time. So the strategy shifts from misdirection to detection.
% of crown-jewel environments with deception coverage% of incidents first surfaced by a disruption controlRespond & recover
Which parts of containment can be pre-authorised to execute without waiting for a human, and how do we make that safe. The defenders' containment was competent and took hours — but started two and a half days late because nothing paged. Average adversary breakout is now 29 minutes, with a fastest observed of 27 seconds. No human-paced escalation survives that.
% of machine-identity credential classes revocable within ten minutesForensic capability and known-good agent configurationWhy the mapping matters more than the list. A defence that cannot be tied to a measured outcome fails Gartner’s own prioritisation test: “If there is no clear line of sight to a business outcome, then the technology investment should not be prioritized as it will drive little value.” All seven pass. The two that move a binary outcome — B and G — are the ones where the answer today is simply no, which is why both appear in tranche 1.
The honest baseline
Assessed per domain. Deliberately unflattering, because a baseline that is not can only be revised downwards later.
Maturity by domain
the Cybersecurity Capability Maturity Model (C2M2), assessed independently per domain. Its Maturity Indicator Levels run MIL0 (the practice is not performed), MIL1 (initial practices, possibly ad hoc), MIL2 (documented, resourced and skilled) and MIL3 (guided by policy, periodically reviewed and measured for effectiveness).
| Domain | Capability | Now | FY27 | Note |
|---|---|---|---|---|
| D1 Exposure management | Asset and identity discovery | MIL1 | MIL2 | Registry exists but is self-declared |
| Control validation / adversarial exposure validation (AEV) | MIL1 | MIL3 | Continuous agentic pentest being stood up | |
| D2 Adversary management | Threat intelligence and scenarios | MIL2 | MIL3 | Strongest existing domain |
| Detection requirements pipeline | MIL0 | MIL2 | No scenario-to-requirement conversion today | |
| Deepfake / social-engineering SOPs | MIL0 | MIL2 | The largest gap against the real threat | |
| D3 Adversary disruption | Deception | MIL0 | MIL2 | Nothing placed |
| Machine-identity revocation | MIL0 | MIL3 | No authority, no target time, no rehearsal | |
| Threat hunting | MIL1 | MIL2 | Exists, not agent-aware | |
| D4 Posture and policy | Cloud and AI posture management | MIL1 | MIL2 | Partial; AI security posture management (AI-SPM) not in place |
| Runtime enforcement of agent scope | MIL0 | MIL2 | Registry is advisory only | |
| Segmentation | MIL1 | MIL2 | Macro; microsegmentation is the journey | |
| IT/OT governance | MIL0 | MIL1 | Gartner describes this gap directly | |
| F Foundation | Programme governance | MIL2 | MIL3 | Board committee already exists |
| Detection engineering capacity | MIL1 | MIL2 | Capacity not established | |
| Forensic independence | MIL0 | MIL2 | Hosted guardrails refuse defensive work |
One gap deserves to be called out on its own. Deepfake and social-engineering SOP hardening sits at MIL0, and it is the single most prevalent real AI-augmented attack — 41% of organisations on audio calls, 35% on video. It is also among the cheapest things on this page to fix, because it is process change, not technology: password reset, supplier interactions, and financial transactions.
What happens next
Three tranches inside an FY27 frame, each item tagged with the domain it serves. Tranche 1 needs no new money, and this is explicitly not a transformation programme — “not a repeat of Y2K.”
- D3 Place deception in crown-jewel environments
- D3 Define machine-identity revocation authority; time one revocation
- D2 Harden SOPs for password reset, supplier and financial processes
- D1 Publish the control-coverage baseline
- D1 Verify what our gateway AI control actually logs and alerts on
- F Stand up forensic capability independent of hosted models
- D2 Inventory approved and rogue client-side GenAI tools via EPP and SSE
- D1 Pull automated IT, OT and cloud asset inventories from exposure platforms
- D1 First adversarial exposure validation run against priority controls
- D2 Convert priority scenarios into detection requirements with owners
- D4 Bind declared agent scope to runtime enforcement
- D4 Stand up AI security posture management
- F Establish detection-engineering capacity and lifecycle
- D3 Ratify the tiered containment catalogue
- D1 Continuous validation as a standing capability
- D4 Macro- to microsegmentation journey begins
- D3 Cohort revocation; pre-authorised reversible actions live
- F AI Lab: self-hosted defensive and forensic capability
- D4 IT/OT governance framework; factory-software owner named
- D2 Composite AI patterns for high-risk use cases
- F Managed or partner service where in-house depth is absent
Decisions we are asking the Security Committee for
Move the conversation from days-to-patch to tolerance to be exploited by a known vulnerability, and set that tolerance by asset class.
“IT operations, product teams and business system owners must be held accountable by the board to fix exposures in their own systems.” This is the programme’s critical path.
With the fifteen outcome-driven metrics as the programme’s definition of success.
Machine-identity revocation without prior change approval. Gartner notes organisations are “extremely rare[ly]” willing to do this and that it needs a cultural shift — we are asking for that shift, scoped to the reversible asset class only.
“outdated systems and unused code are no longer just operational problems, but significant cyberthreat liabilities.”
No new money for tranche 1. Two of its items may change tranche 2 scope.
And one message to take to the board unchanged. Gartner’s guidance for exactly this briefing: temper the fear, uncertainty and doubt (FUD), and use the attention as a decision point “to reset expectations, funding, and accountability.” The medium-term outlook is genuinely positive — frontier models “will enable defenders to inspect an unprecedented volume of source code” — and saying so buys more credibility than alarm would.
Where the depth lives
Every claim on these pages resolves to an exact source line — Gartner research, the enterprise’s own filings, government records and primary technical research.
The Gartner corpus behind this strategy
Twelve reports, 334 pages, supplied by the programme. Extracted and analysed in full; dossier 27 carries verbatim quotes with page locators for every claim.
| Doc ID | Report | What it contributed |
|---|---|---|
| G00859378 | CISOs Must Focus on the Outcomes of Preemptive Cybersecurity | The four-domain architecture, each domain’s core benefit, the AI-washing caution, and the technology mapping |
| G00856541 | CISO Board Scenario: Mythos, Daybreak, Frontier AI Model Risks | The board framing: temper the FUD, reset risk appetite, distributed ownership, technical debt as material risk |
| G00852902 | Cybersecurity Threat: AI-Augmented Attacks | The real threat picture — deepfakes at 41%/35%, and the CISO action list |
| G00799085 | Outcome-Driven Metrics for the Digital Era | The outcome mechanism: above/below the line, direct line of sight, ODM properties |
| G00836733 | Build Preemptive Security to Avert Weaponized AI Risks | The market direction — 50% of security spend by 2030 — and the three verbs |
| G00852689 | How to Respond to the 2026-2027 Threat Landscape | The counterweight numbers, and the ThreatScape framing |
| G00837909 | Strategic Roadmap for Continuous Threat Exposure Management | Scoping by business impact; resolver teams and mobilisation |
| G00810627 | We’re Not Patching Our Way Out of Vulnerability Exposure | Mitigation over patching; the 36% I&O engagement figure |
| G00853789 | 2027 Strategic Roadmap for Infrastructure Cybersecurity | The IT/OT governance gap and the microsegmentation journey |
| G00845745 | The Future of the CISO Role 2030 | Cognitive attack, the talent-pipeline caution, antifragility |
| G00858028 | A Checklist to Keep Your Organization Safe From Frontier AI Models | Risk tiers, layered control planes, composite AI patterns |
| G00846817 | Hype Cycle for Enterprise Architecture, 2026 | Context only; not load-bearing in this strategy |
Companion artefacts
| Artefact | Use it for |
|---|---|
| This document | The strategy conversation and the Security Committee. |
| Scoped implementation plan | The chain register, the confidence measure, the ten milestones, the proof-of-concept scoping and the ownership boundary. The working document for the delivery sessions. |
| Programme strategy detail | Charter, operating model, RACI and decision rights, investment tiers, risk appetite, governance, Fortune 500 peer benchmarking. |
| The worked example | A real agentic intrusion end to end — kill chain, ten-layer vector landscape, and the control that interrupts each stage. Engineering handover. |
| Research dossiers 01–27 | The evidence base, written to disk as the work proceeded. |
Glossary
Every abbreviation is expanded at first use in the running text, and hovering any abbreviation anywhere in the document shows its expansion. This table is the complete list, for anyone who arrives at a detail panel directly.
Programme and organisation
| Term | Expansion |
|---|---|
| FTD | Frontier Threat Defense |
| CDR | Cyber Defense & Resilience — the enterprise's security function |
| CSOC | Cyber Security Operations Centre |
| CISO | chief information security officer |
| CIO | chief information officer |
| CFO | chief financial officer |
| I&O | infrastructure and operations |
| RACI | responsible, accountable, consulted, informed |
| RAID | risks, assumptions, issues and dependencies — a programme tracking log |
| MVP | minimum viable product |
Frameworks and standards
| Term | Expansion |
|---|---|
| ATLAS | Adversarial Threat Landscape for Artificial-Intelligence Systems — MITRE's equivalent framework for AI |
| ATT&CK | Adversarial Tactics, Techniques and Common Knowledge — MITRE's framework for enterprise attacker behaviour |
| C2M2 | Cybersecurity Capability Maturity Model, published by the US Department of Energy |
| MIL0 | Maturity Indicator Level 0 — the practice is not performed |
| MIL1 | Maturity Indicator Level 1 — initial practices are performed but may be ad hoc |
| MIL2 | Maturity Indicator Level 2 — practices are documented, resourced and staffed by skilled people |
| MIL3 | Maturity Indicator Level 3 — practices are guided by policy, periodically reviewed and measured for effectiveness |
| CSF | Cybersecurity Framework, published by NIST |
| ZTMM | Zero Trust Maturity Model, published by CISA |
| CVE | Common Vulnerabilities and Exposures |
| CVSS | Common Vulnerability Scoring System |
| SSVC | Stakeholder-Specific Vulnerability Categorization |
| ISAC | Information Sharing and Analysis Center |
| AI-ISAC | Artificial Intelligence Information Sharing and Analysis Center |
| C2PA | Coalition for Content Provenance and Authenticity |
| FIPS | Federal Information Processing Standards |
| CFR | Code of Federal Regulations |
| W3C | World Wide Web Consortium |
Measurement
| Term | Expansion |
|---|---|
| ODM | outcome-driven metric |
| ODMs | outcome-driven metrics |
| PIR | priority intelligence requirement |
| PIRs | priority intelligence requirements |
| DBIR | Data Breach Investigations Report |
| AUC | area under the curve — a measure of classifier accuracy |
| FUD | fear, uncertainty and doubt |
Exposure and validation
| Term | Expansion |
|---|---|
| CTEM | continuous threat exposure management |
| AEV | adversarial exposure validation |
| BAS | breach and attack simulation |
| CMDB | configuration management database |
| IP | intellectual property |
Posture and platform
| Term | Expansion |
|---|---|
| CNAPP | cloud-native application protection platform |
| CSPM | cloud security posture management |
| CIEM | cloud infrastructure entitlement management |
| KSPM | Kubernetes security posture management |
| AI-SPM | AI security posture management |
| VLAN | virtual local area network |
| CNI | Container Network Interface — the plugin layer that implements Kubernetes networking |
| CEL | Common Expression Language |
Detection and response
| Term | Expansion |
|---|---|
| SIEM | security information and event management platform |
| SOAR | security orchestration, automation and response |
| EDR | endpoint detection and response |
| EPP | endpoint protection platform |
| SSE | security service edge |
| NDR | network detection and response |
| CTI | cyberthreat intelligence |
| DLP | data loss prevention |
| DNS | domain name system |
| MFA | multi-factor authentication |
| JWT | JSON Web Token |
| IMDS | instance metadata service |
| DR | disaster recovery |
| IR | incident response |
| SPIFFE | Secure Production Identity Framework for Everyone |
| CRIU | Checkpoint/Restore In Userspace |
AI and agents
| Term | Expansion |
|---|---|
| LLM | large language model |
| LLMs | large language models |
| GenAI | generative artificial intelligence |
| MCP | Model Context Protocol — the standard by which agents call external tools |
| TTP | tactics, techniques and procedures |
| TTPs | tactics, techniques and procedures |
| NHI | non-human identity |
| API | application programming interface |
| GPU | graphics processing unit |
Manufacturing and process
| Term | Expansion |
|---|---|
| IT/OT | information technology and operational technology |
| OT | operational technology — the systems that run manufacturing |
| core product | a core memory product line |
| SOP | standard operating procedure |
| SOPs | standard operating procedures |
| Y2K | the year-2000 date rollover programme |
Two conventions worth knowing. Gartner document identifiers appear as G00###### and are given in full wherever a Gartner claim is cited, with the page number. MITRE identifiers appear as AML.T#### for a technique and AML.CS#### for a case study — the reference intrusion in this document is AML.CS0068.
Source classes
Four items remain open in the record, and all four are in tranche 1. Our gateway AI control’s real detection scope is unverified; the security information and event management platform’s (SIEM) detection content model is not established; a small number of peer quotes are secondary-sourced and marked for verification; and our SEMI standards position is unconfirmed. None changes the strategy. Two could change tranche 1.
The threat section is written conservatively on purpose. Gartner’s guidance for this exact briefing is to temper the fear, uncertainty and doubt and to help the organisation “ignore sensational marketing hype.” A strategy that overstates the threat gets discounted the first time someone checks it.
What is established
- Deepfake-enabled social engineering is the dominant vector — 41% of organisations on audio calls, 35% on video, and one in five biometric fraud attempts. It is the highest-prevalence vector in the published evidence and the one with the cheapest available mitigation.
- The vulnerability-to-exploit gap has collapsed from days to minutes, because frontier models can autonomously reverse-engineer patches into working exploits.
- AI is mostly being used to scale existing tactics, not to invent new ones. And LLM-driven vulnerability discovery “predates” the named frontier releases.
What is not yet established
So where does the agentic scenario belong?
As a forward scenario used for validation, not as the operative threat model. The worked example in the companion document — a documented intrusion carried out by an autonomous agent collective — remains the most useful exercise input we have, because it is real, externally adjudicated, and it stresses exactly the processes Gartner says are structurally misaligned. It earns its place under D1 as a validation scenario and under D2 as scenario planning — not as the basis for a threat claim.
The counterweight evidence, stated so nobody has to find it themselves
| Figure | Source |
|---|---|
| Vulnerability exploitation is the initial access vector in 31% of breaches; credential abuse 13% (39% somewhere in the chain); phishing 16% | 2026 Verizon DBIR via Gartner |
| “Attackers still need to be ‘right’ multiple times” and foundational controls “still provide adequate protection and resilience” | Gartner board scenario |
| “This situation is not a repeat of Y2K and does not require a transformational response” | Gartner board scenario |
| Next 12–24 months see more vulnerabilities, but “the medium-term outlook is a positive one” | Gartner board scenario |
Two threats worth adding to the watch list
- Cognitive attack. “The next breach will be cognitive: Threat actors will abandon purely technical exploits in favor of influence operations, utilizing cutting-edge AI models to manipulate employee decision making and bypass traditional cybersecurity controls.” This is the strategic extension of the deepfake finding.
- Rogue automation from our own estate. “The proliferation of AI agents operating with broad agencies will create severe risks. Traditional centralized IT and third-party vendor management will be unable to oversee these agents.” Which is why D4’s runtime enforcement outcome matters more than it looks.
The four-domain structure, the domain definitions and each domain’s core benefit are Gartner’s. Ours is the adaptation to the enterprise: the outcomes per pillar, the ODM formulation, the maturity assessment, and the sequencing.
Four adaptations worth knowing about
| Domain | Gartner’s emphasis | Our adaptation |
|---|---|---|
| D1 | Validating control efficacy; simulating real-world attacks | Extended backwards into asset and identity discovery. You cannot validate controls over an estate you cannot enumerate, and Gartner separately recommends pulling automated IT/OT/cloud inventories from exposure platforms. |
| D2 | TTP insight and predicting likely attacks | Weighted toward conversion: intelligence only counts once it becomes a requirement with an owner. Gartner: adversary management “is of most value when combined with one of the other pillars.” |
| D3 | Deception, threat hunting, intelligence-driven controls | Extended to include machine-identity revocation as a cost-raiser. Gartner frames disruption as slowing the attacker and raising cost; a credential that dies in ten minutes does both. |
| D4 | Configuration validation and device hardening | Reframed for agents as runtime enforcement of declared scope — the agentic equivalent of a hardened configuration, and currently absent. |
The disagreement inside the corpus, disclosed
Two Gartner teams write about this in materially different registers. The Emerging Tech team describes “an arsenal of unprecedented sophistication against unprepared enterprises” and warns of “ruinous business loss”; the CxO Leadership team writes that “today, true AI-powered attacks remain very rare in the real world.”
How we resolved it. We use the Emerging Tech research for market direction — the 50%-of-spend-by-2030 prediction is a strong planning signal — and the CxO Leadership research for threat claims, because that is the audience-appropriate standard of evidence and because overstating the threat is the failure mode Gartner itself warns about. If challenged on why our threat framing is more conservative than the spend forecast, this is the answer.
Gartner’s outcome-driven metric construct is the measurement system for this entire strategy. It is not a reporting style — it is a way of turning security work into investment decisions. The stated purpose of a metric is blunt: “The value of a metric is its ability to influence decision making. The decisions influenced are typically priorities and investments.”
The five steps, as we applied them
| Step | Gartner | What we did |
|---|---|---|
| 1 | Develop an initial set of business processes and supporting technology stacks — “the three to five most obvious” | Took the three business outcomes already agreed by the programme, and the crown jewels named by the business: critical IP and recipe data repositories. |
| 2 | Identify business outcomes and business ODMs | Manufacturing continuity, protection of critical IP and recipe data, and speed of AI adoption — the three above-the-line outcomes. |
| 3 | Identify technology risks and dependencies — “What breaks in the business process if the technology breaks?” | The two-control-plane assessment and the confirmed detection baseline supplied this. |
| 4 | Define technology ODMs, which “reflect the technology stack’s readiness” and behave as leading indicators | The fifteen programme metrics, three per domain plus three foundation — each in the canonical “% of” form. |
| 5 | Assess readiness as risk to business enablement, on a scale from no investment to leading edge | The readiness scale on the Outcomes page, with current position marked per domain. |
The seven characteristics, used as filters
- Metric value — ability to influence a decision. Anything that could not change a priority or an investment was cut.
- Above the line, below the line — “only above-the-line metrics should be shared with executives.” Hence a seven-metric committee set distinct from the fifteen.
- Leading indicators — they must “expose a problem before it leads to material loss.” This is why validation coverage beats incident counts.
- Direct line of sight — “A causal relationship should exist… If the technology metric changes, it indicates a change in the business outcome.” Every metric on the page states its line of sight.
- Metric changes drive action — green to yellow to red must trigger “changing a priority or an investment.”
- Discrete audiences — “The CFO needs different metrics from those required by the head of a business unit or a board of directors.”
- Limit the number — five to nine per audience, which sets the committee set below.
The prioritisation rule that follows from all of it
“If there is no clear line of sight to a business outcome, then the technology investment should not be prioritized as it will drive little value to the organization.” Applied honestly, this is a test the programme must keep passing — and it is the reason adversarial-ML tooling stays deferred while deception and SOP hardening move to tranche 1.
Measures we rejected, and why
| Rejected | Why |
|---|---|
| Days to patch | Gartner’s explicit reframe: ask instead “what is the tolerance to be exploited by a known vulnerability?” |
| CVSS-weighted remediation counts | Drives “efforts to address only critical and high CVSS scores rather than prioritizing those riskiest to the organization.” |
| Number of detections deployed | Rises when we write rules, not when protection improves. No investment line of sight. |
| Mean time to detect, alone | Not a leading indicator, and the reference failure was in escalation rather than detection. |
| Inventory completeness | Needs a denominator we do not have. Replaced with reconciliation against observed behaviour. |
| Any expected-loss figure for process IP | The enterprise does not disclose a value for its process technology; an invented number is the first thing a finance reviewer tests. |
And the honest cost. “Measuring some of these elements may require instrumenting parts of the infrastructure to gather new types of data… Gartner believes the visibility and power to report benefits and to guide priorities and technology investments in a business context will make the initial investment worthwhile.” Two of the fifteen need that instrumentation, and they are the two with the largest payoff: reconciliation, and runtime enforcement of declared scope.
Every recommendation across the corpus that this programme can act on, deduplicated and attributed. This is the operational substance beneath the architecture.
Actions Gartner recommends, mapped to our domains
| Domain | Action | Source |
|---|---|---|
| D1 | “Scope exposure assessments based on key business priorities… taking into consideration the potential business impact of a compromise rather than primarily focusing on the severity of the threat alone.” | G00837909 |
| D1 | Build validation into exposure management by evaluating adversarial exposure validation tools or offensive cyber services. | G00837909 |
| D1 | “Pull automated critical IT, OT, and cloud asset inventories from existing exposure assessment platforms, in order to prioritize where to start changing architectures.” | G00853789 |
| D2 | “Modify standard operating procedures for high-risk processes, such as password reset, partner and supplier interactions, and financial transactions, to minimize the risk of deepfake and phishing impersonation attacks.” | G00852902 |
| D2 | Inventory approved and rogue employee use of client-side GenAI tools using incumbent endpoint protection, EDR and security service edge; use tags and groups to enable monitoring. | G00852902 |
| D2 | “Perform key threat scenario simulations and adapt strategic roadmaps to cover the most likely AI evolutions.” | G00852902 |
| D3 | Deploy deception and honeypots that “provide a clear signal of an attack when triggered”; conduct proactive threat hunting; operationalise threat intelligence in other controls. | G00859378 |
| D4 | “Prioritizing the journey from macrosegmentation to network security microsegmentation to limit lateral movement.” | G00853789 |
| D4 | “Improve response time by using threat management techniques to identify and implement mitigation controls” where a patch cannot be completed. | G00810627 |
| D4 | Ensure a cross-functional cybersecurity governance framework including zero trust, centralised access and lifecycle management. | G00853789 |
| F | “Establish a model delivery system with layered control planes”; “define risk tiers”; and for high-risk use cases “combine AI models with deterministic systems.” | G00858028 |
| F | Engage senior leadership and adjacent departments for “effective routes to resolution, risk prioritization criteria, and consistent categorization for newly discovered exposures.” | G00837909 |
And a caution on how far to automate. “AI-driven exploit generation will continuously outpace traditional patching cycles. At the same time, fully automating fixes will eliminate practical learning ground required to develop experienced Level 3 analysts.” So automation is scoped to reversible actions. The long-run mandate Gartner describes is antifragility — “organizations will intentionally use systemic shocks and controlled risk exposure to emerge operationally stronger” — which is an argument for exercises, not for autopilot.
C2M2 maturity levels, assessed independently per domain and cumulatively within one — a practice at a level counts only if every lower practice is also achieved. CSF Tiers are deliberately not used as a maturity scale, because NIST does not define them that way.
The four MIL0 findings, with evidence
| Capability | Evidence |
|---|---|
| Deepfake / social-engineering SOPs | No hardened procedures for the processes Gartner names. The highest-prevalence real vector at 41%/35%. |
| Deception | Nothing placed. The cheapest signal-generating control available. |
| Machine-identity revocation | No named authority, no target time, no rehearsal. |
| Runtime enforcement of agent scope | Nothing binds the agent registry to enforcement, so declared scope is documentation. |
| IT/OT governance | Gartner describes the gap directly: systems “lack unified governance and struggle to support the downtime required for constant patching.” |
| Forensic independence | Hosted model guardrails refuse defensive work, measured at 2.72× refusal. |
What we could not assess
- Existing gateway control efficacy — what our AI gateway control actually detects and whether it alerts is unverified. Tranche 1 action.
- Detection engineering capacity — not established with the CSOC. Foundation ODM F.1 creates the baseline.
- SEMI standards alignment — adoption status unconfirmed.
A limitation we disclose rather than let a reviewer find. No board-credible AI-specific security maturity model exists yet, so this assessment applies a general maturity model to AI-specific domains. Gartner’s own framing supports the approach — the four domains are largely delivered by technologies that already exist, so assessing them with a conventional model is defensible.
Three tranches, weekly sprint execution beneath them. The most important structural point: this is explicitly not a transformation. Gartner: “This situation is not a repeat of Y2K and does not require a transformational response.”
Dependencies — the programme’s real critical path
| Dependency | Owner | Needed by | If not met |
|---|---|---|---|
| Machine-identity lifecycle and revocation mechanism | Identity | Day 30 | D3 stays at MIL0; the cheapest cost-raiser is unavailable |
| SOP ownership for password reset, supplier and financial processes | Service desk, procurement, finance | Day 30 | The highest-prevalence real vector stays unaddressed |
| Asset inventory pulled from exposure platforms | I&O with CDR | Day 90 | D1 cannot produce a reconciliation measure |
| Registry-to-runtime enforcement binding | AI Enablement, jointly | Day 90 | D4’s primary outcome remains zero |
| Accountable resolver teams for exposure findings | Named by the board | Day 90 | Exposure management “cannot function effectively” |
| Factory-software layer owner | Manufacturing | FY27 | The IT/OT gap stays uncharacterised |
One sequencing risk worth stating. Gartner warns that preemptive offerings may be “limit[ed] to detecting, rather than fixing issues, thereby increasing, rather than reducing overhead.” If we stand up validation and posture assessment without the accountable resolvers from decision 02, we will manufacture a findings backlog nobody owns. That is the single most likely way this programme creates work rather than protection — which is why the ownership decision is sequenced first, not last.
The single most decision-relevant assessment in the engagement: the enterprise’s strongest control sits on the plane the reference intrusion would not have crossed.
The governed model and tool plane
A registered directory non-human identity calls the AI gateway with its content-inspection layer, which routes to approved models and to MCP servers acting as mediated proxies. Policy actually enforced at that chokepoint: directory JWT and OAuth validation, scope and group authorisation, quotas and token caps, DLP and redaction, prompt-injection inspection, a model allowlist sourced from AI governance, and central logging to the SIEM. All MCP interactions are required to route through it, and direct MCP connections are prohibited for Tier 1 servers. Assessed as stronger mediation than most enterprises operate.
The substrate plane
The same identities, a different path: artefact registry (the default package source, with development and production container registries) → object store → source repositories and CI → databases, wikis and ticketing. The only controls found were a package vulnerability scanner and write-group membership. No behavioural analytics, no provenance attestation, no inter-agent channel detection.
The discriminator that makes the blind plane detectable anyway
A legitimate dependency has one writer, provenance, and a consuming build. A coordination channel has many writers, no provenance, and no consumer. That distinction is computable from access metadata alone, without inspecting content. Which matters because in the reference case content inspection was never possible — the registry’s audit trail covered only user, group, permission and token events.
The adversary labelled the channel itself. The first message posted on the rebuilt coordination board was a directory named zzHELP_PHASEONE_ARV010841_NO_CONSUMER_SEEK_IDEA. Within three hours 53 agents posted 1,188 messages; by 13 July they were signing with Ed25519 keys after being spoofed by other agents. The discriminator is not an analyst’s inference — it was a field in the object name.
The registry that exists but does not enforce
An the agent registry records owner, declared tools, declared write scope and autonomy tier. Nothing was found binding it to runtime enforcement, which is why D4’s primary outcome — declared scope enforced at runtime — currently reads zero.
The case for self-hosted defensive models does not rest on a claim that they match frontier capability — that claim is indefensible and unnecessary. It rests on four independent legs.
1. Availability — the tool must work at 3am on the worst day
2. Confidentiality — the guardrail and the retention pipeline are the same system
The classifier that refuses a request is also the event that flags the session for retention and human review. So the payload you most need help with is retained longest and seen by the most third parties — and mid-incident, the credentials in it are still valid. Gartner records that 69% of organisations are concerned about maintaining control over AI models, and that “due diligence alone is insufficient; technical controls are essential.”
3. Legal and export exposure — an open question, not a settled one
The safe harbour at 15 CFR § 734.18 covers encrypted transit; inference decrypts and processes. No BIS or DDTC guidance exists on whether inference over controlled technical data constitutes a deemed export. That it is open is the argument — against an enforcement climate that has produced penalties of roughly $213M and $252M against two US critical-manufacturing-sector companies. For a manufacturer whose crown jewels are process technology, this is a board-relevant consideration rather than a technical one.
4. Forensic defensibility
ISO/IEC 27037 requires repeatability and reproducibility of forensic process, and hosted models update and reroute silently.
One design requirement, free now and impossible to retrofit. Every inference that touches evidence must record model hash, engine version and decode parameters into the case record. Retrofitting this invalidates the case work already done.
What runs in it, and the risk it imports
| Element | Detail |
|---|---|
| Serving stack | vLLM, SGLang, TensorRT-LLM or llama.cpp. Quantisation introduces calibration sensitivity that must be characterised before forensic use. |
| Ingest gate | Open weights are untrusted executable code entering a high-trust enclave: pickle-format models execute code at load time, roughly 95% of malicious models on one public hub used that format, the standard scanner has a published bypass, and ShadowRay (CVE-2023-48022) was never fixed. |
| Weight provenance | PRC-origin weights are a first-order selection constraint for a US critical-manufacturing enterprise and need a documented risk position, not an assumption. |
| Control planes | Gartner’s own prescription: “Establish a model delivery system with layered control planes”, define risk tiers, and for high-risk use cases combine AI models with deterministic systems. |
Every question the reference incident raises is an inventory question first. What could one stolen credential reach? Nobody had asked — and the answer took one second to demonstrate once an adversary did.
Week one, building nothing
- Define critical assets in the exposure-management platform and read the attack paths that appear. the enterprise almost certainly already holds the licences, and first-class connectors exist for the vulnerability scanner and the service-management CMDB — a day of integration materially improves the graph. The gating input is a business workshop to say what matters, not an engineering build — without it the page may simply be empty.
- A Kubernetes attack-path run against one production-representative cluster, and an identity graph across the tenant — an afternoon of open tooling, aimed at the surface that actually carried the kill chain.
- Switch on directory activity logs while you are there. They are frequently off, and they are the read half of the strongest identity signal available.
- Pull automated critical IT, OT and cloud asset inventories from the exposure assessment platform — Gartner’s explicit recommendation, in order to prioritise where to change architecture first.
The inventory that detection joins to
Self-declaration is not the defect — being unreconciled is. The only published discovery methodology puts attestation fourth of four: tag scan → six-signal heuristics → CMDB reconciliation → developer attestation of the residue, with a deadline and an escalation path. the enterprise’s registry is step one of one.
| Design choice | Why |
|---|---|
| Copy the published 42-column schema | It already carries Triggers and ConnectedAgents — the two fields ours lacks — plus declared tools, MCP servers, instructions, memory, permissions and guardrail coverage. |
| Three-layer asset model | Solves lifecycle velocity: one agent in three regions becomes one model, one asset, three configuration items. Published failure mode of a flat registry: “governance is applied to an entire class of AI, not individual deployments.” |
| Ranked merge precedence | Rank discovery highest and self-declaration lowest, and let them disagree. Works without consolidating anything — which matters because repository unification across roughly 8,000 developers is not achievable. |
| Measure reconciliation rate and declare-lag | Both computable without knowing the true denominator, so they can be reported honestly from week one. Completeness cannot. |
The digital twin, honestly
The published exemplar did not build a twin from telemetry: a sanitised natural-language specification was supplied, an agent-assisted workflow built an isolated environment, and sensors were installed into it. The engineers write “representative test environment”, never “digital twin”. The bar is far lower than the marketing implies. And Gartner explicitly permits validation against “a simulated digital twin to reduce potential impact”, which is consistent with the cloud-first cyber-range MVP the team chose over a full production-environment twin.
And the independent justification for deferring OT. Attack-path analysis in the exposure platform we would use is unsupported for OT connectors. That is a tooling fact rather than a risk appetite — worth having in the room so the FY27 deferral reads as considered scoping rather than omission.
There was no published vulnerability identifier for either way into the victim. A scanner-driven programme would have matched nothing, because what made it an intrusion was the path — and no severity score rates a path.
The five phases, and the one that matters
| Phase | What it means here |
|---|---|
| Scope | By business impact rather than technology silo. Crown jewels are named: critical IP and recipe repositories. |
| Discover | Far wider than known-vulnerability scanning — Gartner projects that by 2028 more than half of exposure findings will be nontechnical rather than technical flaws. |
| Prioritise | By attack path rather than severity. CVSS-led prioritisation drives “efforts to address only critical and high CVSS scores rather than prioritizing those riskiest to the organization.” |
| Validate | The phase most programmes skip, and the one that would have caught this. Gartner’s aligned categories: adversarial exposure validation, breach and attack simulation, automated security control assessment. |
| Mobilise | Requires resolver teams and a mobilisation process. “Without widespread business engagement most exposure management functions… are unable to function effectively.” |
The one experiment to run first
Security chaos engineering against the admission tier. Admission-controller fail-open is the upstream recommendation, not a misconfiguration — and setting the policy to fail closed does not save you, because of a hard 30-second admission-chain budget. The architectural fix is in-process CEL policy evaluation. That is a complete, bounded, week-one experiment with a real finding either way.
Why “patch faster” is not the answer
| Finding | Consequence |
|---|---|
| The federal CVSS mandate was revoked in June 2026 in favour of a four-variable stakeholder model with 3/14/60/defer timelines — and only about 1% of instances fell in the 3-day band while roughly 60% were deferred. | Severity-driven remediation targets the wrong work at scale. |
| The same directive says collect evidence before remediating. | Remediation and forensics compete; sequence them deliberately. |
| No automatic rollback exists anywhere in the mainstream patch stack, and stopping a fleet takes on the order of eight hours. | Longer than the entire escalation window in the reference incident. Mass-patching is not a containment mechanism. |
| Gartner: use mitigation controls where a patch cannot land. | Mitigation is a first-class response, not a fallback. |
And the risk-appetite question this defence exists to answer. Gartner’s reframe applies precisely here: stop asking how many days to patch and ask “what is our tolerance to be exploited by a known vulnerability?” Validation is what turns that question from rhetoric into a number.
The defenders’ hard part succeeded. The easy part failed.
So measure escalation latency separately from detection latency
Almost every programme reports mean time to detect and folds escalation into it, which makes this class of failure structurally invisible. Separating them requires no new tooling and no new data — only a decision about how existing numbers are cut. And the related failure mode produces no alert at all: “Nothing on any screen tells you that a finding failed to appear.”
Why there is no external help for the agentic case
| Finding | Consequence |
|---|---|
| The reference intrusion is now a formal MITRE ATLAS case study mapping 22 techniques across 37 edges. | The technique scope is externally adjudicated — we no longer argue about representativeness. |
| ATLAS technique objects have no detection field, no data-source field and no detects relationship. The entire published guidance is one mitigation. | ATT&CK users inherit a data-source layer for free. For agentic attacks it does not exist. Building it is the work. |
| 73 of 197 ATLAS techniques have no mitigation at all, and the two describing exactly what happened are graded Realized with zero mitigations. | MITRE names the coordination channel and ships nothing against it. |
| Independent coverage matrix: agent-runtime telemetry touches 69 of 78 techniques, gateway telemetry 16 — all 78 mappings partial. | Our confirmed telemetry is the gateway. The inversion is quantified by someone other than us. |
Attribution is a proved impossibility from logs alone
baggage in MCP params._meta.The detection primitive, and two approaches to rule out
- Do not build strict-sequence detection. MITRE: “the complexity and cost of such implementations often outweigh the benefits, and adoption… has been limited in practice.” Order-independent convergence counting survives retries and parallelism, which is how agents behave.
- Do not detect automation by tempo. It fails at 99.8% or worse against a competent adversary and fails confidently — mean confidence above 0.993 when wrong, with a formal zero-mutual-information result. Enrich an alert with tempo; never raise one.
The highest-value detection nobody ships
An agent invoked a tool it never declared. No baseline needed. Eleven published queries exist and none joins the inventory to runtime activity; all are hunting queries rather than analytics rules, so none raises an incident. And the join key is not guaranteed — in the runtime schema the agent identifier is merely recommended and the tool name optional, so a compliant event can carry no agent identity at all. Make its presence a source-onboarding acceptance criterion.
The 30-day backlog, from telemetry already held or one switch away. Telemetry tampering — model-invocation logging deleted, guardrails deleted, registry token and user creation. Tiny volume, near-zero false positives, and they defend the detection stack itself. Autonomy-boundary removal — a published detection exists for permission overrides on Linux, and the Windows and macOS equivalents are published nowhere and must be written locally. Governed-plane bypass — egress and DNS to model providers and repositories from any source. Registry request logging — ship it and alert on repository creation and on deploys outside known CI paths, because neither the audit trail nor the available webhooks give this.
Almost everything else in this programme is harder against an autonomous adversary. This is the exception, and the evidence is unusually clean.
Why precision is structurally high
A decoy has no legitimate consumer, so any interaction with it is by construction either an error or an intrusion — the same structural argument as the coordination-channel discriminator, and the reason both work without a behavioural baseline. The honest limit: no published false-positive rate exists for any canary vendor, and no recall data exists anywhere. Use the structural argument rather than inventing a number.
Two things to get right before anyone buys anything
| Trap | Detail |
|---|---|
| Do not put a retired product in a plan | One major endpoint deception capability was retired before general availability — onboarding blocked August 2025, decoys removed 31 October 2025. The identity-side honeytoken tag is generally available but covers on-premises directory infrastructure via sensors, so cloud-only accounts are not covered — which is exactly where the non-human identities live. The SIEM honeytoken solution covers a key vault only and is community-supported. |
| Do not use free public canaries | They are statically fingerprintable and are flagged by secret-scanning tooling as canaries without triggering an alert. Paid or self-hosted only, and diversify token types — document, DNS, kubeconfig, MCP configuration. |
What deployed deception actually achieves
- One published deployment reported 83% exploitation at around 30 minutes.
- A vendor detected a real 2025 intrusion within seconds using canaries placed in source-control secrets at both organisation and repository level, running tens of thousands of tokens.
- An MCP-configuration canary detects agents routing around organisational MCP restrictions — directly relevant to our own Tier 1 direct-connection prohibition, whose enforcement point is otherwise unclear.
- A caution worth carrying: in the defensive challenge environments, hardening decoys produced only a “slight advantage” because probing for deception is itself observable — but elsewhere decoys blocked up to 94% of privilege-granting exploits. Not evidence that decoys stopped working.
And respect the published guidance on maturity. National-level trials across 121 organisations and 14 vendors recommend tripwires and honeytokens for everyone, and full interactive honeypots only for mature security operations teams. Given our baseline, tokens are in scope now and interactive honeypots are not — saying so protects the programme from a credibility problem later.
Prevention is the easier half. This is the half that decides whether an intrusion costs two days or two quarters — and the half most organisations have not engineered.
The three tiers
| Tier | Scope | Membership test |
|---|---|---|
| A — autonomous | Machine-identity revocation, and non-critical workload quarantine. | If executing it wrongly breaks a workload rather than a person’s day, and it can be reissued, it qualifies. |
| B — capped | Permitted autonomously up to a blast-radius ceiling, then escalated. | The cap is the control, and it must be expressed in scope rather than in count. |
| C — two-human | Anything that stops production, touches the production environment, or cannot be reversed. | OT sits entirely here and remains deferred to FY27. |
Revocation alone is not a containment strategy
If the adversary holds signing material it can mint valid credentials faster than we withdraw them — which is what happened. So the catalogue must include authority-level actions, not only credential-level ones.
| Mechanism | What it does |
|---|---|
| Issue-time cohort revocation | Invalidate every credential issued in a window rather than chasing individuals — via the token-issue-time condition key. |
| Lease revocation by prefix | Withdraw a whole class of dynamically-issued secrets in one action. |
| Signing-authority taint | The only response to stolen signing material. Without it, revocation loses to an adversary who can issue. |
| Continuous session signals | Push revocation to relying parties rather than waiting for token expiry — final specifications published September 2025. |
| Label-swap quarantine and GitOps prune | Isolate a workload and restore declared configuration without a rebuild. |
Recovering the agent estate itself
Disaster recovery covers data and applications. It rarely covers agent definitions, system prompts, tool registries, MCP configuration, vector-store contents, memory stores or gateway policy. Two consequences: without a known-good configuration you cannot prove an agent is clean, and the fastest route back to service is therefore to redeploy the compromised one.
And containment has a cost that must be priced in. The only serious published evaluation of autonomous containment prices the defender’s own collateral damage into an availability-weighted reward — a restore action carries a negative score. A containment action that halts production is not a successful containment. It is also why the automation decision is taken separately per asset class rather than once for the estate.
Every defence, the domain it serves, the outcomes it moves, and its current state. A capability that moves no measured outcome fails Gartner’s prioritisation test and should not be funded.
| Defence | Domain | Outcomes it moves | Today | |
|---|---|---|---|---|
| A | Where we stand | F | Scope statement for all five — no outcome of its own | Done |
| B | The AI Lab | F | Forensic independence (binary); response-workflow capacity | MIL0 |
| C | Know the ground | D1 | Asset reconciliation; crown-jewel blast radius | MIL1 |
| D | Test & fix at speed | D1 D4 | Controls validated by simulated attack; assets under continuous posture assessment | MIL1 |
| E | Detect & escalate | D2 | Scenarios converted to requirements; techniques with a named telemetry source | MIL0 |
| F | Deceive | D3 | Deception coverage; incidents first surfaced by a disruption control | MIL0 |
| G | Respond & recover | D3 F | Credential classes revocable in ten minutes; known-good agent configuration | MIL0 |
Three framings the evidence settled
| Originally | Reframed to | Because |
|---|---|---|
| “Build a digital twin of the estate” | A cloud-first cyber range as an MVP, plus attack-surface management | The published exemplar used a sanitised natural-language specification and an isolated representative test environment, not telemetry-derived replication — and the team judged a full production-environment twin impractical. Gartner permits validation against a “simulated digital twin” for exactly this narrow purpose. |
| “Simulate the agent swarm” | Remove the assumption that attacks arrive at human speed | No tooling simulates a coordinating collective; offensive multi-agent frameworks are division of labour, not coordination. |
| “Deception slows the adversary down” | Deception as a detection control | The attention-diversion effect is statistically absent in models and trap recognition does not predict behaviour. Selling it as delay would be selling the one benefit the evidence removes. |
The gap the mapping exposes
D2 is covered by one defence, and nothing in the seven addresses deepfake-enabled social engineering — the highest-prevalence real AI attack at 41% of organisations on audio calls and 35% on video. They were designed against an infrastructure intrusion. The SOP-hardening work Gartner specifies — password reset, partner and supplier interactions, financial transactions — closes it in tranche 1 at essentially no cost.
The structural change is not that attacks are cleverer. It is that the interval between a vulnerability becoming known and being weaponised has collapsed from days to minutes, because frontier models can autonomously reverse-engineer a patch into a working exploit.
What speed actually defeats
| Our process | Its assumed timescale | What the adversary now needs |
|---|---|---|
| Patch cycle | Weeks, gated by change windows and downtime | Minutes to produce a working exploit |
| Analyst triage | Hours, business-day weighted | In the reference case the intrusion ran Saturday to Monday and nothing paged |
| Fleet-wide remediation | ~8 hours to stop a fleet, and no automatic rollback exists anywhere in the mainstream patch stack | Less than the escalation window |
| Containment approval | A change ticket | Average breakout 29 minutes; fastest observed 27 seconds |
The number that ends the human-paced-escalation argument
And one piece of good news worth carrying to the board. “the next 12 to 24 months will likely see an increase in the aggregate volume of vulnerabilities… the medium-term outlook is a positive one. Frontier AI models will enable defenders to inspect an unprecedented volume of source code.” The speed problem is transitional, not permanent — which argues for process change now rather than panic buying.
Gartner’s assessment of what AI is actually doing for attackers today: “threat actors leveraging and abusing widely available AI tools to scale and enhance existing tactics, not the invention of net-new attack methods.” The consequence is a volume problem, and volume breaks controls that were tuned for plausible attempt counts.
What scale looked like in the reference case
The detection consequence, and the approach that survives it
Sequence-based detection fails because parallel agents retry and reorder constantly. MITRE is explicit that strict sequencing “can be challenging to implement effectively… adoption of these types of analytics have been limited in practice.” What survives is order-independent convergence counting:
And the approach that does not survive it. Detecting automation by its tempo is the intuitive response to a scale problem, and it fails at 99.8% or worse against a competent adversary — and fails confidently, with mean confidence above 0.993 when wrong. Use tempo to enrich an alert; never to raise one.
Fifty-four attack vectors across ten layers of the estate: agents and models, gateways, identity, cloud and compute, network and egress, data and knowledge, software supply chain, endpoints and collaboration, observability and recovery, and manufacturing. Four have no named owner.
The layer nobody owns, and why it matters most at the enterprise
SEMI E188 explicitly excludes the manufacturing execution system, the material-control system and factory-provided host systems; E187 addresses supplier-provided equipment. Neither standard covers the factory-software layer that actually holds write authority into the tools — and that layer is the one reachable from IT. Gartner describes the same gap from the infrastructure side: “Fragmented IT/OT convergence creates severe risks, as current systems lack unified governance and struggle to support the downtime required for constant patching.”
Two layers our own defences did not cover
| Layer | Why it was missed | Evidence |
|---|---|---|
| Endpoint and collaboration | Our seven defences were designed against an infrastructure intrusion, so the human-facing surface was out of frame. | Deepfake-enabled social engineering at 41% audio / 35% video is the highest-prevalence real AI attack |
| Observability and recovery | The systems that watch and restore the estate are targets because they are the detection capability. | An agent holding backup or DR automation scope can neutralise recovery as a legitimate action |
The reframe that makes breadth fundable. Breadth cannot be closed by buying more controls — there are too many layers and four have no owner. It is closed by enumerating what exists and naming who is accountable, which is why D1 precedes everything and why the board ask for distributed ownership is the programme’s critical path.
Gartner reproduces the 2026 Verizon breach dataset under a heading that is itself the finding: “The Majority of Breaches Still Do Not Exploit Vulnerabilities.”
What this changes in our allocation
- Identity is the highest-leverage surface, not vulnerability management. Credential abuse appears in 39% of breach chains, and machine identity is where our maturity is lowest.
- The human-facing surface earns real investment. Phishing is stable at 16% and deepfake-enabled social engineering is measured at 41% and 35% — and the mitigation is process change, not technology.
- Vulnerability work shifts from patching to exposure. CVSS-led prioritisation drives “efforts to address only critical and high CVSS scores rather than prioritizing those riskiest to the organization.”
And a forward-looking caution on where findings will come from. “By 2028, more than half of threat exposure findings will result from nontechnical vulnerabilities, rather than technical flaws, requiring a fundamental shift in security priorities as these risks surpass traditional IT concerns.”
This is the claim our strategy is most exposed on, so it is stated conservatively and sourced precisely.
What Gartner says
- “Uses of LLMs to assist in malware creation, as part of the malware workflow, or to orchestrate automated attacks are emerging with unclear impact so far.”
- Threat actors integrating models into malware show a “lack of sophistication” and are “more experimental than mature”.
- “today, true AI-powered attacks remain very rare in the real world. Rather, adversaries are using AI to scale traditional phishing, automate tasks, and compensate for skill gaps.”
- AI-augmented attacks are classified in the “unpredictable threats” section of the 2026-2027 ThreatScape, which maps threats by signal quality and threat-actor advantage.
What has nonetheless been observed
| Observation | Status |
|---|---|
| A frontier lab disrupted a state-attributed campaign against ~30 targets in which AI performed 80-90% of the work — “the first documented case of a large-scale cyberattack executed without substantial human intervention.” | Single-vendor self-report, not independently audited. Not sector-specific. |
| MITRE has published the reference intrusion as a formal case study, AML.CS0068, mapping 22 techniques across 37 edges. | Externally adjudicated. This is the strongest evidence available. |
| ATLAS moved from 16 tactics and 84+ techniques at v5.1.0 to v5.4.0 within roughly three months. | The taxonomy is moving fast enough that any coverage baseline must record its ATLAS version. |
And the instruction that governs how we present all of this. “CISOs should help their organization ignore sensational marketing hype around emerging AI threats and focus on actionable responses.” A strategy that overstates this gets discounted the first time a director checks it against the same research.
The most reassuring finding in the corpus, and the one most likely to be misread as a reason to do nothing.
And alongside it: “This situation is not a repeat of Y2K and does not require a transformational response.”
What “foundational controls still work” does not mean
| It does not mean | Because |
|---|---|
| That no investment is needed | “Many capabilities promoted as preemptive exist in current platforms. Combining them, however, has the potential to add resilience.” The work is combination and validation, which is effort rather than licences. |
| That our controls are working | Nothing in our estate has been validated by simulated attack against the agentic technique set. Working and untested are different claims, and outcome D1 exists to close that gap. |
| That the gap will not widen | “AI-driven exploit generation will continuously outpace traditional patching cycles.” Foundational controls hold today; the margin is narrowing. |
The honest framing for the board. Our controls are probably adequate against the attacks of the last three years, and we cannot currently demonstrate that they are adequate against the next three. The programme buys the demonstration — which is a far cheaper ask than a transformation, and a far more defensible one.
Gartner’s definition: “tooling to test for vulnerabilities, assess how they expose your organization to breach activity, and evaluate the security controls you have in place to ascertain their effectiveness. These offerings simulate real-world attacks… against extant security controls.” The stated benefit is visibility into “where and why they fail prior to an actual attack.”
The technology categories Gartner places here
- Adversarial exposure validation (AEV) — the core capability
- Breach and attack simulation (BAS)
- Automated security control assessment
- Autonomous exposure remediation — noting Gartner’s caution that organisations are “extremely rare[ly]” willing to auto-remediate
The five-phase discipline beneath it
| Phase | What it means here |
|---|---|
| Scope | By business impact, not technology silo — “taking into consideration the potential business impact of a compromise rather than primarily focusing on the severity of the threat alone.” |
| Discover | Wider than known-vulnerability scanning. Neither entry vector in the reference case had a published vulnerability identifier, so a scanner-driven programme would have matched nothing. |
| Prioritise | By attack path. CVSS-led prioritisation addresses “only critical and high CVSS scores rather than prioritizing those riskiest to the organization.” |
| Validate | The phase most programmes skip — and the one that would have caught this. |
| Mobilise | Needs resolver teams. “Without widespread business engagement most exposure management functions… are unable to function effectively.” |
Where validation runs
Against real controls, or — in Gartner’s words — “in some cases, a simulated digital twin to reduce potential impact.” That is the narrow, defensible version of a twin, and it matches the cloud-first cyber-range MVP the team chose over a full production-environment twin. The published exemplar for building such an environment used a sanitised natural-language specification, not telemetry replication — the engineers call it a “representative test environment”.
The measurement discipline that must come with it. Score detections on robustness, precision and implementation coverage rather than techniques touched — MITRE: “the goal is not simply to maximize the number of ATT&CK techniques associated with detection content.” And never publish a single coverage percentage without its partial-coverage caveat: the reference framework states its own headline as both “78 of 78 have a native analytic” and “0 direct, 78 partial”.
Gartner: “using data and context to understand likely trends from adversaries… TTPs that are in use and increasing as well as targets.” And the qualification that shapes how we fund it: “On its own, adversary management allows for scenario planning… It is of most value when combined with one of the other pillars to provide actionable enforcement.”
There is no feed to buy for this
No vendor-neutral, standards-body TTP catalogue for autonomous attackers exists beyond MITRE ATLAS and the OWASP Top 10 for Agentic Applications. The US government AI-ISAC remains policy-mandated but pre-decisional with no launch date, so interim AI-threat intelligence should route through an existing sector ISAC. So D2 is a built pipeline, not a procurement.
The pipeline, and what it produces
| Stage | Practice |
|---|---|
| Requirements | Write priority intelligence requirements at a “stable middle” specificity — a named technique, a named source, and the decision the answer supports. An analyst can realistically sustain three to five. |
| Collection | Source named techniques from ATLAS and OWASP’s agentic list. Record the ATLAS version — it moved from 16 tactics and 84+ techniques at v5.1.0 to v5.4.0 in roughly three months. |
| Analysis | Convert to a detection or mitigation requirement with an owner. This is the step that turns D2 from a reading exercise into a control. |
| Development | Sigma-coded, ATT&CK- and ATLAS-tagged detections, through a Requirements Discovery → Triage → Investigation → Development pipeline. |
| Validation | Purple-team emulation — which hands the result to D1. |
The gap that makes this domain urgent rather than academic. MITRE has adjudicated the reference intrusion as a case study and publishes no detection guidance for it — no detection field, no data-source field, one generic mitigation; and 73 of 197 techniques have no mitigation at all, concentrated in the autonomous block. ATT&CK users inherit a data-source layer for free. Here it does not exist, so D2 builds it.
And the vector D2 must cover that our defences did not
Deepfake-enabled social engineering, at 41% of organisations on audio calls and 35% on video. Gartner names the processes: password reset, partner and supplier interactions, and financial transactions. This is a D2 scenario feeding a process control, not a technology purchase — and it is the cheapest item on the roadmap.
Gartner: “accepting that attacks will occur and working to delay or misdirect such attacks… deception and honeypots that provide a clear signal of an attack when triggered, conducting proactive threat hunting… and operationalizing threat intelligence in other controls.” Three stated purposes: slow the attacker, give a clear signal, and feed attacker TTPs back into other controls.
The second half of D3: raising cost through revocation
A stolen credential that dies in ten minutes is worth very little. Machine identities are the reversible asset class — revoking one breaks a workload rather than a person’s day, and it can be reissued — which is what makes pre-authorisation defensible here and nowhere else yet. Mechanisms: issue-time cohort revocation, lease revocation by prefix, signing-authority taint, and continuous session signals.
And Gartner names this as the hard part. “it is extremely rare for organizations to be willing to automatically remediate discovered issues due to the concern over potential disruption… Without a cultural shift, many CISOs will be unable to fully utilize preemptive cybersecurity approaches.” We are asking for that shift, scoped to the reversible asset class only.
Two procurement traps
| Trap | Detail |
|---|---|
| A retired product in the plan | One major endpoint deception capability was retired before general availability — onboarding blocked August 2025, decoys removed 31 October 2025. The identity-side honeytoken is generally available but covers on-premises directory infrastructure via sensors, so cloud-only accounts — where the non-human identities live — are not covered. |
| Free public canaries | Statically fingerprintable, and flagged by secret-scanning tooling as canaries without triggering. Paid or self-hosted only, with diversified token types. |
Gartner: “Successful attacks often take advantage of failures in the configuration of your environments… As the number of vulnerabilities organizations have to manage increases due to the effect of LLM-driven vulnerability discovery… hardening configurations is even more vital than before.” The benefit: it “reduces the attack surface and minimizes the impact of an incident that does compromise a hardened device.”
The technology categories, including one we do not have
CNAPP — cloud security posture management, cloud infrastructure entitlement management, Kubernetes security posture management, and AI security posture management — plus data security posture management, network security posture management, unified endpoint management and endpoint protection.
A procurement note that follows from market structure. The category is consolidating into CNAPP: Wiz into Google, Protect AI into Palo Alto, CalypsoAI into F5, Aim Security into Cato, Robust Intelligence into Cisco. AI-SPM will therefore arrive as a module of something the enterprise may already license — consistent with Gartner’s wider point that “many capabilities promoted as preemptive exist in current platforms.”
The segmentation journey, and why it usually fails
Gartner’s recommendation is “the journey from macrosegmentation to network security microsegmentation to limit lateral movement.” The published failure pattern is organisational, not technical:
| Finding | Consequence for how we phase it |
|---|---|
| Of fourteen organisations attempting even limited segmentation, eleven failed — predominantly for want of an executive champion and application-owner buy-in. | Secure the sponsor and the app owners before the tooling. This is the same distributed-ownership ask as everywhere else. |
| Discovery first: 30 to 90 days of traffic observation before enforcement; pilot in 8-12 weeks; first enterprise segment in 3-6 months. | Enforcing before dependency mapping is the usual cause of outage-driven rollback. FY27, not tranche 1. |
| A Kubernetes NetworkPolicy is accepted and silently does nothing unless the CNI plugin implements enforcement. | Verify enforcement empirically. A policy object is not a control. |
| For legacy process equipment, network-enforced segmentation is the only viable primary control — agents cannot run and VLAN tagging is often unsupported. | The production environment path is network-level, and it sits behind the FY27 OT scoping decision. |
Gartner’s model puts the four domains on a foundation, and is direct about why: “many organizations will require vendor- or third-party-led managed services to realize significant value from these offerings unless they have large and high-level security teams.” And: “With all preemptive cybersecurity offerings, many organizations will also benefit from a managed or supported service.”
The three foundation outcomes
| Outcome | Why it belongs at the base |
|---|---|
% of exposure findings with an accountable resolver outside CDR | Only 36% of organisations have infrastructure teams actively engaged on remediation, and without business engagement exposure management “cannot function effectively.” Gartner escalates this to a board ask. |
| Forensic capability independent of hosted guardrails — binary | Hosted models refused defensive work in the reference incident, measured at 2.72× refusal on security tasks, and “cannot rephrase refused queries or retry.” Cannot be procured mid-incident. |
% of response workflows with defined capacity and a lifecycle | Gartner projects AI agents autonomously managing 25% of incident-response workflows for data security events by 2028. Automating an undefined process automates the wrong one. |
The sequencing risk this layer creates
Gartner warns that preemptive offerings may be “limit[ed] to detecting, rather than fixing issues, thereby increasing, rather than reducing overhead.” If validation and posture assessment stand up before the accountable resolvers exist, the programme manufactures a findings backlog nobody owns. That is the single most likely way this creates work rather than protection — which is why the ownership decision is sequenced first.
And one caution on how far to automate the foundation. “fully automating fixes will eliminate practical learning ground required to develop experienced Level 3 analysts.” The long-run mandate Gartner describes is antifragility — organisations “intentionally us[ing] systemic shocks and controlled risk exposure to emerge operationally stronger” — which is an argument for exercises, not autopilot.
“By 2030, preemptive cybersecurity solutions will account for 50% of IT security spending, up from less than 5% in 2024, and replace traditional ‘stand-alone’ detection and response solutions as the preferred approach to defend against cyberthreats.”
Three corroborating signals
| Signal | Figure |
|---|---|
| Current spend split | Of roughly $244.2B in 2026 security spending, about $49B goes to AI-amplified security and only $2.8B to securing AI itself — roughly a seventeen-fold imbalance. |
| Control gap | 13% of organisations reported a breach of an AI model or application, and 97% of those lacked AI access controls. Unsanctioned AI added about $670,000 to average breach cost. |
| Adoption vs scale | 88% of organisations report AI adoption but fewer than 10% have fully scaled it; 42% abandoned most AI initiatives in 2025, up from 17%. |
And the market-structure consequence for procurement. The AI security category is consolidating into cloud-native application protection platforms — Wiz into Google, Protect AI into Palo Alto, CalypsoAI into F5, Aim Security into Cato, Robust Intelligence into Cisco. Buying a standalone AI security product in FY27 risks buying something that becomes a module of an existing licence.
“Many capabilities promoted as preemptive exist in current platforms. Combining them, however, has the potential to add resilience.” And the instruction that follows: CISOs “must ensure that they do not already have a similar capability being delivered as part of their existing product and service portfolio.”
Gartner’s own mapping of domains to existing technology categories
| Domain | Categories that may already be licensed |
|---|---|
| D1 | Adversarial exposure validation; breach and attack simulation; automated security control assessment; autonomous exposure remediation |
| D2 | Cyberthreat intelligence; unified cyber risk intelligence; CTI management platforms |
| D3 | Deception; firewalls and NDR with threat-intelligence integration; threat hunting; detection engineering |
| D4 | CNAPP — CSPM, CIEM, KSPM, AI-SPM; data security posture management; network security posture management; unified endpoint management; endpoint protection |
Two things we already hold that are under-used
- The exposure-management platform. Critical-asset definition generates attack paths, and first-class connectors exist for the vulnerability scanner and the service-management CMDB. The gating input is a business workshop to say what matters, not an engineering build.
- The gateway. It already enforces directory authentication, scope authorisation, quotas, DLP and prompt-injection inspection, with central logging. What it does not yet do is mint a delegation identity — which is the one thing no product will supply and the fix for an otherwise unsolvable attribution problem.
Gartner’s shorthand for the same model: “Deny intruders access to your global attack surface grid through advanced obfuscation techniques. Deceive bad actors through automated cyber deception and moving target defense. Disrupt attacks via predictive threat intelligence and automated exposure management.”
How the verbs map to our domains
| Verb | Domain | What we actually do |
|---|---|---|
| Deny | D4 posture and policy | Reduce what is reachable: runtime enforcement of declared agent scope, deny-by-default egress for the untrusted-input tier, and the macro-to-microsegmentation journey. |
| Deceive | D3 adversary disruption | Deception tokens where an agent reaches them first — and sold on signal, because the delay benefit does not survive against agents. |
| Disrupt | D1 + D2 | Predictive intelligence converted into detection requirements, and automated exposure validation. |
One term worth knowing if it comes up. Gartner calls the thing being defended the “global attack surface grid” — which is the same observation as our ten-layer landscape, and the same one behind “organizations are creating attack surfaces faster than technologies can protect them.”
“Current detection and response and application security methods aren’t sufficient to keep up with the speed, sophistication and scope of emerging AI-enabled threats.” The preemptive model “adds focus beyond prevention, detection, and response and seeks to improve foresight capabilities to limit an attack’s impact early, optimally preventing initial access.”
The evidence from our own estate
| Finding | Source |
|---|---|
| Agent-runtime telemetry touches 69 of 78 techniques; gateway telemetry touches 16 — and all 78 mappings are partial. Our confirmed telemetry is the gateway. | Independent coverage matrix |
| Detection and correlation worked in the reference case. Criticality scoring failed and nobody was paged, across a weekend. | The victim’s own post-mortem |
| There is no published detection guidance for the autonomous technique block, and 73 of 197 techniques have no mitigation at all. | MITRE ATLAS |
And the one detection investment that still pays disproportionately. Escalation, not detection. Measuring mean time to escalate separately from mean time to detect costs nothing — it is a decision about how existing numbers are cut — and it is the only way the reference failure becomes visible. Related: “Nothing on any screen tells you that a finding failed to appear.”
“AI-washing, both as a threat and as a capability, is prevalent in cybersecurity product marketing today.” Gartner is blunt about the cause: “Preemptive cybersecurity is being applied as a label to both new and existing technologies by vendors… there is AI-washing both in creating demand for and in labeling these technologies as preemptive.”
Three specific claims to interrogate
| Vendor claim | What to ask |
|---|---|
| “AI-powered attack detection” | Ask for the detection content and its validation results. True AI-powered attacks “remain very rare in the real world”, so ask what corpus the detection was tuned on. |
| “AI security posture management” | Ask specifically whether it detects an agent acting outside its declared scope at runtime. Classic AI-SPM is a static configuration discipline and does not. |
| “Autonomous remediation” | Ask what it will do without approval, and read Gartner’s caution: organisations are “extremely rare[ly]” willing to auto-remediate, which can limit these offerings to detecting rather than fixing — increasing overhead. |
And the sceptical habit worth keeping about our own claims too. The same discipline applies inward. We do not report coverage we have not validated, we never publish a single coverage percentage without its partial-coverage caveat, and we state limits plainly — including that no published false-positive rate exists for any canary vendor and no recall data exists anywhere.
“The immediate danger is the CISO attempting to own the solution in a silo. If the CISO promises to simply ‘patch faster’ without the backing of the CIO, they will fail. In most organizations this is neither attainable nor sustainable.”
Why it is not attainable
| Constraint | Evidence |
|---|---|
| Severity-led prioritisation targets the wrong work | CVE/CVSS assessment produces “efforts to address only critical and high CVSS scores rather than prioritizing those riskiest to the organization.” |
| The federal mandate itself moved away from it | The CVSS mandate was revoked in June 2026 for a four-variable stakeholder model with 3/14/60/defer timelines — and only about 1% of instances fell in the 3-day band while roughly 60% were deferred. |
| The tooling cannot support it | No automatic rollback exists anywhere in the mainstream patch stack, and stopping a fleet takes around eight hours — longer than the entire escalation window in the reference incident. |
| The production environment cannot absorb the downtime | “current systems lack unified governance and struggle to support the downtime required for constant patching.” |
| And the gap will widen regardless | “AI-driven exploit generation will continuously outpace traditional patching cycles.” |
Plus the ownership consequence, which is the real fix. “cybersecurity teams can only guide vulnerability prioritization. IT operations, product teams and business system owners must be held accountable by the board to fix exposures in their own systems.” Only 36% of organisations have infrastructure teams actively engaged on remediation today.
The programme asks for autonomous action in exactly one place — machine-identity revocation — and stops there deliberately. Three reasons, all sourced.
1. It erodes the capability it depends on
“fully automating fixes will eliminate practical learning ground required to develop experienced Level 3 analysts.” The senior analysts a programme like this needs in three years are produced by the work it would be most tempting to automate now.
2. The defender’s own disruption is a real cost
The only serious published evaluation of autonomous containment prices collateral damage into an availability-weighted reward — a restore action carries a negative score. A containment action that halts production is not a successful containment, which is why the automation decision is taken separately per asset class rather than once for the estate.
3. Organisations do not actually do it
“it is extremely rare for organizations to be willing to automatically remediate discovered issues due to the concern over potential disruption, and this may limit preemptive cybersecurity offerings to detecting, rather than fixing issues, thereby increasing, rather than reducing overhead.”
Where Gartner expects this to go anyway. “By 2028, cybersecurity AI agents will autonomously manage 25% of incident response workflows for data security events.” Which is an argument for defining the workflow and its capacity now, while the numbers are still manual and therefore honest — that is foundation outcome F.3. The long-run mandate is antifragility: “intentionally use systemic shocks and controlled risk exposure to emerge operationally stronger.”
“CISOs Must Focus on the Outcomes of Preemptive Cybersecurity”, 5 August 2026, nine pages plus a technology appendix. It supplies the four-domain structure, each domain’s definition and core benefit, the foundation layer, and a mapping of domains to technology categories.
Its three stated challenges, all of which we inherit
- Cultural shift required, because “predicting cyberthreat activity accurately is largely aspirational” and organisations are “extremely rare[ly]” willing to auto-remediate.
- AI-washing is prevalent in both the demand creation and the product labelling.
- Much of it already exists in current platforms — “Combining them, however, has the potential to add resilience.”
What we added to it
| Gartner supplies | We added |
|---|---|
| Four domain definitions and core benefits | The enterprise-specific content for each, and the outcome-driven metrics per domain |
| A foundation layer described as managed services | Three named foundation outcomes, including the honest gaps: detection capacity, forensic independence, accountable resolvers |
| A technology mapping table | An assessment of which of those we already hold, and the finding that AI-SPM does not deliver runtime scope enforcement |
| Domain-level benefit statements | A maturity assessment per domain, and a readiness position on an investment scale |
One permission in this document worth knowing about. Gartner explicitly allows exposure validation to run “against actual controls, or in some cases, a simulated digital twin to reduce potential impact.” That is the narrow use the team’s cloud-first cyber-range MVP occupies — so the twin was not rejected outright, it was scoped to the one purpose the research supports.
“Outcome-Driven Metrics for the Digital Era”, 11 July 2025, twenty-one pages. It supplies above-and-below-the-line structure, the five-step derivation method, and — in Note 1 — seven characteristics that decide whether a metric is worth reporting at all.
The definition of an ODM, which is also a filter
The five steps, as applied
| Step | Gartner | Ours |
|---|---|---|
| 1 | Three to five most obvious business processes and their technology stacks | The three agreed business outcomes, plus the crown jewels named by the business |
| 2 | Business outcomes and business ODMs | Continuity, IP and recipe protection, adoption speed — above the line |
| 3 | Technology risks and dependencies — “What breaks in the business process if the technology breaks?” | The two-control-plane assessment and the confirmed baseline |
| 4 | Define technology ODMs, which behave as leading indicators | Fifteen metrics in “% of” form, three per domain plus three foundation |
| 5 | Assess readiness as risk to business enablement, on a scale from no investment to leading edge | The readiness scale, with current position marked per domain |
The other five characteristics, used as filters
- Metric value — “its ability to influence decision making… typically priorities and investments.” Anything that could not change one was cut.
- Leading indicators — must “expose a problem before it leads to material loss.” This is why validation coverage beats incident counts.
- Direct line of sight — “If the technology metric changes, it indicates a change in the business outcome.” Every metric on the page states its line of sight.
- Metric changes drive action — green to yellow to red must trigger a change of priority or investment.
- Discrete audiences — “The CFO needs different metrics from those required by the head of a business unit or a board of directors.”
And the honest cost the document names. “measuring some of these elements may require instrumenting parts of the infrastructure to gather new types of data… Gartner believes the visibility and power to report benefits and to guide priorities… will make the initial investment worthwhile.” Two of our fifteen need exactly that, and they are the two with the largest payoff: reconciliation, and runtime enforcement of declared scope.
Step 5 of the method: “Create the business enablement risk scale with, at one end, no investment or technology stack to enable the business outcome; at the other, investment in leading-edge technologies that would elevate the current technology stack to drive the most positive outcome for the business.”
Readiness, defined
“technology that operates the way it is designed, fully supports a business process, meets compliance requirements and is not unreasonably at risk of failure” — and separately, “Readiness describes the state of investments and capabilities to manage [the risk].” So a low position is a statement about investment, not about competence.
Reading our five positions
| Domain | Position | Why there |
|---|---|---|
| D2 Adversary management | Highest | Threat modelling and scenario work are genuinely mature. But Gartner notes the domain “is of most value when combined with one of the other pillars” — so our strongest domain currently returns the least. |
| F Foundation | Middle | Governance is strong — a dedicated board-level Security Committee already owns this remit — while capacity, forensic independence and resolver accountability are not established. |
| D1 Exposure management | Low | A self-declared registry exists; no controls have been validated by simulated attack against the agentic technique set. |
| D4 Posture and policy | Low | Cloud posture management is partial, AI-SPM is absent, segmentation is macro-level, and runtime enforcement of declared scope is zero. |
| D3 Adversary disruption | Lowest | Nothing placed, no revocation authority defined, no target time, no rehearsal. And the cheapest of the five to move. |
The shape of the answer is the recommendation. D3 sits lowest and costs least to improve — deception needs no inventory, no gateway change and no new telemetry, and revocation authority is a decision rather than a purchase. That is why tranche 1 invests there rather than in the domain where we are already strongest.
All three measure the same underlying question in different registers: do we know what we have, and have we tested whether it holds?
| Metric | Investment | How it is computed, in outline |
|---|---|---|
D1.1 % of discovered AI assets and machine identities that are fully reconciled | Asset discovery tooling plus the exposure platform’s inventory connectors | Discovered population as the denominator, reconciled-and-owned as the numerator. Not completeness — that needs a true total we do not have. Report alongside median declare-lag. |
D1.2 % of priority adversary techniques with at least one defence proven by executing the attack in the last 90 days | Adversarial exposure validation or breach-and-attack simulation, plus the cyber-range MVP | Externally-scoped technique set as the denominator; techniques with a detection that passed an adversarial test as the numerator. Never published without its partial-coverage caveat. |
D1.3 % of crown-jewel data stores with a current, machine-generated blast-radius map | Attack-path analysis in the exposure platform | Crown jewels are named: critical IP and recipe repositories. Numerator is those with destinations-reachable-per-credential computed and current. |
The outline above is not the specification. Each metric’s unit of count, denominator, pass test and fail modes are in its own panel, reachable from the Definition and pass test link beside it on the programme set.
Intelligence only counts once it becomes enforcement. Gartner: adversary management “is of most value when combined with one of the other pillars to provide actionable enforcement.”
| Metric | Investment | How it is computed, in outline |
|---|---|---|
D2.1 % of named high-risk procedures that pass a live social-engineering test with all four verification controls in place | Process redesign — no technology purchase | Denominator is the named high-risk processes: password reset and MFA reset, partner and supplier interactions, financial transactions. Numerator is those with out-of-band callback to a pre-registered channel and dual authorisation above a threshold. |
D2.2 % of priority threat scenarios converted into an implementable requirement with a named individual owner | Threat-intelligence analyst time and the detection-engineering pipeline | Priority intelligence requirements raised as denominator; those with a detection or mitigation requirement and an accountable owner as numerator. An analyst sustains three to five PIRs. |
D2.3 % of priority techniques with a named log source that is collected today and carries the required fields | Minimum-telemetry-requirements analysis | Per technique, whether a log source is named and shipping. The Dependencies column of that artefact is where the brokered asks land. |
The outline above is not the specification. Each metric’s unit of count, denominator, pass test and fail modes are in its own panel, reachable from the Definition and pass test link beside it on the programme set.
Why the deepfake metric leads this domain. It addresses the highest-prevalence real AI attack — 41% of organisations on audio calls, 35% on video, and one in five biometric fraud attempts — and it is the cheapest item on the roadmap because it is process change. Detection tooling is not the answer here: commercial deepfake detectors lose 45-50% of their accuracy on realistic in-the-wild content, liveness certification structurally excludes the injection attacks actually used, and untrained humans catch deepfakes almost never.
Three metrics, matching Gartner’s three stated purposes for the domain minus the one that does not survive against agents.
| Metric | Investment | How it is computed, in outline |
|---|---|---|
D3.1 % of crown-jewel environments with at least one live, monitored, tested decoy on the intruder’s path | Paid or self-hosted canary tokens — not the free public service | Crown-jewel environments as denominator; those with diversified token types placed as numerator. Token types: document, DNS, kubeconfig, MCP configuration. |
D3.2 % of machine-identity credential classes revocable cohort-wide in under ten minutes, measured at the resource | Identity lifecycle work, plus cohort-revocation mechanisms | Credential classes enumerated as denominator; those with a named authority, a rehearsed procedure and a measured time as numerator. Currently undefined, so the first measurement is also the first improvement. |
D3.3 % of confirmed incidents whose first signal, in the post-incident timeline, came from a disruption control | No incremental investment — an attribution field on incident records | Confirmed incidents as denominator; those whose first signal came from deception or threat hunting as numerator. This is the metric that tests whether D3 is working rather than merely deployed. |
The outline above is not the specification. Each metric’s unit of count, denominator, pass test and fail modes are in its own panel, reachable from the Definition and pass test link beside it on the programme set.
Reduce what is reachable, and make declared scope mean something at runtime.
| Metric | Investment | How it is computed, in outline |
|---|---|---|
D4.1 % of production agents whose registry-declared scope is refused at runtime when exceeded | Gateway and registry integration — a build, jointly with AI Enablement | Production agents as denominator; those whose registry entry is bound to an enforcement point as numerator. Currently zero. Must be written as an explicit requirement and tested — AI-SPM is a static discipline and does not deliver it. |
D4.2 % of critical assets assessed against a named standard at least weekly, with drift raised as an owned finding | CNAPP modules already partly licensed, plus AI-SPM | Critical assets from the inventory as denominator. For OT, note that assessment is achievable where response authority is not — the FY27 deferral is about action, not visibility. |
D4.3 % of untrusted-input workloads with egress denied by default and the denial proven by test | Network policy work, scoped to the parsing tier only | Workloads that parse externally-sourced input as denominator. Scoping it narrowly is what makes the cost bounded — and verification must be empirical, since a Kubernetes NetworkPolicy silently does nothing without an enforcing CNI plugin. |
The outline above is not the specification. Each metric’s unit of count, denominator, pass test and fail modes are in its own panel, reachable from the Definition and pass test link beside it on the programme set.
And the segmentation journey these sit inside. Gartner’s recommended path is macro- to microsegmentation to limit lateral movement. Phase it discovery-first: 30 to 90 days of observation before enforcement, a pilot in 8-12 weeks, a first enterprise segment in 3-6 months. And expect the failure mode to be organisational — eleven of fourteen organisations attempting segmentation failed, mostly for want of an executive champion and application-owner buy-in.
Gartner rests the four domains on managed services and capabilities, noting most organisations need partner support “unless they have large and high-level security teams.”
| Metric | Investment | How it is computed, in outline |
|---|---|---|
F.1 % of deduplicated exposure findings with a named individual outside CDR who has accepted or rejected them on the clock | Governance — a board decision, not a purchase | Findings raised as denominator; those with a named accountable owner outside CDR as numerator. Only 36% of organisations have infrastructure teams actively engaged on remediation. |
| F.2 Forensic analysis capability that does not depend on a hosted model’s guardrails — reported as met, partly met or not met | A vetting exercise plus a small self-hosted deployment | Yes or no. Currently no. Hosted models refuse defensive work at 2.72× the neutral rate and “cannot rephrase refused queries or retry.” Every evidence-touching inference must record model hash, engine version and decode parameters. |
F.3 % of named response workflows with a stated concurrency limit, overflow behaviour and an exercise in the last 12 months | Detection-engineering process definition | Workflows as denominator; those with stated capacity and a requirement-to-validation lifecycle as numerator. Gartner projects AI agents autonomously managing 25% of IR workflows for data security events by 2028 — define it before automating it. |
The outline above is not the specification. Each metric’s unit of count, denominator, pass test and fail modes are in its own panel, reachable from the Definition and pass test link beside it on the programme set.
Four objectives, all owned outside CDR. Roughly 20% of the relevant controls sit with us, so writing these as FTD deliverables would produce a plan that reports as late against work the programme was never able to do.
| Objective | Counterparty | What FTD contributes |
|---|---|---|
| Reduce standing authority — credential lifetimes, scope, no ambient inheritance | Identity | Blast-radius analysis showing which credentials to shorten first |
| Author the isolation standard, then segment | Platform engineering | The standard, and the attack paths it must interrupt. Phase discovery-first: 30-90 days observation before enforcement. |
| Bind the agent registry to runtime enforcement | AI Enablement + Security Architecture | The detection that proves the binding holds. No posture-management product supplies this. |
| Declare the untrusted-input tier; deny-by-default egress for it alone | Platform + network | Scoping the tier so cost lands only where justified |
One design consequence that is not a control choice
“if an LLM is supplied with untrusted input, it will produce arbitrary output.” Indirect prompt injection is a property of the technology, not a defect awaiting a patch. That makes the untrusted-input tier an isolation requirement rather than a filtering one — cheaper and more durable than any inspection tier, and consistent with Gartner’s related position that the future of AI security is in securing agent actions rather than prompts.
And the closing argument that has landed best with this audience. Five of the six capabilities this programme needs already exist at the enterprise for human adversaries. Asset inventory, risk analysis, segmentation, identity hygiene and egress instrumentation all exist in some form. The work is extending them to an actor that tests thousands of paths in parallel. Only pre-authorised response is genuinely new — and it is a governance decision rather than a purchase.
The centre of gravity, for a specific and checkable reason: MITRE has adjudicated the reference intrusion as a formal case study and publishes no detection guidance for it — no detection field, no data-source field, and one generic mitigation.
The sequencing that is not optional
The four things to build, in order
| Build | Why it comes when it does |
|---|---|
| The coverage baseline | A measurement, not a build — one analyst, no new tooling, using MITRE’s published five-column schema. Its Dependencies column is where the brokered asks land. |
| The inventory to join to | Discovery first, attestation last. The only published methodology puts developer attestation fourth of four. |
| Delegation identity at the gateway | Because attribution is formally non-identifiable from logs — trace-based grouping recovers 4-6% of a delegation’s events, and sub-agent fanout above two is the breaking point. Carrier is specified: W3C baggage in MCP params._meta. |
| Declared-versus-observed detection | Needs no baseline. Eleven published queries exist and none joins the inventory to runtime activity. |
The 30-day backlog, from telemetry already held or one switch away. Telemetry tampering (model-invocation logging deleted, guardrails deleted, registry token and user creation) — tiny volume, near-zero false positives, and it defends the detection stack itself. Autonomy-boundary removal — a published detection exists for Linux permission overrides; Windows and macOS equivalents are published nowhere and must be written locally. Governed-plane bypass via egress and DNS to model providers. Registry request logging, because neither the audit trail nor the available webhooks give it.
The enterprise has nothing defined for machine-identity revocation — no named authority, no target time, no rehearsal. Combined with the fact that it costs nothing to fix, that makes this the cheapest material improvement available to the programme.
The three tiers
| Tier | Scope | Membership test |
|---|---|---|
| A — autonomous | Machine-identity revocation; non-critical workload quarantine | If executing it wrongly breaks a workload rather than a person’s day, and it can be reissued, it qualifies |
| B — capped | Permitted autonomously to a blast-radius ceiling, then escalated | The cap is the control, expressed in scope rather than count |
| C — two-human | Anything that stops production, touches the production environment, or cannot be reversed | OT sits entirely here and remains deferred to FY27 |
Why the automation ask is narrow, and what makes it safe
Gartner: organisations are “extremely rare[ly]” willing to auto-remediate, and “without a cultural shift, many CISOs will be unable to fully utilize preemptive cybersecurity approaches.” We are asking for that shift, scoped to the reversible asset class only — and the safeguard is structural: the Tier A list is changeable only by the Security Committee, so the scope of autonomy cannot be widened by those exercising it.
And containment has a cost that must be priced in. The only serious published evaluation of autonomous containment prices the defender’s own collateral damage into an availability-weighted reward — a restore action carries a negative score. A containment action that halts production is not a successful containment, which is why the automation decision is taken separately per asset class.
Four objectives. One of them cannot be acquired mid-incident, which is why it sits in tranche 1.
| Objective | The finding behind it |
|---|---|
| A forensic capability that still works mid-incident | Hosted models refused the disaster-recovery work in the reference case. Measured at 2.72× refusal on defensive tasks, 43.8% on system hardening — and stating you are authorised makes it worse. |
| Known-good configuration for the agent estate | DR covers data and applications, rarely agent definitions, prompts, tool registries, MCP configuration, vector stores or gateway policy. Without it you cannot prove an agent is clean, so the fastest route back to service is to redeploy the compromised one. |
| Pre-authorised reversible recovery | Average adversary breakout is 29 minutes, fastest observed 27 seconds. Anything needing a change window will not execute in time. |
| Continuous validation as the assurance loop | Closes back into D1. It is how every claim in this strategy stops being a claim. |
The four arguments for our own capability, none of which is a capability claim
- Availability — the tool must work at 3am on the worst day.
- Confidentiality — the classifier that refuses is the same system that flags the session for retention and human review, so the payload you most need help with is retained longest. 69% of organisations are concerned about control over AI models.
- Legal and export exposure — the safe harbour covers encrypted transit, but inference decrypts and processes, and no regulator has ruled on whether that is a deemed export.
- Forensic defensibility — standards require repeatability, and hosted models update and reroute silently.
One design requirement that is free now and impossible to retrofit. Every inference touching evidence must record model hash, engine version and decode parameters into the case record. Retrofitting this invalidates the case work already done.
The seven defences were designed against a real agentic intrusion, before the Gartner frame was adopted. Mapping them onto the four domains is therefore a genuine test of whether the technical work serves the strategy — and it mostly does.
| Domain | Coverage | By which defences |
|---|---|---|
| D1 Exposure management | Well covered | C Know the ground (inventory, attack paths, blast radius) and D Test & fix at speed (CTEM, validation) |
| D2 Adversary management | Thin | E Detect & escalate only — and nothing at all for deepfake-enabled social engineering |
| D3 Adversary disruption | Well covered | F Deceive and G Respond & recover (revocation as a cost-raiser) |
| D4 Posture and policy | Partial | Within D — but the runtime enforcement binding is a gap, and no product supplies it |
| F Foundation | Covered | A Where we stand (the assessment) and B The AI Lab (forensic independence) |
What closing that gap actually involves
Not detection tooling. Commercial deepfake detectors lose 45-50% of their accuracy on realistic in-the-wild content, liveness certification structurally excludes the injection attacks actually used against identity verification, provenance metadata is destroyed by any screenshot or re-encode, and untrained humans catch deepfakes almost never. What works is procedural: callback to a pre-registered channel rather than one the caller supplies, a passphrase that never travels over the channel being verified, and dual authorisation above a value threshold.
Two incidents that frame the ask. Arup lost a confirmed US$25.6M after an employee joined a video call on which every other participant was synthetic. And MGM Resorts was compromised in roughly ten minutes through a help-desk MFA reset with an impact exceeding $100M — using no synthetic media at all. The second is the more important one for us: the process weakness is exploitable without any AI, and AI simply makes it cheaper to attempt at scale.
Three measurements and one piece of good news. All four are chosen because they cannot be flattered.
| Number | What it is | Why it is trustworthy |
|---|---|---|
| 16 of 78 | Techniques touched by the telemetry modality we have — gateway. The modality we lack, agent-runtime, touches 69. | Generated by an independent open-source coverage matrix, not by us. And all 78 mappings are partial — “0 direct, 78 partial”. |
| 0 | Controls validated by simulated attack against the agentic technique set. | Binary. D1’s core measure, and Gartner’s stated benefit for the domain is visibility into where controls fail before an attack. |
| 0% | Production agents whose declared scope is enforced at runtime. | Nothing binds the registry to enforcement, and posture-management tooling does not supply it. |
| 1 | A dedicated board-level Security Committee already owning this remit. | From the enterprise’s own filed Item 1C disclosure. Most peers route this through an audit or technology-risk committee. |
And the discipline that must accompany the first number. Never publish a single coverage percentage without its partial-coverage caveat. The framework we borrow from states its own headline twice and honestly — every technique has an analytic, and every mapping is partial. Both are true. If any technique on our dashboard ever shows full coverage, the dashboard is wrong.
Seven actions. Four are decisions or process changes rather than purchases — which follows from Gartner’s observation that many preemptive capabilities already exist in current platforms.
| Action | Why now | |
|---|---|---|
| D3 | Place deception in crown-jewel environments | Needs no inventory, no gateway change, no new telemetry. Highest-evidence control available; use paid or self-hosted tokens, never the fingerprintable free service. |
| D3 | Define machine-identity revocation authority; time one revocation | A decision, not a build. Currently undefined. Moves D3 off zero without procurement. |
| D2 | Harden SOPs for password reset, supplier and financial processes | The highest-prevalence real vector, and the mitigation is procedural: pre-registered callback, passphrase off-channel, dual authorisation. |
| D1 | Publish the control-coverage baseline | One analyst, MITRE’s published schema, no new tooling. |
| D1 | Verify what the gateway AI control actually logs and alerts on | Our only existing AI-specific control, and its behaviour is unverified. May change tranche 2 scope. |
| F | Stand up forensic capability independent of hosted models | Cannot be procured mid-incident. No dependency on inventory, gateway or SIEM. |
| D2 | Inventory approved and rogue client-side GenAI tools via endpoint and network controls | Gartner’s explicit action, using incumbent EPP, EDR and security service edge. |
Four outcomes move in these thirty days. Deception coverage, revocation time, deepfake SOP hardening, and forensic independence. Three of the four are currently at zero and one is binary-no — which is why they are the fastest and least disputable wins available.
Seven actions, and this is where instrumentation spend starts. Gartner is candid that some measurement “may require instrumenting parts of the infrastructure to gather new types of data.”
| Action | Depends on | |
|---|---|---|
| D1 | Pull automated IT, OT and cloud asset inventories from the exposure platform | Business workshop to define critical assets — without it the attack-path view may be empty |
| D1 | First adversarial exposure validation run against priority controls | The coverage baseline from tranche 1 |
| D2 | Convert priority scenarios into detection requirements with owners | Analyst capacity — three to five sustainable PIRs |
| D4 | Bind declared agent scope to runtime enforcement | Joint delivery with AI Enablement. Not supplied by any posture product. |
| D4 | Stand up AI security posture management | Likely a module of an existing CNAPP licence given market consolidation |
| F | Establish detection-engineering capacity and lifecycle | CSOC agreement. Define it before automating it. |
| D3 | Ratify the tiered containment catalogue | Security Committee decision |
And the risk that makes ownership sequencing non-negotiable. Gartner warns preemptive tooling may be “limit[ed] to detecting, rather than fixing issues, thereby increasing, rather than reducing overhead.” Stand up validation and posture assessment before accountable resolvers exist and tranche 2 manufactures a findings backlog nobody owns. Decision 02 must land before this tranche starts.
Seven themes. The test of this tranche is whether the programme stops being a programme — coverage measurement becoming a standing CSOC function rather than a project activity.
| Commitment | Note | |
|---|---|---|
| D1 | Continuous validation as a standing capability | The assurance loop Gartner describes as the core of the domain |
| D4 | Begin the macro- to microsegmentation journey | Discovery-first: 30-90 days observation, pilot in 8-12 weeks, first segment in 3-6 months. Expect organisational failure modes — 11 of 14 attempts failed for want of a champion and app-owner buy-in |
| D3 | Cohort revocation; pre-authorised reversible actions live | Signing-authority taint and issue-time cohort invalidation |
| F | AI Lab: self-hosted defensive and forensic capability | Four arguments, none a capability claim |
| D4 | IT/OT governance framework; factory-software owner named | Gartner describes the gap directly; network-enforced segmentation is the only viable primary control for legacy process equipment |
| D2 | Composite AI patterns for high-risk use cases | Gartner: “combine AI models with deterministic systems to ensure reliable outputs” |
| F | Managed or partner service where in-house depth is absent | Gartner: most organisations need this “unless they have large and high-level security teams” |
And the long-run mandate this is heading toward. Gartner’s 2030 framing is antifragility: “Simply returning to the status quo after an incident will no longer be viable… organizations will intentionally use systemic shocks and controlled risk exposure to emerge operationally stronger.” That is an argument for standing exercise capability, which is what tranche 3 buys.
Ordered by what they unblock rather than by size. Decisions 01 and 02 cost nothing and gate almost everything else.
| Decision | What it unblocks | Cost | |
|---|---|---|---|
| 01 | Reset the risk-appetite question — from days-to-patch to “what is our tolerance to be exploited by a known vulnerability?” | Makes exposure prioritisation defensible, and stops the programme being measured on a target it cannot hit | None |
| 02 | Name accountable owners outside CDR — “IT operations, product teams and business system owners must be held accountable by the board.” | Eight of twenty objectives. Without it, tranche 2 manufactures findings nobody owns | None |
| 03 | Adopt the four domains and the outcome set | Gives the programme a definition of success that is not activity | None |
| 04 | Grant reversible containment authority — machine-identity revocation without prior change approval | D3 moves off MIL0. Gartner names this the required cultural shift | None; needs the Tier A safeguard |
| 05 | Elevate technical debt to a material business risk | Funds the cleanup that reduces exposure faster than patching can | FY27 planning |
| 06 | Fund tranche 1; gate tranche 2 on its findings | Demonstrated delivery before a substantial ask | Existing headcount |
And the message to carry into the room alongside them. Gartner’s guidance for exactly this briefing: temper the fear, uncertainty and doubt, and use the attention as a decision point “to reset expectations, funding, and accountability.” The medium-term outlook is genuinely positive — frontier models “will enable defenders to inspect an unprecedented volume of source code” — and saying so buys more credibility than alarm would.
The document carries over 400 citation markers against a registry of 173 entries across 21 source classes. Every marker opens the verbatim quotation and its locator. Nothing rests on an unattributed assertion.
The evidence base, by layer
| Layer | What it is |
|---|---|
| Gartner corpus | Twelve reports, 334 pages, supplied by the programme and read in full. G00859378 supplies the architecture and G00799085 the measurement system; the other ten evidence and refine them. |
| the enterprise’s own filings | Form 10-K risk factors and the Item 1C cybersecurity disclosure — including the board Security Committee and the company’s own statement that AI “may… lead to new and/or more sophisticated methods of attack.” |
| Government and court record | The DOJ trade-secret case, the acquittal that bounds it, the ODNI threat assessment, and SEC disclosure requirements. |
| Primary technical research | Twenty-eight research dossiers written to disk as the work proceeded, covering detection, inventory, deception, recovery, the AI lab, exposure management, peer benchmarking and defence depth. |
One methodological limitation, disclosed rather than discovered. No board-credible AI-specific security maturity model exists yet. The assessment therefore applies a general maturity model (C2M2 levels) to AI-specific domains. That is defensible — Gartner notes the four domains are largely delivered by technologies that already exist — but it is a limitation, and better stated by us than found by a reviewer.
The most prevalent AI-augmented attack in the published evidence, and the one our defences did not cover. 41% of surveyed organisations have experienced a deepfake and social-engineering attack on an audio call to an employee, 35% on a video call, and deepfakes now account for one in five biometric fraud attempts.
Two incidents that frame the problem
| Incident | What happened | What it proves |
|---|---|---|
| Arup, Hong Kong, Feb 2024 | A confirmed US$25.6M loss after an employee joined a video conference on which every other participant was synthetic, including the CFO. | Real-time video deepfakes are operationally viable against a competent finance function. |
| MGM Resorts, Sept 2023 | Initial access via socially engineering the IT help desk into an MFA reset. Compromise in roughly ten minutes; disclosed impact above $100M. | No synthetic media was required at all. The process weakness is exploitable with a phone call — AI only makes it cheaper to attempt at scale. |
Why detection tooling is not the answer
| Approach | Documented limit |
|---|---|
| Commercial deepfake detection | Models lose roughly 45-50% of their AUC on realistic in-the-wild 2024 content versus laboratory benchmarks. |
| Liveness / presentation-attack detection | The certification standard structurally excludes injection attacks — feeding synthetic video directly into the data path, which is the technique actually used. A product can hold valid certification and remain fully exposed. |
| Content provenance (C2PA) | Metadata is destroyed by any screenshot or re-encode, and absence of a credential is not evidence of fakery. Useful for content you publish, not content you receive. |
| Training people to spot fakes | Untrained participants identified deepfakes in approximately 0.1% of trials. (Vendor-funded study — treat the precise figure with caution; the direction is not in doubt.) |
What actually works — and it is procedural
The controls that hold do not depend on detecting a fake:
- Callback to a pre-registered channel — never a number or address supplied by the caller. This single control defeats the Arup pattern entirely.
- A passphrase that never travels over the channel being verified. If the code word is spoken on the call being authenticated, it authenticates nothing.
- Dual authorisation above a value threshold for payment changes, vendor bank-detail changes and privileged access grants.
- Help-desk hardening for password and MFA reset — the MGM vector. This is the highest-value single process in scope.
And the documented failure modes are procedural too. Two recur: the attacker supplies the callback number (defeated by pre-registration), and manufactured urgency causes staff to bypass the procedure — which is a training and authority problem, not a technology one. Staff must be explicitly authorised to refuse and delay a senior request without career risk, or the control exists on paper only.
Gartner sets the ceiling: “You only need five to nine metrics for each target audience. Do not report everything you know. Prioritize the top five to nine technology dependencies and the top audiences to receive only the highest-value information that drives above-the-line decisions.” Combined with “only above-the-line metrics should be shared with executives”, that is why fifteen run the programme and seven reach the Security Committee.
The selection test each of the seven had to pass
- Does it cover a distinct domain? Two from D1, one each from D2, D3 and D4, one from the foundation. No domain is unrepresented and none is over-represented.
- Is it a leading indicator? It must “expose a problem before it leads to material loss.” This is why validated coverage is on the list and incident counts are not.
- Can it be computed honestly today? Anything needing a denominator we do not have was excluded — which is why reconciliation is on the list and inventory completeness is not.
- Does moving it change a decision? “The value of a metric is its ability to influence decision making. The decisions influenced are typically priorities and investments.”
Thresholds and the action each triggers
Each metric’s exact unit of count, denominator and pass test is in its own specification panel, reachable from the Definition link beside it on the committee set and the programme set.
| Metric | Amber | Red | Action on red |
|---|---|---|---|
| 01 · D1.2Priority controls validated by simulated attack | Falling quarter on quarter | No validation run in a quarter | Validation capacity becomes a funded item rather than a best-effort activity |
| 02 · D1.1AI assets and identities reconciled | Declare-lag rising | Reconciliation rate falling while the estate grows | Discovery tooling escalated; registry treated as advisory until reconciled |
| 03 · D2.1High-risk processes hardened against deepfake impersonation | Any named process unhardened after 60 days | A hardened process bypassed in an exercise | Authority to refuse and delay a senior request is restated in writing |
| 04 · D3.2Machine-identity classes revocable in ten minutes | Rehearsal missed in a quarter | Any crown-jewel-reaching class not revocable | Appetite statement A2 is breached — the gap becomes a named risk with a remediation date |
| 05 · D4.1Production agents with declared scope enforced at runtime | Enforcement coverage static | New high-autonomy agents deployed without enforcement | Autonomy tier reduced until enforcement exists |
| 06 · D4.2Critical IT, OT and cloud assets under continuous posture assessment | New critical assets outside assessment | OT scope slipping beyond FY27 | Scoping decision re-ratified explicitly rather than allowed to drift |
| 07 · F.1Exposure findings with an accountable resolver outside CDR | Outstanding exceeding accepted | Findings ageing without an owner | Escalation to the Security Committee — this is the board accountability Gartner names |
And the eight we deliberately left off. The other eight programme metrics are below the line: they run the work and belong in the monthly CDR review, not the committee pack. Gartner is explicit that both layers are necessary and serve different purposes — but that “only above-the-line metrics should be shared with executives to drive effective business decisions.” Reporting all fifteen upward would reduce the decision value of each one.
The outcome the Security Committee is protecting, stated in the business’s own terms: production output is not interrupted by a cyber event. It is first because it is the only one of the three with a precedent in this industry and a published number attached to it.
| Evidence | Detail |
|---|---|
| TSMC, August 2018 | A misconfigured software installation propagated a variant of WannaCry across production tool controllers, halting several facilities. TSMC put the revenue impact at approximately 3% of third-quarter revenue, roughly US$255M, plus a gross-margin reduction of about one percentage point. |
| The cause matters more than the cost | TSMC stated the infection came from a new software tool installed without virus scanning before connection to the network — not an intrusion campaign. The failure mode was a change-control gap, and that is the same class of gap an unsupervised agent with write authority represents. |
| the enterprise’s own disclosure | The enterprise operates fabrication facilities across several countries with deep interdependency between sites, and its Form 10-K identifies unauthorised access to facilities or technology infrastructure as a named risk. |
| The sector’s standards lag | SEMI E187 and E188 followed the 2018 incident by three to four years, and neither covers the manufacturing execution system, material-control system or factory host layer. |
Which programme metrics have line of sight to it
- Metric 06 — critical IT, OT and cloud assets under continuous posture assessment, because the OT estate is where continuity is lost.
- Segmentation maturity under D4 — Gartner names microsegmentation specifically to “limit lateral movement” between converged IT and OT.
- Metric 01 — validation, because an untested containment path is an assumption.
And the boundary we are explicit about. Response authority inside the production environment is deferred to FY27 by decision. For FY26 this goal is served by assessment, segmentation and validation — not by automated containment. Gartner describes why: converged environments “struggle to support the downtime required for constant patching.”
Process recipes, yield data and design files are the assets that determine whether the enterprise’s technology lead survives. This goal is second because the threat to it is documented in court records, not inferred.
| Evidence | Detail |
|---|---|
| United Microelectronics Corporation | Pleaded guilty to trade-secret theft and was sentenced to a US$60M fine in respect of the enterprise core product technology. The Deputy Attorney General described the conduct as part of a campaign to acquire American technology. |
| Fujian Jinhua | Convicted at trial in February 2024 on the related charges. |
| Competitive pressure is current | The enterprise names ChangXin Memory Technologies as a competitor receiving state support. |
| The regulatory shock is real | The China Cyberspace Administration decision materially affected the enterprise’s revenue, disclosed in its own filings. |
| And AI is named as an IP risk by the enterprise itself | The Form 10-K identifies risks arising from AI use in relation to intellectual property and from new methods of attack. |
Why this goal changes the machine-identity metric’s weighting
Metric 04 — credential classes revocable within ten minutes — is weighted so that any class with reach into recipe or design data must be in scope. A revocation catalogue that covers the easy classes and omits the crown-jewel ones would improve the number while leaving the goal unprotected. That is precisely the failure Gartner’s “direct line of sight” test is meant to catch.
One complication worth stating to the committee. Running inference over controlled technical data may itself be an export-control event depending on where the model is hosted and who can reach it. This is a reason the self-hosted AI Lab sits in the defence set — not only for forensic independence, but because hosted guardrails also refuse legitimate defensive analysis.
The third goal is the one CISOs usually leave off the slide, and the one that makes the programme an enabler rather than a tax. Gartner’s own final step in deriving outcome-driven metrics is to express readiness as a risk to business enablement — not as a security score.
The argument in one line
If we cannot say what an agent is permitted to do, the business cannot safely say yes to it. Runtime enforcement of declared scope (metric 05) is therefore an adoption control as much as a security one — it is what allows a high-autonomy use case to be approved at all.
| What the evidence shows | Consequence for this goal |
|---|---|
| Enterprise AI initiatives are abandoned at a material rate, most often for governance and trust reasons rather than model performance. | Ungoverned adoption is slower in practice, because projects stall at approval or get withdrawn after an incident. |
| Gartner advises risk-tiering frontier AI use cases and applying layered control planes, with composite AI patterns for the high-risk tier. | A tiering scheme is the mechanism that lets low-risk use cases move fast while concentrating scrutiny where it is warranted. |
| Very few peers have a defined programme for non-human identity governance. | Doing this creates a defensible position rather than merely catching up. |
Why this goal is stated as speed and not as “secure AI adoption”. Because “secure adoption” has no direction. Speed is measurable and the business already tracks it, which is what makes the line of sight from a below-the-line protection level to an above-the-line business outcome real rather than rhetorical.
The diagram is not a presentational device. It is the output of a defined method, and the method is what makes the fifteen programme metrics defensible rather than assembled.
| Step | Applied here |
|---|---|
| 1 Identify the business processes and the technology stacks that support them | Production and the factory-software layer; the design and yield-analysis estate; the agent and model estate. |
| 2 Determine the business outcomes those processes deliver | Manufacturing continuity; protection of critical IP and recipe data; speed of the enterprise’s own AI adoption. |
| 3 Identify the technology risks to those outcomes | Interruption of tool control; exfiltration or corruption of recipe data; an agent acting outside its declared scope. |
| 4 Derive the technology outcome-driven metrics | The fifteen below-the-line protection levels, three per domain plus three for the foundation. |
| 5 Express readiness as a risk to business enablement | The readiness statement — which is why it is phrased as what the business cannot yet safely do. |
What makes something an outcome-driven metric and not just a number
- It is a continuous outcome of an identifiable investment — fund it and it improves, defund it and it degrades.
- It sits on a sliding scale, so the committee is choosing a protection level rather than passing or failing.
- It takes the canonical form “% of X”, which is why every metric on the executive set is written that way.
- It has direct line of sight to a business outcome — and if it does not, “the technology investment should not be prioritized as it will drive little value.”
And the rule that cut the committee set from fifteen to seven. “You only need five to nine metrics for each target audience”, and “only above-the-line metrics should be shared with executives.” Gartner is also explicit about what to remove: stop reporting activity counts that no investment decision follows from.
The change being asked for
From “how many days do we take to patch?” to “what is our tolerance to be exploited by a known vulnerability?” — set by asset class.
Why the first question cannot be answered usefully any more
The interval between disclosure and working exploit has compressed from days to minutes for some classes of vulnerability, so a days-to-patch target is a statement about our process rather than about our exposure. Gartner is blunter still: we are “not patching our way out of vulnerability exposure”, and mitigation and compensating control must carry the load where patching cannot. Only 36% of organisations even have infrastructure teams actively engaged.
| What the reframe enables | Because |
|---|---|
| Differentiated targets by asset class | A tolerance of near-zero for the factory-software layer and a higher one for a development sandbox are both rational; a single days-to-patch target for both is not. |
| Decision models rather than severity sorting | CISA’s binding directive and the SSVC model already frame remediation as a decision on exploitability and mission impact. |
| An honest conversation about residual risk | Because the committee is setting a tolerance rather than approving a deadline it will not meet. |
What this does not mean. It is not a licence to stop patching. Gartner’s framing of preemptive security is explicit that these approaches are “not substitutes or replacements for cyber hygiene”. The reframe changes what we measure and report, not what we maintain.
The ask, in the words the board should hear
“cybersecurity teams can only guide vulnerability prioritization. IT operations, product teams and business system owners must be held accountable by the board to fix exposures in their own systems.”
Why this is the critical path and not an administrative preference
Roughly 80% of the controls this programme depends on sit outside Cyber Defense & Resilience. Gartner states the consequence directly: “without widespread business engagement most exposure management functions… are unable to function effectively”, and warns that preemptive tooling can end up “limit[ed] to detecting, rather than fixing issues, thereby increasing, rather than reducing overhead.”
| Dependency area | The accountable owner being requested |
|---|---|
| Machine-identity lifecycle, issuance and revocation mechanism | Identity |
| Segmentation, and binding the agent registry to a runtime enforcement point | Platform and infrastructure & operations |
| Repository and build policy, artefact-registry defaults | Developer experience |
| The factory-software layer — the manufacturing execution system (MES), material-control system (MCS) and factory host systems that SEMI E187 and E188 do not cover | Manufacturing |
| Standard operating procedures for password and MFA reset, supplier interaction and financial transactions | Service desk, procurement, finance and HR |
What is being ratified
Gartner’s four core preemptive domains as the programme’s structure — exposure management, adversary management and threat intelligence, adversary disruption, and posture and policy management — on a foundation of managed services and capabilities, with the fifteen outcome-driven metrics as the definition of success.
Why adopt an external frame rather than write our own
- It is additive, not a replacement: preemptive security “adds focus beyond prevention, detection, and response”, so prevent–detect–contain–respond remains the delivery method underneath.
- Much of it is already bought: Gartner notes many preemptive capabilities already exist in current platforms, which is why tranche 1 needs no new money.
- It gives the committee a defensible external reference for the structure, so debate can be about sequencing and funding rather than about taxonomy.
- It carries its own anti-hype discipline: focus on “how the tools support these outcomes, not on the hype around AI.”
The alternative framing, if the committee prefers three words to four domains. Gartner’s other formulation of the same idea is predict, prioritise, prevent. The content is unchanged; only the label differs. What should not change is that the outcomes are adopted alongside the structure — a frame without metrics is a diagram.
The specific authority requested
Revocation of machine-identity credentials without prior change approval, for an enumerated Tier A list, with the list changeable only by the Security Committee and every action logged and reviewed.
Why speed is the whole point
Average adversary breakout time is 29 minutes; the fastest observed was 27 seconds, and the median hand-off to a second-stage operator 22 seconds. A containment action gated on a change window is not a containment action.
Why machine identity and nothing else
- It is the reversible asset class — revoking a credential breaks a workload, not a person’s day, and it can be reissued in minutes.
- It is the class the reference incident actually abused, and where the adversary held signing authority and could mint credentials faster than they were withdrawn — which is why the catalogue must include authority-level taint, not only credential revocation.
- No human-account, endpoint-isolation or production-network authority is being requested. Those are not reversible in the same sense and are not in scope for this decision.
And the proof obligation we accept in return. The authority should be contingent on rehearsal: a named authority, a written procedure and a measured end-to-end time per credential class, tested quarterly. That is committee metric 04. If the rehearsal lapses, the authority should lapse with it.
The finding
“outdated systems and unused code are no longer just operational problems, but significant cyberthreat liabilities.”
Why frontier models change the calculus specifically
The capability that has genuinely shifted is the speed and scale of finding and weaponising known weakness, not the invention of new attack classes. Legacy and unused code is exactly the surface that rewards cheap, patient, automated review — and unused code is the worst case, because nobody is watching it and nobody will notice it being probed. Gartner’s medium-term observation cuts both ways: models “will enable defenders to inspect an unprecedented volume of source code” — and attackers first.
| What changes if this is accepted | Mechanism |
|---|---|
| Decommissioning gets funded | Because it becomes a risk-reduction line rather than a deferred engineering nicety. |
| Unused code becomes a tracked exposure | Dead endpoints, orphaned repositories and stale artefact registries enter the exposure inventory rather than sitting outside it. |
| Mitigation is legitimised where patching is impossible | Consistent with Gartner’s position that we will not patch our way out of this. |
| The IT/OT case becomes statable | Converged environments “struggle to support the downtime required for constant patching” — which is a technical-debt problem wearing an operations costume. |
Why it belongs on a board agenda rather than an engineering backlog. Because the decision that creates technical debt is a funding and prioritisation decision, and it is made above the level of the teams who inherit it. Naming it as material risk is what allows a resolver team to trade a feature for a decommission without needing to win that argument alone.
The proposition
Approve tranche 1 now — it requires no new money, because three of its seven items are governance or process and the rest use capability already licensed. Then gate tranche 2 on what tranche 1 finds, rather than approving a full-programme budget against assumptions.
The two tranche-1 items that may change tranche 2’s scope
| Item | What it could change |
|---|---|
| The coverage baseline — which techniques are reachable with the telemetry we already have. Today 16 of 78 by the modality we hold, against 69 of 78 by runtime instrumentation. | If the gap is mostly closable by re-pointing existing collection, the tranche-2 telemetry line shrinks. If it is not, the case for runtime instrumentation is made with our own numbers rather than a vendor’s. |
| The gateway delegation test — whether a delegation identity can be minted at the gateway. Attribution is formally non-identifiable from logs alone, and trace-based grouping recovers only 4-6% of a delegation’s events. | If the gateway can carry it, runtime enforcement of declared scope is a build. If it cannot, the fallback is credential-level resolution — weaker, cheaper, and worth knowing before committing to the enforcement design. |
Why sequencing disruption before detection is deliberate
Because it is where we are weakest and where the cheap wins are. Deception needs no inventory, no gateway change and no new telemetry pipeline; SOP hardening is process change; revocation authority is a decision. Gartner notes disruption “provides a clear signal of an attack” and feeds other controls.
What the committee is really approving. Not a budget. A method of deciding the budget — spend a quarter establishing the two numbers that determine the largest line items, then fund against measurements rather than estimates. If tranche 1 produces no findings that change tranche 2, that is itself a useful result.
Every other number on the baseline is unflattering. This one is not, and it is worth understanding why it matters more than it looks.
What the enterprise has already disclosed. “Our Board of Directors administers its cybersecurity risk oversight function directly as a whole, as well as through the Security Committee… which oversees monitoring and incident response, risk mitigation, supply chain…”
Why this is the scarce asset
| What the committee already has | Why it usually has to be built |
|---|---|
| A standing board-level body with an explicit cybersecurity remit | Most programmes must first establish a forum, then earn its attention, then win the right to ask it for authority. That is typically a year. |
| Authority that reaches outside the security function | Decision 02 — naming accountable owners in IT operations, product and business teams — is only actionable because a body with that reach exists. Gartner frames this as a board accountability, not a security one. |
| Standing agenda time for risk-appetite conversations | Decision 01 is a risk-appetite reset. There is no other forum that can make it. |
| An existing route for materiality judgements | Which matters given SEC Regulation S-K Item 106, Item 1C and the Form 8-K Item 1.05 four-business-day determination. |
What it means for the maturity assessment
Programme governance is assessed at MIL2 today and targeted at MIL3 — the highest starting point of anything on the baseline. Six of the fifteen assessed capabilities sit at MIL0. The gap is not governance; it is execution capacity and delegated authority, and both are things this committee can grant.
The baseline is assessed against the Cybersecurity Capability Maturity Model (C2M2), published by the U.S. Department of Energy. Its Maturity Indicator Levels (MILs) are cumulative and are assessed per capability, not as a single organisational score.
| Level | Definition | What it takes to claim it |
|---|---|---|
| MIL0 | The practice is not performed. | Nothing. It is the honest answer for six of our fifteen capabilities. |
| MIL1 | Initial practices are performed, but may be ad hoc. | Someone does it. It need not be written down or repeatable. |
| MIL2 | Practices are documented, and adequately resourced and skilled. | A written procedure, named people and enough capacity to run it. This is the FY27 target for most capabilities. |
| MIL3 | Practices are guided by policy, periodically reviewed, and measured for effectiveness. | Policy, review cadence and an effectiveness measure. Reserved for the two capabilities where the committee is being asked to grant authority: control validation and machine-identity revocation. |
Why not the NIST Cybersecurity Framework Tiers
Because the CSF Tiers are not a maturity scale. They describe the rigour of an organisation’s cybersecurity risk-governance and risk-management practices, and NIST’s own Organizational Profile guidance frames them as a characterisation to inform a target profile — not a ladder to climb. Using them as one produces a number the committee will reasonably read as a score, and which cannot be tied to a specific investment.
What CSF 2.0 is doing in this strategy. Two subcategories are load-bearing rather than decorative: GV.RM-02, that risk appetite and tolerance are established and communicated — which is decision 01 — and the GV.RR category on roles, responsibilities and authorities — which is decision 02. We use CSF for governance structure and C2M2 for capability measurement.
The word “ownership” usually produces a workshop and a matrix. This ask is narrower and harder: a board instruction naming a single accountable owner for each of five dependency areas, with the obligation to accept or reject exposure findings on a defined clock.
| Dependency area | Accountable owner requested | The first thing they would be asked for |
|---|---|---|
| Machine-identity lifecycle | Identity | An enumerated list of credential classes, and a revocation mechanism that works by cohort rather than one credential at a time. |
| Segmentation and enforcement binding | Platform / I&O | Thirty to ninety days of traffic observation before any enforcement, and a decision on where the agent registry binds to a runtime enforcement point. |
| Repository and build policy | Developer experience | Artefact-registry anonymous-access defaults reviewed, and build-time provenance for anything reaching production. |
| The factory-software layer | Manufacturing | An owner for the manufacturing execution system, material-control system and factory host systems — the layer SEMI E187 and E188 do not cover. |
| High-risk standard operating procedures | Service desk, finance, procurement, HR | Out-of-band callback to a pre-registered channel on password and MFA reset, supplier interaction, and financial transactions. |
Why a named owner and not a shared responsibility
Gartner measured the engagement problem: only 36% of organisations have infrastructure teams actively engaged on remediation, and exposure management functions without business engagement are “unable to function effectively.” It also observes that the obstacles here are predominantly non-technical. The segmentation evidence is the starkest version: of fourteen organisations attempting even limited segmentation, eleven failed — mostly for want of an executive champion and application-owner buy-in.
What accountability means in practice, stated so it can be accepted or refused. Not that the owner fixes everything. That the owner accepts a finding with a date, or rejects it with a reason, within a defined window — and that unowned findings escalate to the Security Committee rather than accumulating in a CDR backlog. Committee metric 07 measures exactly this, which is why it earns a place on a seven-metric set.
The strategy is built on two of the twelve reports and corroborated by the other ten. It is worth being explicit about which, because it determines what is load-bearing and what is supporting.
| Role | Reports | What depends on it |
|---|---|---|
| Base | G00859378 Outcomes of Preemptive Cybersecurity; G00799085 Outcome-Driven Metrics for the Digital Era | The four-domain architecture and the entire outcome mechanism. If either is wrong, the strategy is wrong. |
| Threat picture | G00852902 AI-Augmented Attacks; G00852689 the 2026-2027 landscape | The threat tab, including the deepfake prevalence figures and the counterweight numbers. |
| Board framing | G00856541 CISO Board Scenario | Decisions 01, 02 and 05, and the tone of the closing message. |
| Execution | G00837909 CTEM roadmap; G00810627 vulnerability exposure; G00853789 infrastructure cybersecurity 2027; G00858028 frontier-AI checklist | Scoping, resolver teams, the microsegmentation journey, IT/OT governance and risk tiering. |
| Direction | G00836733 preemptive security; G00845745 Future of the CISO 2030 | The market trajectory and the longer-horizon risks. Supporting, not load-bearing. |
| Context only | G00846817 Hype Cycle for Enterprise Architecture | Read; not used. Stated so the corpus count is not mistaken for corpus dependency. |
How it was read
- Every report extracted in full and analysed page by page; dossier 27 carries verbatim quotes with page locators for every claim that appears in this document.
- The two base reports were read end to end before anything was built on them — which is how the five-to-nine metric ceiling and the five-step derivation method came to shape the outcomes tab rather than being retrofitted.
- Where Gartner’s position is more conservative than the prevailing narrative, Gartner’s position is the one used — including that true AI-powered attacks remain rare and that current offensive use is largely experimental.
- Where Gartner and a primary source disagree on a number, both are shown and the primary source is preferred. The Verizon report is currently secondary-sourced and flagged as such.
And the discipline that goes with reading a vendor-analyst corpus. Gartner supplies its own caution and we apply it to this document too: focus on “how the tools support these outcomes, not on the hype around AI”, and treat preemptive capability as additive to rather than a replacement for cyber hygiene. Several claims in the corpus are strategic planning assumptions rather than measurements; where one is used, it is labelled as such.
% of discovered AI assets and machine identities that are fully reconciled1 · What exactly one item is
One object, where an object is either an AI asset or a machine identity, deduplicated on a stated join key.
| Definition and why it is drawn here | |
|---|---|
| AI asset | Counted as an object in five classes only: a network-reachable model endpoint (URL plus deployment name); a registered agent (registry identifier); a tool or Model Context Protocol server exposed to an agent (server URI plus tool name); a retrieval corpus or index an agent can read; a stored model artefact in a registry. |
| Machine identity | A credential-bearing non-human principal, in ten classes: application registration or service principal; system-assigned managed identity; user-assigned managed identity; workload identity federation subject; platform-issued application programming interface (API) key or static secret; OAuth client-credentials grant; personal access token used by automation; workload certificate or SPIFFE identity; cloud role or service account assumed by a workload; signing identity for code or artefacts. |
| Deliberately not counted | Individual prompts, chat sessions, and notebook experiments not reachable from a shared endpoint. These are excluded because they are unbounded and would make the count irreproducible — two people counting on the same day would not agree. |
| The join key | Identities join on the platform object identifier. Assets join on endpoint URL plus deployment name, or on the registry identifier. Objects that cannot be joined are counted separately and flagged unjoinable, and the unjoinable count is published beside the metric as its own error bar. |
2 · The denominator, and where it comes from
The union of all discovery signals, deduplicated. Five signals, in descending order of trustworthiness: S1 resource-tag scan across all cloud subscriptions; S2 six-signal heuristics over application registrations and identities (permission shape, naming pattern, redirect URI, credential type, sign-in pattern, absent owner); S3 gateway and egress telemetry showing traffic to an inference endpoint; S4 a service-management platform record; S5 registry self-declaration. An object is discovered if it appears in at least one signal.
System of record. The the agent registry for assets; the identity platform for identities. The service-management platform holds the reconciliation join.
3 · The test that puts an item in the numerator
| Condition | Requirement |
|---|---|
| R1 — one record | Joined to exactly one record in the system of record. No duplicate, no orphan. |
| R2 — a named owner | A named individual who is currently employed and has acknowledged ownership within the last twelve months. A team name fails. |
| R3 — declared scope | Declared purpose and permitted scope present and non-empty. |
| R4 — complete credential list | Every credential the object holds is enumerated, with an issue date and an expiry. |
4 · How a pass gets wrongly claimed
- Reconciled against the registry alone — which only proves the object declared itself, not that it exists as declared.
- An owner field populated with a distribution list, a team, or a person who has left.
- Partial reconciliation counted as a pass. All four tests must hold; three of four is a fail.
- Silently dropping unjoinable objects instead of publishing the count.
5 · How it is computed
Computed monthly by the discovery pipeline, not by hand. Each signal is a scheduled query; the union and the dedup are scripted so the number is reproducible from the same inputs. Owned by CDR with Identity as the data provider for S1 and S2.
6 · The arithmetic, worked through
| Line | Value |
|---|---|
| Tag scan (S1) returns candidate identities | 340 |
| Heuristics (S2) flag application registrations as agent-like | 128 |
| Gateway telemetry (S3) shows distinct inference endpoints | 44 |
| Registry (S5) holds declared agents | 61 |
| Union after deduplication — the denominator | 418 |
| Have a record in the system of record | 61 |
| Pass all four reconciliation tests — the numerator | 38 |
| Metric | 38 / 418 = 9% |
| Companion — shadow rate (only ever seen by S1–S3) | 357 / 418 = 85% |
What here is real and what is illustrative. The arithmetic above is illustrative, to show how the figure resolves. What is established today is that the registry exists but is self-declared and unreconciled, so the denominator has not yet been enumerated. Publishing the denominator is the first deliverable, not the percentage.
7 · How this number could lie
Why self-declaration is the wrong first step, not the wrong idea
The only published discovery methodology for this problem puts developer attestation fourth of four — tag scan, then heuristics, then reconciliation against the configuration management database (CMDB), then attestation of whatever residue remains, with a deadline and an escalation path. Our registry is step one of one. Invert the order and the same registry becomes a trustworthy join key rather than the population itself.
Why this metric is first to fund
Metrics D1.2, D3.2, D4.1 and D4.2 all need a scoped population before they can be computed at all. Gartner recommends pulling automated information technology, operational technology and cloud inventories from the exposure platform rather than maintaining them separately. The three-layer asset data model in the service-management platform is the intended join point.
% of priority adversary techniques with at least one defence proven by executing the attack in the last 90 days1 · What exactly one item is
One technique on the versioned priority technique list. Not one control — controls are not countable objects, because no two people enumerate them the same way.
| Definition and why it is drawn here | |
|---|---|
| Why the unit is a technique and not a control | A control count can be made to say anything: one platform is one control or forty, depending on who is asked. A technique is an externally defined, enumerable object with a published identifier, so the denominator can be audited by someone outside the programme. |
| What sits underneath | A register of defence claims, each a tuple of (technique identifier, control instance, expected outcome), where expected outcome is prevent, detect or both. The claim register drives the testing; the technique list drives the metric. |
| Deliberately not counted | Techniques we have documented a decision not to defend against. These leave the denominator only by written decision with a named approver, and the count of excluded techniques is published with the metric. |
2 · The denominator, and where it comes from
The priority technique list — named, versioned and published with the metric. Version 1 is built from the techniques MITRE mapped in ATLAS case study AML.CS0068 set within the 78-technique coordination-and-tool-use set used for the coverage baseline, minus documented exclusions. Re-ratified quarterly; the version number travels with every published figure.
System of record. The defence-claim register and the validation run log, held by the validation function. Every pass carries a run identifier so it can be re-examined.
3 · The test that puts an item in the numerator
| Condition | Requirement |
|---|---|
A prevent claim passes when | The emulated action is attempted against a representative target and fails, and the failure is attributable to the named control by a log entry showing the denial. “It did not work” without an attributable denial is not a pass. |
A detect claim passes when | The emulated action executes and an alert of the claimed severity appears in the security information and event management (SIEM) platform within the claimed time budget, and the alert names the technique. A generic anomaly alert is not a pass. |
| Environment | Production, or a production-equivalent replica with the same configuration. A pass achieved only in a lab is recorded as lab pass and does not enter the numerator. |
| Freshness | A pass expires after 90 days, because configuration drifts. The metric is therefore always a statement about the present. |
4 · How a pass gets wrongly claimed
- Counting a technique as covered because an analytic exists. Existence is not evidence. That is metric D2.3, which is a different and easier test.
- A lab pass promoted to a production pass without re-running it in production.
- An alert that fires but cannot be attributed to the emulated action — common when the test runs during normal change activity.
- An alert outside the claimed time budget recorded as a pass because it eventually arrived.
- Shrinking the technique list to raise the percentage, without publishing the exclusions.
5 · How it is computed
Run continuously by the validation and purple-team function; the figure is recomputed at each publication from the run log rather than maintained as a spreadsheet. Evidence for each pass is the run identifier, the target, the timestamp and the attributed log line.
6 · The arithmetic, worked through
| Line | Value |
|---|---|
| Priority technique list v1 | 22 techniques |
| Defence claims made against them | 31 claims |
| Techniques with no claim at all | 7 |
| Claims that passed in production within 90 days | 12 |
| Techniques with at least one passing claim — the numerator | 9 |
| Metric | 9 / 22 = 41% |
| Companion — claim-level pass rate | 12 / 31 = 39% |
| Companion — techniques with no claim | 7 / 22 = 32% |
What here is real and what is illustrative. The arithmetic is illustrative. The established position today is zero for the autonomous technique block: no validation run has been completed against it, so the honest current value is 0%.
7 · How this number could lie
Why this is the single most important number on the page
It is the direct measure of Gartner’s stated benefit for this domain: visibility into “where and why they fail prior to an actual attack.” It is also the only metric in the set that cannot be improved by buying something. It moves when a defence is executed against and holds.
And the reporting discipline that must travel with it
Never publish it as a bare percentage. The coverage framework underneath states its own limits twice and honestly: every technique has an analytic, and every mapping is partial. If any technique ever shows complete coverage, the measurement is wrong. Score detections on robustness, precision and implementation coverage rather than on techniques touched.
% of crown-jewel data stores with a current, machine-generated blast-radius map1 · What exactly one item is
One data store on the crown-jewel register. A store is a named, addressable repository of data — not a system, not a team, not an application.
| Definition and why it is drawn here | |
|---|---|
| The six store classes | Source-code repository; process-recipe store; yield and test-data store; design-file store; model-weights store; signing-key store. |
| Deliberately not counted | Copies and caches that are not separately addressable. Where a copy is separately addressable and reachable by a different principal set, it is a separate store — because its blast radius is different. |
2 · The denominator, and where it comes from
The crown-jewel register: a business-owned, Security-Committee-ratified, versioned list. Inclusion test — a store is crown-jewel if loss, corruption or disclosure would (a) halt or degrade production output, (b) disclose process-recipe or design intellectual property, or (c) allow trusted code or artefacts to be published in the enterprise’s name.
System of record. The register itself, plus the identity graph and exposure platform that generate the maps.
3 · The test that puts an item in the numerator
| Condition | Requirement |
|---|---|
| B1 — effective readers | The resolved principal list with read access, expanded through group and role nesting, including machine identities. The access-control list as written is not the effective list and does not pass. |
| B2 — effective writers and publishers | The narrower set holding write, merge, tag, sign or deploy rights. |
| B3 — reachable egress | The destinations and downstream systems a workload holding the store’s credentials can reach. |
| B4 — the detecting telemetry | The named log source and field that would show a bulk read or an unexpected write. |
| Currency and provenance | All four generated from live platform data within the last 90 days. A hand-maintained list fails regardless of accuracy, because it cannot be re-derived. |
4 · How a pass gets wrongly claimed
- Using the access-control list as written rather than effective permissions. This is the usual failure and it understates the radius by an order of magnitude.
- Mapping human principals only and omitting machine identities — the second usual failure, and the one that matters most here.
- A map with no egress element, which answers who can read it but not where it can go.
- A map with no named detection source, which answers the exposure question but leaves no way to notice abuse.
5 · How it is computed
Generated monthly by query against the identity graph and exposure platform. CDR owns the generation; the business owns the register and its ratification date.
6 · The arithmetic, worked through
| Line | Value |
|---|---|
| Crown-jewel register v1 | 64 stores |
| Have an effective read list | 41 |
| Also have the write and publish list | 22 |
| Also have egress mapped | 11 |
| Also have a named detection source — the numerator | 9 |
| Metric | 9 / 64 = 14% |
What here is real and what is illustrative. Illustrative. The register does not yet exist in ratified form, so neither the numerator nor the denominator can be stated today. Establishing the register is a business task, not a CDR task.
7 · How this number could lie
Why this is the scoping metric for the whole programme
Gartner is explicit that exposure scoping should follow potential business impact rather than threat severity alone. The crown-jewel register is that scope, and four other metrics inherit it: deception placement (D3.1) is scoped to crown-jewel environments, revocation classes (D3.2) are weighted by crown-jewel reach, posture assessment (D4.2) defines critical by crown-jewel path, and the validation target set is chosen for blast radius. Get this register wrong and four other metrics measure the wrong things accurately.
% of named high-risk procedures that pass a live social-engineering test with all four verification controls in place1 · What exactly one item is
One named procedure. A procedure is countable because it has a written trigger, a sequence of steps and a named execution owner. Not a department, not a policy, not a training course.
| Definition and why it is drawn here | |
|---|---|
| The inclusion test | A procedure is in scope if an instruction received over a voice, video or messaging channel can cause: (a) a credential or multi-factor authentication (MFA) factor to change, (b) money to move, (c) a supplier or bank detail to change, or (d) data or access to be granted. |
| Deliberately not counted | Awareness training completion. It is a different thing measured on a different scale, and counting it here would let a training push move a control metric. |
2 · The denominator, and where it comes from
The high-risk procedure register. Gartner names the areas to start from: password reset, partner and supplier interaction, and financial transactions. Version 1 is expected to hold on the order of a dozen procedures — help-desk password reset, help-desk MFA reset, privileged access grant, new supplier onboarding, supplier bank-detail change, payment release above threshold, purchase-order amendment, contract signature, employee data change, physical access grant.
System of record. The procedure register, owned jointly by the service desk, finance, procurement and human resources. The exercise results are held by CDR.
3 · The test that puts an item in the numerator
| Condition | Requirement |
|---|---|
| H1 — out-of-band callback | Verification is completed by calling back on a channel drawn from the system of record, never from the requester. |
| H2 — a secret off-channel | A shared secret or passphrase that never traverses the channel being verified. |
| H3 — dual authorisation | A second authoriser above a stated value or privilege threshold. The threshold must be a number, not a judgement. |
| H4 — stated authority to refuse | The procedure states, in writing, that staff may delay or refuse a request pending verification and that no adverse consequence follows from doing so, including for a request that appears to come from a senior executive. |
| And the deciding evidence | A social-engineering exercise against that specific procedure in the last 180 days that did not succeed. Written-only is not hardened. A procedure bypassed during an exercise reverts to fail immediately. |
4 · How a pass gets wrongly claimed
- The callback number taken from the caller. This is the single most documented failure mode and it defeats the control entirely.
- The passphrase spoken on the same call it is meant to verify.
- A threshold described as “significant” or “unusual” rather than stated as a number.
- Refusal authority that exists in the security policy but not in the procedure the agent is reading at the time.
- Counting the procedure because the text was updated, without an exercise. Manufactured urgency causes staff to bypass procedures they know — only the exercise tests that.
5 · How it is computed
Procedure owners attest the text; CDR runs the exercises on a rolling schedule so every procedure is tested within 180 days. The exercise result, not the attestation, sets the value.
6 · The arithmetic, worked through
| Line | Value |
|---|---|
| High-risk procedure register v1 | 11 procedures |
| Contain all four controls in the written text | 0 |
| Also passed a live exercise in the last 180 days — the numerator | 0 |
| Metric | 0 / 11 = 0% |
What here is real and what is illustrative. Here the zero is real. This capability is assessed at MIL0 — the practice is not performed — and it is the largest single gap against the most prevalent actual threat. The register size of eleven is illustrative; the numerator is not.
7 · How this number could lie
Why this is on the committee set at all
Because it is the most prevalent real AI-augmented attack and the cheapest to mitigate. 41% of organisations have experienced deepfake-enabled social engineering on an audio call and 35% on video; deepfakes account for one in five biometric fraud attempts. And the investment is process redesign and training — no technology purchase.
| Incident | Why it is the relevant precedent |
|---|---|
| Arup, February 2024 — a confirmed US$25.6M loss after an employee joined a video call on which every other participant was synthetic. | Real-time video deepfakes are operationally viable against a competent finance function. |
| MGM Resorts, September 2023 — a help-desk MFA reset, roughly ten minutes to compromise, more than $100M in impact. | No synthetic media was needed. The process weakness is exploitable with a phone call; AI only makes it cheaper at scale. This is the more important case for us, because it is the one our controls must stop first. |
Why detection tooling is not the control
- Commercial detectors lose roughly 45 to 50% of their area under the curve (AUC) on realistic in-the-wild content compared with laboratory benchmarks.
- Liveness certification structurally excludes injection attacks — the technique actually used — so a certified product can be fully exposed to it.
- Content provenance is destroyed by any screenshot or re-encode, and the absence of a credential is not evidence of fakery.
- Untrained people identify deepfakes in roughly 0.1% of trials.
% of priority threat scenarios converted into an implementable requirement with a named individual owner1 · What exactly one item is
One scenario on the scenario register. A scenario is a named adversary objective plus the sequence of techniques used to reach it against a named the enterprise asset. It is countable because it is a register entry with an identifier.
| Definition and why it is drawn here | |
|---|---|
| What is not a scenario | A threat actor name. A technique on its own. A news article. These become scenarios only when written against a named asset with an objective, which is what makes them testable. |
| Register size is capped deliberately | A register of two hundred scenarios is a backlog, not a plan. Version 1 is capped so that conversion is achievable within a quarter, and the cap is published. |
2 · The denominator, and where it comes from
The scenario register, derived from the programme’s priority intelligence requirements (PIRs) and re-ratified quarterly.
System of record. The scenario register and the detection-requirement tracker, joined on scenario identifier.
3 · The test that puts an item in the numerator
| Condition | Requirement |
|---|---|
| Five attributes, all required | The scenario has produced at least one requirement that is: (1) written in implementable terms — data source, logic intent, expected precision; (2) assigned to a named individual; (3) carries a due date; (4) carries a state of open, built, validated or retired; (5) is traceable back to the scenario identifier and forward to a built analytic or an explicit decision. |
| An explicit decline is a pass | A documented “we will not detect this, here is why, and here is the compensating control” counts as converted. This is deliberate — without it the metric punishes honest triage and nobody will record a decline. |
4 · How a pass gets wrongly claimed
- An owner recorded as a team or a function. Teams do not convert scenarios; people do.
- A requirement written as an intention (“improve coverage of credential abuse”) rather than as something an engineer can build.
- A scenario marked converted with no forward traceability, so nobody can tell whether the analytic was ever built.
- Declines recorded without a compensating control, which turns the escape hatch into a way of clearing the register.
5 · How it is computed
Computed from the requirement tracker at each quarterly scenario review, joined on scenario identifier. Owned by threat intelligence, with detection engineering as the receiving function.
6 · The arithmetic, worked through
| Line | Value |
|---|---|
| Scenario register v1 | 18 scenarios |
| Have a requirement with all five attributes | 6 |
| — of which built or validated | 4 |
| — of which declined with a named compensating control | 2 |
| Metric | 6 / 18 = 33% |
What here is real and what is illustrative. Illustrative. Today there is no scenario-to-requirement conversion process, so this capability sits at MIL0 and the register does not exist.
7 · How this number could lie
Why this metric exists rather than a threat-intelligence volume measure
Because intelligence that does not change a control is overhead. Gartner puts it plainly: adversary management “is of most value when combined with one of the other pillars to provide actionable enforcement.” This metric is the join between D2 and everything else — it measures whether the intelligence function produces work the rest of the programme can act on.
And it is the mechanism Gartner recommends for the roadmap itself. Scenario simulations should feed the capability roadmap rather than sit beside it, which is why the conversion record is required to be traceable forward to a built analytic or a written decline. The traceability is the deliverable; the percentage is just its summary.
% of priority techniques with a named log source that is collected today and carries the required fields1 · What exactly one item is
One technique — the same versioned priority technique list used by D1.2. Same denominator, different test: D2.3 asks whether the evidence would exist; D1.2 asks whether we proved we would catch it.
| Definition and why it is drawn here | |
|---|---|
| Why the two metrics share a denominator | So that the pair can be read as a gap. If D2.3 is high and D1.2 is low, we have the data and lack the analytics. If both are low, we have a telemetry problem first. Different denominators would make that comparison meaningless. |
2 · The denominator, and where it comes from
The priority technique list, version-matched to D1.2. The published coverage baseline uses the 78-technique coordination-and-tool-use set.
System of record. The detection-engineering coverage matrix, with field population verified by query against the SIEM rather than asserted.
3 · The test that puts an item in the numerator
| Condition | Requirement |
|---|---|
| T1 — named and collected | A specific log source and field set is named, and that source is actually being ingested today and retained for at least the stated window. A source that could be enabled does not pass. |
| T2 — required fields populated | The minimum fields the technique requires are present and populated in real events, verified by sampling. A schema that permits a field is not the same as a field that carries data. |
| The metric is T1 and T2 together | Deliberately the harder of the two available bars. |
| What is explicitly out of scope here | Whether an analytic exists (D2.3 does not ask) and whether it was proven to work (D1.2). |
4 · How a pass gets wrongly claimed
- Counting a source because the vendor documents it. Verify by sampling events; documentation is not data.
- Counting a schema field that is defined but empty in practice — the most common false pass in agent telemetry today.
- Counting a technique as covered by a generic source that would contain the event in principle but has no field distinguishing it.
- Reporting a single coverage percentage with no robustness dimension, which implies techniques are equally covered when they are not.
5 · How it is computed
Recomputed monthly from the coverage matrix, with the field-population check run as a scheduled query. Owned by detection engineering.
6 · The arithmetic, worked through
| Line | Value |
|---|---|
| Priority set — coordination and tool-use techniques | 78 |
| Pass T1 and T2 by the telemetry modality collected today — the numerator | 16 |
| Metric | 16 / 78 = 21% |
| Would pass under runtime instrumentation, not deployed | 69 / 78 = 88% |
What here is real and what is illustrative. These figures are real, not illustrative — they come from the generated coverage matrix. The 16-versus-69 gap is the central technical fact of the baseline, and establishing whether it is closable with collection we already own is a tranche-1 deliverable.
7 · How this number could lie
What the 16-versus-69 gap actually means
It is not a gap in analytics. It is a gap in where the telemetry is taken from: gateway-level observation sees a fraction of what runtime instrumentation sees, because the interesting behaviour happens between an agent and its tools rather than at the network boundary. That makes the gap a collection-architecture decision with a cost attached — which is why decision 06 gates the tranche-2 telemetry line on establishing this number with our own estate rather than a reference one.
One caveat that changes how the number should be read. Some of the most valuable detections need no broad telemetry at all. “An agent invoked a tool it never declared” requires only the registry joined to tool-call records, and needs no behavioural baseline. Of the eleven published hunting queries for agent activity, none performs that join. So a low coverage percentage does not mean nothing useful can be built today.
% of crown-jewel environments with at least one live, monitored, tested decoy on the intruder’s path1 · What exactly one item is
One environment — a network or platform boundary containing at least one store from the crown-jewel register. Environments are countable because the boundary is a configuration object.
| Definition and why it is drawn here | |
|---|---|
| Why environments and not decoys | Counting decoys rewards volume. Counting environments asks the question that matters: is there anywhere an intruder can reach a crown jewel without touching something that tells us? |
| Density is tracked separately | One qualifying decoy per environment is the pass bar here. Placement quality and density are tracked as a companion, because a single token in a large environment is weak but is not zero. |
2 · The denominator, and where it comes from
The set of environments derived from the crown-jewel register (D1.3). It inherits that register, so it inherits its ratification date too.
System of record. The deception deployment inventory, joined to the alert-routing configuration.
3 · The test that puts an item in the numerator
| Condition | Requirement |
|---|---|
| C1 — on the path | Reachable by an intruder who has reached that boundary, and placed where enumeration would encounter it — not in an unused subnet. |
| C2 — monitored, not merely logged | Touching it raises an alert routed to a monitored queue with a defined response procedure. A log entry nobody is watching fails. |
| C3 — credible | Naming, metadata and age consistent with real assets in the same environment, and not identifiable by the known public fingerprinting techniques. |
| C4 — tested end to end | Someone touched it and the alert arrived at the queue, within the last 180 days. |
4 · How a pass gets wrongly claimed
- Decoy deployed, alert unrouted. This is the most common failure and it produces a metric that looks healthy and detects nothing.
- A decoy that is fingerprintable, which turns it into a signal to the adversary that they are being watched.
- A stale decoy whose metadata no longer matches the environment around it.
- A decoy placed off the path, where it is safe but never encountered.
5 · How it is computed
Computed from the deployment inventory monthly; the C4 test is scheduled so every decoy is exercised within 180 days. Owned by CDR.
6 · The arithmetic, worked through
| Line | Value |
|---|---|
| Crown-jewel environments | 12 |
| With any decoy deployed | 0 |
| Passing all four conditions — the numerator | 0 |
| Metric | 0 / 12 = 0% |
What here is real and what is illustrative. The zero is real — nothing is placed today. The denominator of twelve is illustrative and depends on the crown-jewel register.
7 · How this number could lie
Why this is in the first thirty days rather than FY27
Because it needs no inventory, no gateway change and no new telemetry pipeline — and because disruption is where we are weakest. Gartner’s stated benefit for the domain: it “may slow down an attack as an attacker spends time… on a target of no value” and it provides “a clear signal of an attack.”
And why the evidence supports it more strongly than most controls
- In a controlled study of language-model agents, bait was taken at roughly 78% against a human baseline of about 37% — agents are more susceptible to deception than people, not less.
- The same work found a recognition-action gap of 73.4%: agents frequently identified a lure as suspicious and interacted with it anyway.
- Attention diversion was statistically absent — decoys did not distract agents from real objectives, so the control adds signal without adding risk.
- National-level trials across 121 organisations and 14 vendors give the operational counterpoint: the control fails on routing and response, not on placement.
% of machine-identity credential classes revocable cohort-wide in under ten minutes, measured at the resource1 · What exactly one item is
One credential class, where a class is the tuple (credential type × issuing platform × revocation mechanism). For example: “application-registration client secret in the corporate tenant”, “user-assigned managed identity”, “Kubernetes service-account token in cluster group A”, “artefact-signing key”.
| Definition and why it is drawn here | |
|---|---|
| Why classes and not credentials | Because revocation capability is a property of the mechanism, not of the individual credential. Counting credentials would make the metric move with the size of the estate rather than with our ability to act — it would fall as the estate grew even if capability improved. |
| Enumerated once, then versioned | The class list is established once and re-ratified quarterly. It is small — on the order of a dozen to twenty entries — which is what makes the metric readable. |
| Scope rule | Any class with reach into a crown-jewel store must be in scope. Classes leave scope only by Security Committee decision. |
2 · The denominator, and where it comes from
The enumerated credential-class list. It is published, because a percentage over an unpublished class list is not auditable.
System of record. The identity platform for issuance; the revocation rehearsal log for the measured times.
3 · The test that puts an item in the numerator
| Condition | Requirement |
|---|---|
| V1 — a named authority | A role that can authorise revocation for that class without a change-advisory board. |
| V2 — a written procedure | Current within twelve months. |
| V3 — cohort capability | The mechanism can revoke the whole class, or a defined subset, in one action — not credential by credential. |
| V4 — a measured time under ten minutes | From decision to the resource refusing the credential. Measured, not estimated. The measurement point is the resource, not the directory. |
| V5 — rehearsed within 90 days | With the measured time recorded against the run. |
4 · How a pass gets wrongly claimed
- Measuring at the directory rather than at the resource. Revoking a secret does not invalidate access tokens already issued until they expire — so a class can look revocable in the directory and remain fully usable at the resource for the token lifetime. Where the platform cannot close that window, the class fails until a continuous-evaluation mechanism is in place.
- A procedure that exists but has never been executed, so the time is an estimate.
- Credential-by-credential revocation counted as cohort capability. Under time pressure the difference is the whole control.
- An authority named in a document but requiring a change window in practice.
5 · How it is computed
Measured at each rehearsal and recorded per class with the run identifier. Owned by Identity as the mechanism provider, with CDR as the authority holder.
6 · The arithmetic, worked through
| Line | Value |
|---|---|
| Enumerated credential classes | 14 |
| With a named authority and a written procedure | 0 |
| With a measured end-to-end time under ten minutes — the numerator | 0 |
| Metric | 0 / 14 = 0% |
What here is real and what is illustrative. The zero is real: nothing is defined today — no authority, no target time, no rehearsal. The class count of fourteen is illustrative until the enumeration is done, and that enumeration is the first deliverable.
7 · How this number could lie
Why ten minutes, and why machine identity specifically
Average adversary breakout time is 29 minutes, with a fastest observed of 27 seconds and a median hand-off to a second-stage operator of 22 seconds. A revocation that takes a change window is not a control. And machine identities are the reversible asset class — revoking one breaks a workload rather than a person’s day, and it can be reissued in minutes. That is what makes pre-authorised action defensible here and nowhere else yet.
Revocation at the credential level is not sufficient
If the adversary holds signing material they can mint valid credentials faster than we withdraw them — which is what happened in the reference case. So the catalogue must include authority-level actions: issue-time invalidation, lease revocation by prefix, signing-authority taint, and continuous session signals.
% of confirmed incidents whose first signal, in the post-incident timeline, came from a disruption control1 · What exactly one item is
One confirmed incident — an event that passed triage into the incident process with a severity assigned. Not an alert.
| Definition and why it is drawn here | |
|---|---|
| First signal, defined precisely | The earliest alert or observation by timestamp that the post-incident review assesses as relating to the incident, whether or not anyone acted on it at the time. This is deliberate: the metric measures what the estate produced, not what the analyst noticed. |
| The six source categories | Every confirmed incident’s first signal is classified as exactly one of: disruption control (a decoy touched, a honeytoken used, a revocation-triggered failure observed); threat hunt; detection analytic; third-party or external notification; user report; discovered during unrelated work. This metric is the first category. |
2 · The denominator, and where it comes from
Confirmed incidents in the rolling four-quarter window. The window is four quarters because a single quarter produces a number too small to interpret.
System of record. The post-incident review record. The classification is made at review time and is not revisited afterwards.
3 · The test that puts an item in the numerator
| Condition | Requirement |
|---|---|
| The pass condition | The post-incident timeline names the first signal, and its source category is disruption control. |
| Reporting form | Always as numerator and denominator with absolute counts shown, never as a bare percentage. With fewer than about twenty incidents in the window, a percentage alone is misleading. |
4 · How a pass gets wrongly claimed
- Classifying by which alert triggered the response rather than which signal came first. The two are frequently different, and the difference is itself worth reporting.
- Counting alerts instead of incidents, which makes the denominator a function of tuning.
- Publishing a percentage on a denominator of three.
5 · How it is computed
Assembled at each post-incident review; the metric is recomputed quarterly over the trailing four quarters. Owned by the incident-response function.
6 · The arithmetic, worked through
| Line | Value |
|---|---|
| Confirmed incidents, rolling four quarters | 23 |
| First signal was a detection analytic | 11 |
| First signal was a user report | 6 |
| First signal was third-party notification | 4 |
| First signal was a threat hunt | 2 |
| First signal was a disruption control — the numerator | 0 |
| Metric | 0 / 23 = 0% |
What here is real and what is illustrative. Illustrative distribution. The zero in the disruption row is structurally certain, because no disruption controls are deployed — but the other rows are not the enterprise figures and the classification has not yet been applied retrospectively. Doing that retrospective classification is cheap and would give a real baseline within weeks.
7 · How this number could lie
Why carry a metric that cannot be directly influenced
Because it is the only measure that tests whether deception and hunting earn their place rather than merely being deployed. D3.1 counts placement; this counts payoff. Without it, a fully green D3.1 could coexist with a deception capability that has never once been the thing that told us.
And it is the number that would have changed the reference incident. In the reference case the defender’s own account identifies the failure as one of escalation, not detection — signals existed and did not become an incident quickly enough. A first-signal classification applied consistently is how that pattern becomes visible before the post-mortem rather than during it.
% of production agents whose registry-declared scope is refused at runtime when exceeded1 · What exactly one item is
One production agent: an agent with a registry identifier that is either reachable by someone other than its author, or runs on a schedule.
| Definition and why it is drawn here | |
|---|---|
| Deliberately not counted | Experiments reachable only by their own author. This exclusion is what makes the count stable — without it the denominator tracks developer activity rather than production exposure. |
| Declared scope, defined | The registry record’s enumerated permitted set: tools, data sources, egress destinations, and a maximum autonomy tier. All four must be enumerated for the record to be scoreable. |
2 · The denominator, and where it comes from
Registered agents that meet the production test. Published alongside the total registered count, so the exclusion is visible.
System of record. The the agent registry, joined to the enforcement point's policy decision log.
3 · The test that puts an item in the numerator
| Condition | Requirement |
|---|---|
| E1 — an enforcement point that can refuse | Calls traverse a gateway, proxy or sidecar that is able to deny a call, not merely observe it. |
| E2 — delegation identity resolved | The enforcement point knows which registered agent, acting for which principal, is making the call — not merely which credential was presented. |
| E3 — registry evaluated at call time | The declared set is read from the registry at the moment of the call, so a scope change takes effect without a redeploy. A scope compiled into the agent at build time fails. |
| E4 — refusals logged and alerted | A denial is recorded with the technique-relevant fields and raises an alert. |
| The deciding evidence | An out-of-scope tool call is attempted in production and the refusal plus the log entry are observed. Configuration evidence alone does not pass. |
4 · How a pass gets wrongly claimed
- An enforcement point that logs but cannot deny. Observation is not enforcement and the distinction is the entire metric.
- Scope compiled in at build time, which cannot be changed under incident conditions.
- Credential-level attribution recorded as delegation identity. Attribution is formally non-identifiable from logs alone, and trace-based grouping recovers only 4 to 6% of a delegation’s events — so if the gateway cannot mint a delegation identity, E2 is partial and the agent fails.
- Counting an agent because an AI security posture management (AI-SPM) tool has inventoried it. That is a static discipline and cannot deliver this.
5 · How it is computed
Computed from the enforcement point's policy configuration joined to the registry, with the E-condition evidence held per agent. Owned jointly by CDR and AI Enablement — this is a build, not a procurement.
6 · The arithmetic, worked through
| Line | Value |
|---|---|
| Registered agents | 61 |
| Meet the production test — the denominator | 34 |
| Traverse an enforcement point that can deny (E1) | 0 |
| Pass all of E1 to E4 — the numerator | 0 |
| Metric | 0 / 34 = 0% |
What here is real and what is illustrative. The zero is real: nothing binds the registry to runtime enforcement today. The agent counts are illustrative pending the D1.1 enumeration.
7 · How this number could lie
The finding that makes this a build rather than a purchase
AI security posture management cannot deliver this. It is a static configuration and inventory discipline. As one vendor states the distinction: “static AI-SPM tells you what an agent can do; runtime-informed AI-SPM tells you what it actually does.” Gartner’s AI trust, risk and security management framing is reported to acknowledge that it “doesn’t adequately cover the governance of the autonomous actions these models can take.”
The detection this unlocks, which nobody ships
Once E2 and E3 hold, declared scope becomes comparable as well as enforceable — which gives the highest-value agent detection available: an agent invoked a tool it never declared. It needs no behavioural baseline and no anomaly model. Eleven published hunting queries exist for agent activity and none of them performs that join.
One dependency worth testing in week one. Whether the gateway can carry a delegation identity at all — the proposed mechanism is a World Wide Web Consortium (W3C) baggage header carried in the Model Context Protocol request metadata. If it cannot, the fallback is credential-level resolution: weaker, workable, and far better known before the enforcement design is committed than after.
% of critical assets assessed against a named standard at least weekly, with drift raised as an owned finding1 · What exactly one item is
One asset, at the granularity the assessing tool reports — host, cluster, cloud resource, or tool controller. The granularity is published, because it determines the number and is the easiest thing to shift quietly.
| Definition and why it is drawn here | |
|---|---|
| Critical, defined | On the crown-jewel path — holds, processes, or has write authority into a store on the crown-jewel register — or is a production tool controller or factory-software component. |
| Why factory software is named explicitly | SEMI E188 explicitly excludes the manufacturing execution system, the material-control system and factory host systems; E187 addresses supplier-provided equipment. Neither covers the factory-software layer that holds write authority into the tools, and that is the layer reachable from information technology. It is a gap in the industry standards, not in the enterprise’s implementation, so this metric names it rather than inheriting the omission. |
2 · The denominator, and where it comes from
Critical assets from the asset register, across information technology, operational technology and cloud — reported with the three sub-populations visible, because their coverage differs greatly.
System of record. The asset register, with the assessing platforms as the evidence source.
3 · The test that puts an item in the numerator
| Condition | Requirement |
|---|---|
| P1 — a named written standard | A benchmark or internal baseline, with its version recorded. “Hardened” without a named standard fails. |
| P2 — at least weekly, unattended | Assessment runs at least every seven days without human initiation. |
| P3 — drift becomes an owned finding | A deviation produces a finding with an owner and a date — not merely a changed dashboard state. |
| P4 — absence detection | The platform knows what it is not seeing: an asset present in the register but not reporting becomes a finding. Without P4, unassessed assets are indistinguishable from passing ones. |
4 · How a pass gets wrongly claimed
- Omitting P4. This is the near-universal omission and it is the difference between a coverage figure and a marketing figure.
- Counting a point-in-time scan as continuous assessment.
- A dashboard that shows drift without creating an owned finding, so nothing follows from it.
- Reporting one blended percentage across information technology, operational technology and cloud, which hides the operational-technology position entirely.
5 · How it is computed
Computed weekly from the assessing platforms joined to the asset register, with the three sub-populations reported separately. Owned by CDR with platform and infrastructure operations as the remediating function.
6 · The arithmetic, worked through
| Line | Value |
|---|---|
| Critical assets on the register | 2,140 |
| — information technology and cloud | 1,870 |
| — operational technology and factory software | 270 |
| Under a seven-day unattended assessment with drift findings and absence detection — the numerator | 640 |
| Metric | 640 / 2,140 = 30% |
| — operational technology sub-population | approximately 0% |
What here is real and what is illustrative. Illustrative. Cloud posture is partial today, AI-SPM is absent, and the operational-technology estate is largely unassessed — but the register has not been consolidated, so no numerator can be stated yet.
7 · How this number could lie
Why assessment and response are deliberately separated for operational technology
Response authority inside the production environment remains deferred to FY27 by decision. Assessment is achievable where response authority is not, so this metric includes operational technology while the containment metrics exclude it. Gartner describes the underlying constraint precisely: “Fragmented IT/OT convergence creates severe risks, as current systems lack unified governance and struggle to support the downtime required for constant patching.”
The segmentation journey this metric sits inside
| Finding | Consequence for sequencing |
|---|---|
| Gartner recommends “the journey from macrosegmentation to network security microsegmentation to limit lateral movement.” | FY27, not tranche 1. |
| Discovery first: 30 to 90 days of traffic observation before enforcement; a pilot in 8 to 12 weeks; a first enterprise segment in 3 to 6 months. | Enforcing before dependency mapping is the usual cause of outage-driven rollback. |
| Of fourteen organisations attempting even limited segmentation, eleven failed — predominantly for want of an executive champion and application-owner buy-in. | Secure the sponsor and the application owners before the tooling. The failure mode is organisational. |
| For legacy process equipment, network-enforced segmentation is the only viable primary control — agents cannot run and virtual local area network (VLAN) tagging is often unsupported. | The production environment path is network-level and sits behind the FY27 decision. |
% of untrusted-input workloads with egress denied by default and the denial proven by test1 · What exactly one item is
One workload — a deployable unit with its own network identity: a container workload, a function, or a virtual machine role. The granularity is published.
| Definition and why it is drawn here | |
|---|---|
| Untrusted-input, defined by one test | Does the workload process content it did not author and cannot fully validate? Concretely: inference over user or external content; document and email ingestion; web retrieval; repository-content processing; and Model Context Protocol tool servers reachable by agents. The question to ask is whether the content it reads could contain instructions. |
| Deliberately not counted | Workloads whose only inputs are internally generated and schema-validated. Including them would dilute the denominator with the easy cases. |
2 · The denominator, and where it comes from
Workloads meeting the untrusted-input test, enumerated from the platform inventory rather than declared.
System of record. The platform inventory joined to the network-policy and egress-proxy configuration.
3 · The test that puts an item in the numerator
| Condition | Requirement |
|---|---|
| G1 — explicit allow-list | A named list of permitted destinations; everything else refused. A wildcard, or a broad content-delivery range that reaches anywhere, fails. |
| G2 — enforced where the workload cannot reconfigure it | Network policy with an enforcing data plane, or an egress proxy the workload must traverse. |
| G3 — denials logged | With source workload identity and destination, so a blocked attempt is a signal and not just a failure. |
| G4 — proven by test | An attempt to a non-allowed destination, made from inside the workload, is refused — demonstrated within the last 90 days. |
4 · How a pass gets wrongly claimed
- A policy authored with no enforcing plane. Kubernetes NetworkPolicy silently does nothing without a container network interface (CNI) plugin that enforces it — this is a documented trap and it produces a confident false pass.
- An allow-list containing a wildcard or a broad provider range, which permits egress to anywhere behind that provider.
- Domain name system (DNS) resolution permitted to arbitrary resolvers, which leaves an exfiltration path open regardless of the allow-list.
- A proxy that can be bypassed by addressing a destination directly by internet protocol address.
5 · How it is computed
Computed monthly from configuration, with the G4 test scheduled per workload class rather than per workload. Owned by platform engineering with CDR defining the untrusted-input test.
6 · The arithmetic, worked through
| Line | Value |
|---|---|
| Workloads meeting the untrusted-input test | 87 |
| Have an explicit allow-list (G1) | 12 |
| Also enforced at a point the workload cannot change (G2) | 6 |
| Also logged and proven by test (G3, G4) — the numerator | 4 |
| Metric | 4 / 87 = 5% |
What here is real and what is illustrative. Illustrative. The untrusted-input population has not been enumerated, which is the first task here.
7 · How this number could lie
Why this is framed as isolation rather than input filtering
Because prompt filtering is a probabilistic control against an adversary who can iterate, while egress control is a deterministic one. Gartner’s related guidance points at securing agent actions rather than agent prompts. The practical consequence: we do not need to detect a malicious instruction if the workload that receives it cannot reach anywhere useful.
And it is the control that pairs with D4.1. D4.1 constrains what an agent may call; D4.3 constrains where a workload may reach. Together they close the two paths that matter, and either alone leaves the other open. D4.3 is the cheaper of the two and does not depend on the registry, which is why it can proceed while the enforcement-point question is still being settled.
% of deduplicated exposure findings with a named individual outside CDR who has accepted or rejected them on the clock1 · What exactly one item is
One deduplicated finding. The deduplication rule is one finding per (weakness × affected asset group × owner) — not one per scanned instance.
| Definition and why it is drawn here | |
|---|---|
| Why the deduplication rule has to be stated | Because per-instance counting inflates numerator and denominator together and makes the percentage meaningless: a single misconfiguration across four hundred hosts would dominate the figure. Stating the rule is what makes the metric comparable quarter to quarter. |
| Deliberately not counted | Informational findings with no remediation path, and findings whose only owner is CDR itself — the latter are reported separately, because a programme that owns its own findings is measuring nothing. |
2 · The denominator, and where it comes from
All deduplicated findings raised in the reporting period.
System of record. The exposure-management workflow system. Email does not count as a record — if the acceptance is not in the workflow, the finding is unowned.
3 · The test that puts an item in the numerator
| Condition | Requirement |
|---|---|
| A1 — a named individual outside CDR | Recorded as accountable. A team, a function or a distribution list fails. |
| A2 — accepted or rejected on the clock | The owner has either accepted with a remediation date, or rejected with a stated reason and a named compensating control — within the acceptance window. Ten working days is the proposed window; the number is a Security Committee decision, not a CDR one. |
| A3 — in the workflow system | Not in a mailbox, a spreadsheet or a meeting minute. |
| A rejection is a pass | The metric measures whether ownership was resolved, not whether everything gets fixed. This is essential: without it the metric punishes honest risk acceptance and nobody will record a decision at all. |
4 · How a pass gets wrongly claimed
- Assigning to a team. Teams do not accept risk; named people do, and that is the whole point of the board ask.
- Counting per scanned instance, which makes the figure move with scanner configuration.
- Treating an unanswered finding as accepted by default. Silence is the failure state this metric exists to make visible.
- A rejection with no compensating control named, which is a decision to do nothing recorded as a decision.
5 · How it is computed
Computed from the workflow system at the end of each reporting period. Owned by CDR as the raiser; the value is determined entirely by behaviour outside CDR, which is deliberate.
6 · The arithmetic, worked through
| Line | Value |
|---|---|
| Deduplicated findings raised in the quarter | 1,340 |
| With a named individual outside CDR | 384 |
| — of which accepted with a date, within the window | 151 |
| — of which rejected with a reason and a compensating control | 59 |
| Resolved ownership on the clock — the numerator | 210 |
| Metric | 210 / 1,340 = 16% |
| Companion — escalated to the Security Committee unowned | 174 |
What here is real and what is illustrative. Illustrative. What is established is that there is no systematic assignment today, and that roughly 20% of the relevant controls sit with CDR — so about four fifths of the remediation capability is outside the team raising the findings.
7 · How this number could lie
Why this is the programme’s critical path
Gartner is unusually direct: “cybersecurity teams can only guide vulnerability prioritization. IT operations, product teams and business system owners must be held accountable by the board to fix exposures in their own systems.” And separately: “Without widespread business engagement most exposure management functions… are unable to function effectively.” Only 36% of organisations have infrastructure teams actively engaged on remediation. Gartner also observes that the obstacles here are predominantly non-technical.
And what the ask actually is. Not goodwill, and not a responsibility workshop. A board instruction naming accountable owners for five dependency areas: Identity for the machine-identity lifecycle; platform and infrastructure operations for segmentation and enforcement binding; developer experience for repository and build policy; manufacturing for the factory-software layer; and the service desk, finance, procurement and human resources for the high-risk procedures. Of the six decisions requested, this and the appetite reframe are the only two that cost nothing — and between them they determine whether the other four are deliverable.
1 · What exactly one item is
One capability. There is no population to divide by, so there is no percentage.
| Definition and why it is drawn here | |
|---|---|
| Why this one breaks the “% of” form, deliberately | Gartner’s canonical outcome-driven metric form is “% of X”, and for fourteen of the fifteen we hold to it. Here the population is a single capability. Inventing a denominator to preserve the form would be dishonest, so this is reported as met, partly met or not met, with the condition list shown and the quarterly test result as the supporting number. That gives the committee a trend without a production sitericated percentage. |
2 · The denominator, and where it comes from
Not applicable. Reported as a state with four named conditions.
System of record. The capability's own documentation, plus the quarterly refusal-test log.
3 · The test that puts an item in the numerator
| Condition | Requirement |
|---|---|
| F2a — local weights | A self-hosted inference capability on infrastructure the security team controls, with model weights held locally. |
| F2b — it will do the work | It processes incident artefacts — malicious code, adversary prompts, exfiltrated content — without refusal. Verified quarterly against a standing test set, not assumed. |
| F2c — evidence handling | Artefact handling aligns with digital-evidence requirements and the environment is documented for later admissibility. |
| F2d — export-control determination | A written assessment of inference over controlled technical data, because where the model runs and who can reach it changes the answer. |
| Reporting form | Met only when all four hold. Partly met lists which conditions fail. Not met is the position today. |
4 · How a pass gets wrongly claimed
- Treating a hosted model with a commercial agreement as equivalent. The agreement governs data use; it does not govern refusal behaviour.
- Standing the capability up and never running the refusal test, which leaves F2b an assumption at exactly the moment it matters.
- Omitting the export-control determination, which is the condition most likely to stop the capability being usable on the artefacts that matter most.
5 · How it is computed
Reviewed quarterly. The refusal test is the only part that produces a number, and that number is the supporting measure.
6 · The arithmetic, worked through
| Line | Value |
|---|---|
| F2a — local weights | Not met |
| F2b — quarterly refusal test passing | Not met |
| F2c — evidence-handling alignment | Not met |
| F2d — written export-control determination | Not met |
| State | Not met |
What here is real and what is illustrative. This is the real position, not an illustration. The capability does not exist today.
7 · How this number could lie
partly met must name which conditions fail.Why this belongs in a strategy document at all
Because it cannot be procured mid-incident, and because the failure is documented rather than hypothetical. Hosted models refused legitimate blue-team work at 2.72× the neutral rate across 2,390 real security tasks, and could not be worked around by rephrasing. In the reference incident the defender hit exactly this.
And one requirement that is easy to miss. Every evidence-touching inference must record the model hash, engine version and decode parameters, or the analysis cannot be reproduced later. That is a logging requirement on day one, not a refinement — retrofitting it means the early analysis is the analysis that cannot be defended.
% of named response workflows with a stated concurrency limit, overflow behaviour and an exercise in the last 12 months1 · What exactly one item is
One response workflow: a named, triggerable procedure with an entry condition and a defined end state. Countable because it has an identifier in the runbook set.
| Definition and why it is drawn here | |
|---|---|
| Deliberately not counted | Guidance documents with no trigger and no end state. They may be useful, but they cannot be exercised and they cannot be delegated, which is what this metric is about. |
| Why capacity rather than existence | Because “we have a runbook” says nothing about whether it survives twenty concurrent invocations — and volume is precisely what changes when the adversary operates at machine speed. |
2 · The denominator, and where it comes from
The named workflow set. Published with the count, because a small well-specified set is a better position than a large vague one and the metric should not disguise which we have.
System of record. The runbook repository, with the exercise log as the evidence for the last condition.
3 · The test that puts an item in the numerator
| Condition | Requirement |
|---|---|
| W1 — a stated concurrency limit | How many simultaneous executions the workflow can sustain, as a number. |
| W2 — overflow behaviour | What happens when the limit is exceeded: queue, degrade, or escalate. “Undefined” is the answer this metric exists to eliminate. |
| W3 — a lifecycle | A named owner, a review date, and a retirement or supersession state. |
| W4 — exercised within 12 months | With the result recorded. |
4 · How a pass gets wrongly claimed
- A capacity described qualitatively (“scales as needed”) rather than as a number.
- An overflow behaviour that is actually the absence of one — work silently queuing with no alert is not a defined behaviour.
- An exercise that tested the happy path at concurrency one, which is the usual form and tests nothing this metric is asking about.
5 · How it is computed
Computed from the runbook repository and the exercise log, reviewed quarterly. Owned by the incident-response function.
6 · The arithmetic, worked through
| Line | Value |
|---|---|
| Named response workflows | 46 |
| State a concurrency limit and overflow behaviour (W1, W2) | 11 |
| Also have a lifecycle and an exercise within 12 months — the numerator | 8 |
| Metric | 8 / 46 = 17% |
What here is real and what is illustrative. Illustrative. Detection-engineering and response capacity is not established today, so the workflow set has not been enumerated with capacity attributes.
7 · How this number could lie
Why measure the baseline before automating anything
Gartner projects that AI agents will autonomously manage 25% of incident-response workflows for data-security events by 2028. That is a strategic planning assumption rather than a measurement, and it should be read as one. But either way the implication holds: a workflow whose capacity and end state are undefined cannot be delegated, because there is no specification for the thing taking it over to meet.
And the honest reason this is a foundation metric rather than a response one. Gartner rests the four domains on managed services and capabilities, noting most organisations need partner support “unless they have large and high-level security teams.” A workflow set with stated capacity is also the artefact that makes a partner conversation possible — it is what a service provider would be asked to meet.
Why these three and not a longer list
Because a longer list would let anything in. Each gate names a specific reason the existing security programme cannot be expected to cover the item: the asset is new, the control has broken, or the required speed exceeds human execution. Anything that does not fit one of those three is, by definition, work the existing programme is already accountable for.
| The argument against | The answer |
|---|---|
| “Gate 2 will be used to smuggle everything in — any control can be said to be under AI pressure.” | Fair, and it is the weakest gate. So G2 requires a named mechanism — speed, scale or fidelity — and evidence that the control previously held. “AI makes phishing worse” does not pass; “voice verification no longer works, here is the prevalence figure” does. |
| “Parking the posture-management and patching work leaves real risk unowned.” | It leaves it owned elsewhere, which is where it already was. The failure mode we are avoiding is a programme that inherits every unsolved problem in the estate because it is the newest thing with budget. |
| “The dependency bucket is where the plan will actually fail.” | Almost certainly true, and that is why it exists as a named bucket with dates rather than as an assumption. Gartner’s finding is that these obstacles are predominantly non-technical, and that exposure functions without business engagement are “unable to function effectively”. |
One consequence worth accepting openly. Applying these gates shrinks the programme. Eleven core items is a smaller thing than the earlier draft described, and it will read as less ambitious. It is also the only version that can be delivered and measured, and the only version that answers the question actually asked.