MDR Detection Maturity

Detection at scale, prerequisites, and the five-level maturity ladder — for MDR builders, MDR buyers, and SOC leaders making the build-vs-partner call.

70 / 25 / 5
universal / tenant-profile / custom rule mix
<24h
mature TTP-to-fleet-rule deployment
<1%
false-positive rate ceiling per rule per tenant
5
levels of detection-engineering maturity
Three audiences: (1) MDR providers assessing their own detection-engineering maturity; (2) buyers evaluating MDR vendors during RFP/procurement; (3) in-house SOC leaders making the build-vs-partner decision. The vocabulary below is shared; the questions differ. This playbook delivers a maturity scorecard, capability gap analysis, and a remediation roadmap.
Core thesis: The structural fact that shapes everything in MDR detection engineering is the fleet-scale constraint. You can't manually tune per tenant when you have 500–5,000 of them. You don't get customer-specific context for new tenants on day one. Your false-positive math is multiplied by N customers. Most "detection engineering" wisdom from in-house SOC literature doesn't transfer, because in-house SOC has N=1 and MDR has N=large.

What Detection at Scale Actually Looks Like

The four structural constraints that an MDR detection program must work within. Every architectural decision below traces back to one of these.

1. False-positive economics multiply

At 1,000 tenants, a rule firing one false positive per day per tenant = 1,000 tickets/day from a single rule. Math collapses unless: (a) the rule is tuned to <1% FP rate, OR (b) auto-triage absorbs FPs at the agent layer before a human sees them. This single constraint forces nearly every detection-design trade-off in mature MDRs.

2. Can't manually tune per tenant

In-house SOCs whitelist specific behaviors for specific users ("the dev team uses Postman; ignore their anomalous Graph calls"). MDR can't — whitelist sprawl at 1,000 tenants becomes unmaintainable. Either (a) auto-baseline behavior per tenant (UEBA), or (b) write rules that are tenant-agnostic from the start, or (c) charge premium tiers explicitly for customer-specific tuning.

3. Telemetry heterogeneity

Every customer's log sources, retention, sampling, schema, and vendor mix differs. A rule that says "look at Defender for Endpoint table X" doesn't work for the 30% of customers on CrowdStrike. Mature MDRs invest heavily in telemetry normalization — Sigma cross-platform rules, parser libraries, vendor-agnostic data models (Splunk CIM, Elastic ECS, OCSF).

4. Time-to-detect / time-to-respond SLAs

Contractually committed. Detection must fire within minutes, not hours. Drives toward streaming detection over batch, drives investment in real-time correlation engines, drives the prioritization of universal rules over customer-specific (universal rules deploy faster).

The Three-Tier Detection Model

Every mature MDR converges on roughly this volume split. The ratios shift but the layers are universal — because they're driven by the scale constraints above, not by style preference.
Tier What it is Example Volume
Universal / baseline Patterns true regardless of tenant; ship fleet-wide on day one, no per-tenant config authenticationProtocol = "deviceCode" from IP outside named-locations ~70%
Tenant-profile Rule shape universal, baseline per-tenant; requires behavioral state but no per-tenant authoring Inbox-rule creation by an account that hasn't created one in 90 days ~25%
Customer-specific Bespoke detection for high-tier customers (regulated sector, known APT exposure, insider risk) Detection of named threat actor TTPs for a financial-sector client ~5%
Where MDR R&D actually lives: moving things from tier 2 to tier 1 (making more behaviors detectable universally without per-tenant baseline) and from tier 3 to tier 2 (productizing custom detections for fleet use). Per-tenant tuning is the cost-of-goods problem; productized detection is the margin.

The Six Detection Sources Mature MDRs Blend

No single source covers the full ATT&CK matrix. Mature MDRs combine all six with explicit ownership and quality controls per source.

1. Vendor-provided detections

Microsoft Defender XDR built-ins, Sentinel content hub, Splunk ES content packs, CrowdStrike Falcon detections. The floor — not the ceiling. Coverage lags novel TTPs by weeks to months. Free if you have the vendor, but you don't own the rule and can't tune it without copying.

2. Open-source rule libraries

Sigma (SIEM-agnostic, the lingua franca for portable rules), Sublime Security (email, MIT-licensed, 3,478-commit library), Elastic SIEM rules, SOC Prime TDM, SOCFortress, mthcht repos. Sigma is how mature MDRs share and port rules across customer SIEMs.

3. Threat intel-driven detections

CTI report drops translate into rules within hours. Storm-2372 / EvilTokens / FlowerStorm publishes → device-code-flow rule deployed fleet-wide within 24–48h. Median time from public TTP disclosure to fleet deployment is the single highest-signal MDR maturity metric. Mature: <24h. Average: 1–2 weeks. Immature: never deployed at all.

4. Internal red team / threat research

MDRs that do their own attacker emulation feed observations back into detection. Atomic Red Team and Caldera for technique replay; internal pentesters for adversary emulation; threat-researcher hires whose job is to find new TTPs and write detections for them. This is where AI-assisted detection generation (LLM → Sigma rule from CTI text) is currently emerging.

5. Customer-incident-driven (the flywheel)

Every IR produces post-mortem detection rules: "the attacker did X — write a rule for X." The MDR caseload becomes the competitive moat: more cases → more observed TTPs → more rules → better coverage → wins more cases. Eye Security's 279-case BEC dataset and 79% MFA-bypass insight is exactly this flywheel; CrowdStrike's annual Global Threat Report is the same pattern at FAANG-scale.

6. UEBA / behavioral analytics

Defender for Identity, Splunk UBA, Exabeam, Vectra, Darktrace. ML-baselined per identity, host, or asset. High false-positive rate without tuning, but it's the only path for attacks where rules-based detection fundamentally misses — behavioral anomalies that look like normal operations one signal at a time. Critical for token theft, NHI compromise, slow-burn lateral movement.

Detection-source maturity test: ask the MDR which of these six they own (own = have a named team, deployment pipeline, and quality metric for). Immature MDRs own 1–2 (vendor + maybe OSS). Mature MDRs own all 6 with named ownership and explicit per-source SLAs.

Prerequisites — What You Need Before Detection at Scale Is Possible

If any of these is missing or weak, the rest of the program collapses. Build (or buy) them in this order — each unlocks the next.

1. Detection-as-code pipeline

Detection rules in git. PR-reviewed. CI-tested against synthetic + replayed real-attacker data. Deployed via automation with explicit rollback. The same engineering rigor as product code, because that's what it is. Tools: Tines, Anvilogic, custom GitHub Actions pipelines, Panther's detection-as-code IDE, Splunk's Content Management Platform.

Without this: every detection change is a manual deployment; rollback is "edit and pray"; no test coverage; rule-quality regressions go undetected; new hires can't ship rules safely.

2. Telemetry normalization layer

Customer log sources reduced to a common schema before detection runs. Sigma + per-vendor parsers, Elastic Common Schema (ECS), Splunk Common Information Model (CIM), or the newer OCSF (Open Cybersecurity Schema Framework). Lets one rule run across all customer telemetry without per-vendor variants.

Without this: every rule needs N vendor-specific copies; rule sprawl is multiplicative; coverage gaps are vendor-shaped not threat-shaped; cross-customer threat correlation is impossible.

3. False-positive triage layer (analyst tier OR AI triage)

Because fleet-scale FP economics don't work with raw alerts hitting humans. Either (a) a Tier-1 analyst team sized to handle the FP volume (expensive, scales linearly), or (b) AI-triage absorbing the noise before humans see it (capex-heavy, scales sublinearly — the AI-in-SOC investment thesis). The mature pattern is increasingly (b) with (a) for escalation.

Without this: either burn out analysts on a 1,000-tickets-per-day floor, or tune rules so conservatively that real threats fall through.

4. ATT&CK coverage measurement

Every detection tagged to MITRE ATT&CK T-numbers. Coverage matrix maintained: which techniques have detections, which don't, what fraction of fleet receives each detection. ATT&CK Evaluations (or internal red team simulations) measure actual detection rate per technique. Mature MDRs publish (at least internally) a coverage map updated quarterly.

Without this: "we have 10,000 rules!" becomes the marketing pitch instead of "we cover 87% of ATT&CK Enterprise with median detection latency 4 minutes." Rule count is a noise metric; coverage is the signal metric.

5. Customer-incident → detection rule flywheel

Every IR engagement produces a post-mortem rule. Operationally: SOC handover into detection-engineering with a defined hand-off ritual. Specifically: the analyst who closed the incident drafts a Sigma rule (or escalates "this is detectable but I don't see how"); detection-engineering team reviews and deploys. This is the moat — the longer you operate, the more attacker TTPs you've seen and encoded.

Without this: incidents close without producing reusable detection content; each customer's compromise teaches you nothing transferable; the caseload moat doesn't compound.

6. Time-to-detect SLAs as a metric, not vibes

Mean time from attacker action to first alert, measured continuously, broken down per detection source and ATT&CK technique. Reported to customers monthly. Used as an engineering target — "T1078.004 has 8-minute median; bring it to <3." Without explicit SLAs, latency drifts upward silently as alert volume grows.

Without this: nobody owns the latency curve; rule additions silently slow the pipeline; the "we detect within minutes" marketing line is unfalsifiable.

The Five-Level MDR Detection Maturity Ladder

Where most MDRs in the market sit today: levels 1–3, clustered at the top of 2. Levels 4–5 are the differentiated tier. The buyer's job is to figure out which level a vendor actually operates at — not which level they claim.

Level 1 Ad-hoc / Vendor-Rule Consumer

Detection content is whatever the SIEM/EDR vendor ships. No internal rule authoring. New TTPs detected only when the vendor publishes an update. No detection-as-code — rules edited in vendor UIs. No ATT&CK mapping. FP volume managed by tuning thresholds or muting noisy rules.

Signals: "we use Microsoft 365 Defender" as the answer to "how do you detect" • rule count is the marketing metric • ATT&CK coverage map doesn't exist or is vendor-generated • new TTP coverage measured in weeks-to-months

Level 2 Repeatable / OSS-Rule Curator

Internal team curates and deploys OSS rules (Sigma library, Sublime, Elastic SIEM rules). Rules are version-controlled but deployment is partly manual. Some ATT&CK tagging, no formal coverage measurement. FP triage via analyst-tier with rough rule-quality feedback. TTP-to-deployment latency 3–14 days.

Signals: "we maintain a Sigma rule library" • some custom rules, mostly vendor + OSS • analyst-tier-driven tuning • "we cover ATT&CK" without a number behind it

Level 3 Defined / Detection-as-Code + ATT&CK-Mapped

Full detection-as-code: git, PR review, CI tests, automated deployment, rollback. Every rule tagged to ATT&CK; coverage matrix maintained internally. Customer-incident → rule flywheel is functional (some incidents produce rules; not all). Telemetry normalization layer exists. Time-to-deploy: 1–3 days for known-public TTPs. FP triage is analyst-tier, occasionally augmented with automation.

Signals: "our detection content is in git" • can answer "what fraction of ATT&CK Enterprise do you detect" with a number • has named detection-engineering function distinct from analyst-tier • deploys CTI-driven rules within 1–3 days

Level 4 Managed / Fleet TTP-Velocity + FP-Rate SLAs

Median time from public TTP disclosure to fleet rule deployment is measured and committed-to (<24h). Per-rule FP-rate SLAs (<1% per tenant per rule). UEBA layer in production for branches that rules-based detection misses. Internal red team or attacker emulation feeds detection R&D. Customer-incident flywheel is systematic — every IR closes with a hand-off ritual to detection-engineering. Coverage map published quarterly internally; sometimes externally.

Signals: publishes time-to-deploy metric • per-customer detection-latency reporting • has internal red team or named threat researcher hire • UEBA in production not just on the roadmap • can answer "what's your median time from Microsoft publishing a TTP to your fleet deploying detection"

Level 5 Optimized / AI-Triage + Caseload-Flywheel Moat

AI-triage layer absorbs FPs at fleet scale, making detection sensitivity tunable past what an analyst-only tier could process. Caseload-driven detection (proprietary observations from N years of incidents) is the differentiating moat — not vendor rules, not OSS. Detection-engineering velocity is competitive grade; new TTPs deploy within hours. Coverage and latency published externally as proof. Sometimes contributes back to OSS (Sigma, Sublime) as a credibility move.

Signals: AI-triage in production, with measured analyst-time reduction • original threat research published • OSS contributions to Sigma/Sublime/ATT&CK • coverage map is public • new TTPs detected within hours, not days
Realistic distribution today: majority of MDRs operate at L2 with marketing claims of L4. Genuinely-L4 vendors are a minority. L5 is rare and usually requires either (a) FAANG-scale telemetry, or (b) deep AI-in-SOC investment, or (c) a competitive sector moat (Eye Security's BEC dataset + AIVD-origin threat research is structurally a path to L5).

Buyer Questions — What to Ask an MDR During Evaluation

Procurement-grade questions designed to surface actual maturity vs marketed maturity. Each question has a green-flag answer (real maturity), a yellow-flag answer (level 2–3), and a red-flag answer (level 1 with marketing).

Detection-content questions

"What's your median time from public TTP disclosure to fleet detection deployment?"
Green flag: a specific number in hours • Yellow: "depends on the TTP, usually within a week" • Red: "we deploy as we see them" or no answer
"What fraction of ATT&CK Enterprise techniques do you have detection coverage for?"
Green flag: a specific percentage with a coverage-map artifact • Yellow: "most" / "the relevant ones" • Red: rule count quoted instead
"How many detections originated from your own incident caseload vs vendor/OSS?"
Green flag: a ratio with examples • Yellow: "a lot of them" • Red: doesn't understand the question
"Show me a detection you wrote in the last 30 days for a TTP that was first publicly disclosed in the same window."
Green flag: produces the Sigma rule and the CTI source • Yellow: describes one verbally • Red: can't produce an example

Operational-maturity questions

"Walk me through your detection-as-code pipeline. Where does a new rule start, what tests does it pass, how is it deployed, how is it rolled back?"
Green flag: specific tools (git, CI tests, deployment pipeline, rollback mechanism) • Yellow: "we use version control" • Red: rules edited in vendor UI
"What's your false-positive rate per rule per tenant, and how do you measure it?"
Green flag: <1%, with a defined measurement methodology • Yellow: "low" • Red: never thought about it as a per-rule per-tenant metric
"What's your mean time from alert generation to analyst eyes-on, and what absorbs FP volume before that?"
Green flag: specific time, plus a named auto-triage or AI layer • Yellow: tiered-analyst model with no automation • Red: "all alerts go straight to analysts"

Source-mix questions

"Of the six detection sources — vendor, OSS, threat intel, internal research, customer incident, UEBA — which do you own with a named team and explicit pipeline?"
Green flag: 5–6 with named owners • Yellow: 3–4 • Red: 1–2 (usually vendor + OSS)
"Do you contribute back to OSS detection content (Sigma, Sublime, ATT&CK), and where can I see those contributions?"
Green flag: public PRs or named research • Yellow: "we share with peer MDRs" • Red: doesn't contribute

Anti-pattern questions (red-flag-elicitation)

"How many detection rules do you have?" — the question itself is a red flag. The right metrics are coverage and latency. If the vendor's first metric is rule count, they're optimizing the wrong thing.
"How big is your threat intel team?" — team size is a vanity metric. The question that matters is: what's the throughput (TTPs translated to detections per month) and the latency (TTP disclosure to fleet deployment).
"Can I see your ATT&CK coverage map and your time-to-deploy report for the last quarter?" — vendors who can produce these are at L3+. Those who say "we don't share that externally" are usually at L2.

Engagement Economics

Three engagement shapes depending on whether you're an MDR building, a buyer evaluating, or a SOC leader making the build-vs-partner decision.
Audience Engagement Duration Investment
MDR provider Maturity assessment + capability gap analysis + roadmap to next level (L2→L3, L3→L4) 5–10 days €6,250 – 12,500
MDR buyer / procurement Vendor-evaluation framework + RFP question set + scoring rubric + meeting-by-meeting evaluation support 3–5 days €3,750 – 6,250
SOC build-vs-partner Internal SOC maturity baseline + economic comparison vs MDR options + transition plan 5–8 days €6,250 – 10,000
Strategic retainer Ongoing maturity advisory + quarterly review + RFP support as needed 2 days/quarter €10,000/yr
Build-vs-partner economics, briefly: running an internal SOC at L3+ requires at minimum a detection-engineering function (3–5 FTE, ~€300–500K/yr loaded), telemetry-normalization tooling (~€50–100K/yr), and 24x7 analyst tier (~€500K–1M/yr depending on shift model). Floor ~€1M/yr for a small-but-real SOC. Most organizations below ~500 employees should partner; most above ~5,000 employees should consider building. The middle is the real decision space and where this engagement focuses.

OSS Detection Content

Coverage Measurement + Attack Emulation

Detection-as-Code Platforms

  • Tines — workflow + detection automation with CE tier
  • Anvilogic — detection content marketplace + governance
  • Panther — cloud-native SIEM with detection-as-code IDE