What Detection at Scale Actually Looks Like
1. False-positive economics multiply
At 1,000 tenants, a rule firing one false positive per day per tenant = 1,000 tickets/day from a single rule. Math collapses unless: (a) the rule is tuned to <1% FP rate, OR (b) auto-triage absorbs FPs at the agent layer before a human sees them. This single constraint forces nearly every detection-design trade-off in mature MDRs.
2. Can't manually tune per tenant
In-house SOCs whitelist specific behaviors for specific users ("the dev team uses Postman; ignore their anomalous Graph calls"). MDR can't — whitelist sprawl at 1,000 tenants becomes unmaintainable. Either (a) auto-baseline behavior per tenant (UEBA), or (b) write rules that are tenant-agnostic from the start, or (c) charge premium tiers explicitly for customer-specific tuning.
3. Telemetry heterogeneity
Every customer's log sources, retention, sampling, schema, and vendor mix differs. A rule that says "look at Defender for Endpoint table X" doesn't work for the 30% of customers on CrowdStrike. Mature MDRs invest heavily in telemetry normalization — Sigma cross-platform rules, parser libraries, vendor-agnostic data models (Splunk CIM, Elastic ECS, OCSF).
4. Time-to-detect / time-to-respond SLAs
Contractually committed. Detection must fire within minutes, not hours. Drives toward streaming detection over batch, drives investment in real-time correlation engines, drives the prioritization of universal rules over customer-specific (universal rules deploy faster).
The Three-Tier Detection Model
| Tier | What it is | Example | Volume |
|---|---|---|---|
| Universal / baseline | Patterns true regardless of tenant; ship fleet-wide on day one, no per-tenant config | authenticationProtocol = "deviceCode" from IP outside named-locations |
~70% |
| Tenant-profile | Rule shape universal, baseline per-tenant; requires behavioral state but no per-tenant authoring | Inbox-rule creation by an account that hasn't created one in 90 days | ~25% |
| Customer-specific | Bespoke detection for high-tier customers (regulated sector, known APT exposure, insider risk) | Detection of named threat actor TTPs for a financial-sector client | ~5% |
The Six Detection Sources Mature MDRs Blend
1. Vendor-provided detections
Microsoft Defender XDR built-ins, Sentinel content hub, Splunk ES content packs, CrowdStrike Falcon detections. The floor — not the ceiling. Coverage lags novel TTPs by weeks to months. Free if you have the vendor, but you don't own the rule and can't tune it without copying.
2. Open-source rule libraries
Sigma (SIEM-agnostic, the lingua franca for portable rules), Sublime Security (email, MIT-licensed, 3,478-commit library), Elastic SIEM rules, SOC Prime TDM, SOCFortress, mthcht repos. Sigma is how mature MDRs share and port rules across customer SIEMs.
3. Threat intel-driven detections
CTI report drops translate into rules within hours. Storm-2372 / EvilTokens / FlowerStorm publishes → device-code-flow rule deployed fleet-wide within 24–48h. Median time from public TTP disclosure to fleet deployment is the single highest-signal MDR maturity metric. Mature: <24h. Average: 1–2 weeks. Immature: never deployed at all.
4. Internal red team / threat research
MDRs that do their own attacker emulation feed observations back into detection. Atomic Red Team and Caldera for technique replay; internal pentesters for adversary emulation; threat-researcher hires whose job is to find new TTPs and write detections for them. This is where AI-assisted detection generation (LLM → Sigma rule from CTI text) is currently emerging.
5. Customer-incident-driven (the flywheel)
Every IR produces post-mortem detection rules: "the attacker did X — write a rule for X." The MDR caseload becomes the competitive moat: more cases → more observed TTPs → more rules → better coverage → wins more cases. Eye Security's 279-case BEC dataset and 79% MFA-bypass insight is exactly this flywheel; CrowdStrike's annual Global Threat Report is the same pattern at FAANG-scale.
6. UEBA / behavioral analytics
Defender for Identity, Splunk UBA, Exabeam, Vectra, Darktrace. ML-baselined per identity, host, or asset. High false-positive rate without tuning, but it's the only path for attacks where rules-based detection fundamentally misses — behavioral anomalies that look like normal operations one signal at a time. Critical for token theft, NHI compromise, slow-burn lateral movement.
Prerequisites — What You Need Before Detection at Scale Is Possible
1. Detection-as-code pipeline
Detection rules in git. PR-reviewed. CI-tested against synthetic + replayed real-attacker data. Deployed via automation with explicit rollback. The same engineering rigor as product code, because that's what it is. Tools: Tines, Anvilogic, custom GitHub Actions pipelines, Panther's detection-as-code IDE, Splunk's Content Management Platform.
2. Telemetry normalization layer
Customer log sources reduced to a common schema before detection runs. Sigma + per-vendor parsers, Elastic Common Schema (ECS), Splunk Common Information Model (CIM), or the newer OCSF (Open Cybersecurity Schema Framework). Lets one rule run across all customer telemetry without per-vendor variants.
3. False-positive triage layer (analyst tier OR AI triage)
Because fleet-scale FP economics don't work with raw alerts hitting humans. Either (a) a Tier-1 analyst team sized to handle the FP volume (expensive, scales linearly), or (b) AI-triage absorbing the noise before humans see it (capex-heavy, scales sublinearly — the AI-in-SOC investment thesis). The mature pattern is increasingly (b) with (a) for escalation.
4. ATT&CK coverage measurement
Every detection tagged to MITRE ATT&CK T-numbers. Coverage matrix maintained: which techniques have detections, which don't, what fraction of fleet receives each detection. ATT&CK Evaluations (or internal red team simulations) measure actual detection rate per technique. Mature MDRs publish (at least internally) a coverage map updated quarterly.
5. Customer-incident → detection rule flywheel
Every IR engagement produces a post-mortem rule. Operationally: SOC handover into detection-engineering with a defined hand-off ritual. Specifically: the analyst who closed the incident drafts a Sigma rule (or escalates "this is detectable but I don't see how"); detection-engineering team reviews and deploys. This is the moat — the longer you operate, the more attacker TTPs you've seen and encoded.
6. Time-to-detect SLAs as a metric, not vibes
Mean time from attacker action to first alert, measured continuously, broken down per detection source and ATT&CK technique. Reported to customers monthly. Used as an engineering target — "T1078.004 has 8-minute median; bring it to <3." Without explicit SLAs, latency drifts upward silently as alert volume grows.
The Five-Level MDR Detection Maturity Ladder
Level 1 Ad-hoc / Vendor-Rule Consumer
Detection content is whatever the SIEM/EDR vendor ships. No internal rule authoring. New TTPs detected only when the vendor publishes an update. No detection-as-code — rules edited in vendor UIs. No ATT&CK mapping. FP volume managed by tuning thresholds or muting noisy rules.
Level 2 Repeatable / OSS-Rule Curator
Internal team curates and deploys OSS rules (Sigma library, Sublime, Elastic SIEM rules). Rules are version-controlled but deployment is partly manual. Some ATT&CK tagging, no formal coverage measurement. FP triage via analyst-tier with rough rule-quality feedback. TTP-to-deployment latency 3–14 days.
Level 3 Defined / Detection-as-Code + ATT&CK-Mapped
Full detection-as-code: git, PR review, CI tests, automated deployment, rollback. Every rule tagged to ATT&CK; coverage matrix maintained internally. Customer-incident → rule flywheel is functional (some incidents produce rules; not all). Telemetry normalization layer exists. Time-to-deploy: 1–3 days for known-public TTPs. FP triage is analyst-tier, occasionally augmented with automation.
Level 4 Managed / Fleet TTP-Velocity + FP-Rate SLAs
Median time from public TTP disclosure to fleet rule deployment is measured and committed-to (<24h). Per-rule FP-rate SLAs (<1% per tenant per rule). UEBA layer in production for branches that rules-based detection misses. Internal red team or attacker emulation feeds detection R&D. Customer-incident flywheel is systematic — every IR closes with a hand-off ritual to detection-engineering. Coverage map published quarterly internally; sometimes externally.
Level 5 Optimized / AI-Triage + Caseload-Flywheel Moat
AI-triage layer absorbs FPs at fleet scale, making detection sensitivity tunable past what an analyst-only tier could process. Caseload-driven detection (proprietary observations from N years of incidents) is the differentiating moat — not vendor rules, not OSS. Detection-engineering velocity is competitive grade; new TTPs deploy within hours. Coverage and latency published externally as proof. Sometimes contributes back to OSS (Sigma, Sublime) as a credibility move.
Buyer Questions — What to Ask an MDR During Evaluation
Detection-content questions
Operational-maturity questions
Source-mix questions
Anti-pattern questions (red-flag-elicitation)
Engagement Economics
| Audience | Engagement | Duration | Investment |
|---|---|---|---|
| MDR provider | Maturity assessment + capability gap analysis + roadmap to next level (L2→L3, L3→L4) | 5–10 days | €6,250 – 12,500 |
| MDR buyer / procurement | Vendor-evaluation framework + RFP question set + scoring rubric + meeting-by-meeting evaluation support | 3–5 days | €3,750 – 6,250 |
| SOC build-vs-partner | Internal SOC maturity baseline + economic comparison vs MDR options + transition plan | 5–8 days | €6,250 – 10,000 |
| Strategic retainer | Ongoing maturity advisory + quarterly review + RFP support as needed | 2 days/quarter | €10,000/yr |
OSS Detection Content
- Sigma — SIEM-agnostic detection rule format + community ruleset
- Sublime Security — Email detection platform (MIT)
- Elastic Detection Rules — ML + traditional rules
- SOC Prime — commercial + free TDM
- mthcht awesome-lists — curated detection content collections
- SOCFortress — community-driven detection content
Coverage Measurement + Attack Emulation
- MITRE ATT&CK — framework + Navigator for coverage mapping
- Atomic Red Team — technique replay library
- CALDERA — automated adversary emulation
- Center for Threat-Informed Defense — ATT&CK Evaluations methodology
- OCSF — Open Cybersecurity Schema Framework
- Atomic Red Team (canonical repo)
Detection-as-Code Platforms
- Tines — workflow + detection automation with CE tier
- Anvilogic — detection content marketplace + governance
- Panther — cloud-native SIEM with detection-as-code IDE
- Splunk SOAR / Content Management
- Sigma rule packs via GitHub Actions deployment
- Custom: git + GitHub Actions + Sigma compilers + per-SIEM deployers
Companion Playbooks
Detection at scale is the operational backbone for everything the threat-specific playbooks describe. Where detection-engineering meets the specific threats:
Identity Hardening
The six-branch identity attack tree — which branches are tractable as universal rules, which require tenant-profile baselines, which need UEBA.
Identity Breach Response
Where the incident-driven detection flywheel actually closes: every IR produces a post-mortem rule for the detection-engineering pipeline.
BEC Defense
Email detection content is the largest single OSS rule library (Sublime Security, MIT-licensed, 3,478 commits). Universal tier territory.
AI Agent Security
The AI-triage layer that makes fleet-scale FP economics work — level 5 territory.