Detection Engineering at Scale

From Artisan Rules to Deployable Artifacts

SQR
Signal Quality Ratio
The metric that matters
3
Detection Tiers
Universal → Profiled → Research
CI/CD
Detection Pipeline
Version → Test → Deploy → Measure

The KPI That Matters

Signal Quality Ratio

Signal quality ratio is the fraction of findings that someone actually acts on, divided by the total findings surfaced. Not mean time to respond. Not alerts processed per shift. Not coverage percentage across the MITRE matrix. Those metrics are easy to game and they incentivise the wrong behavior.

MTTR rewards closing tickets fast. That means auto-closing, bulk-closing, and building workflows that mark things resolved before an analyst has formed an opinion. Coverage percentage rewards writing detections against every technique in a framework, regardless of whether those detections produce actionable signal in the environment they are deployed to. Both metrics look good in quarterly reports. Neither tells you whether the detection program is actually finding threats.

SQR improves quarter over quarter when the detection engineering program is healthy. It goes up when noisy detections get tuned or retired instead of left running. It goes up when customer profiling removes detections that generate noise in environments where the underlying technology does not exist. It goes up when detection authors have access to outcome data — did this detection lead to a real investigation, or did it just create work?

The uncomfortable truth: Most SOCs measure how fast they close tickets. That incentivises closing tickets, not finding threats. A detection program measured by signal quality ratio cannot succeed by generating noise and resolving it efficiently — it has to produce findings that are worth investigating.

Three Detection Tiers

Not every detection requires the same engineering investment. The tiers below separate commodity detections from competitive advantage. Misallocating engineering effort across these tiers is the most common failure mode in detection programs.

1 Universal Baseline

Ransomware staging patterns. Credential dumping from LSASS. Known command-and-control infrastructure. Mass phishing campaigns. These are the detections that every environment needs and that every mature platform already provides. They justify the XDR subscription.

The correct engineering posture for universal detections is buy or standardise. Sigma rules for SIEM-layer detections. YARA rules for file and memory scanning. Platform-native detections from the EDR or XDR vendor. There is no competitive advantage in writing a custom Kerberoasting detection when CrowdStrike, SentinelOne, and Defender all ship one that has been tested against millions of endpoints.

Engineering effort at this tier goes into validation, not creation. Confirming that the vendor detection fires in the specific environment. Confirming that the log sources required are actually being collected. Confirming that the alert routes to a human who knows what to do with it. The baseline tier is about coverage assurance, not rule authorship.

2 Sector and Profile-Specific

A healthcare organization with on-premise PACS servers and HL7 interfaces has a different threat surface than a SaaS company running entirely in AWS. A 50-person company with a single Microsoft 365 tenant has different noise patterns than a 2,000-person company with hybrid Active Directory. Detection engineering that ignores customer profiling generates noise proportional to the mismatch between the detection set and the environment.

Customer context — industry vertical, company size, cloud footprint, regulatory requirements, technology stack — needs to be a first-class attribute in the detection data layer. The XDR platform needs to know that a particular tenant is a law firm running Citrix so that it can activate detections for legal-sector targeting campaigns and Citrix-specific exploitation, while suppressing detections for AWS services that do not exist in that environment.

This is where detection engineering becomes a product function. The act of profiling a customer, mapping their environment to the right detection sets, and continuously adjusting as their environment changes — that is product work. It requires a data model, a curation workflow, and feedback loops. It does not fit inside a SOC shift rotation.

3 Emerging and Research-Driven

When a vulnerability researcher publishes a proof-of-concept for a new exploitation technique, the detection for it does not exist yet. When threat intelligence identifies a new campaign targeting a specific sector, the behavioral indicators may not map to any existing detection logic. These are detections that have to be authored from primary research, and they carry the highest risk of false positives because they have not been tested against diverse production environments.

Emerging detections need a CI/CD-style pipeline. Version control so that every change to the detection logic is tracked and reversible. Testing against representative environments — not just the researcher's lab, but customer environments that approximate real-world noise. Canary rollout to a subset of tenants before broad deployment. Rollback capability when a detection generates unexpected volume. Lifecycle management so that detections are retired when the underlying vulnerability is patched or the campaign ends.

The organizations that do this well treat detection rules as deployable artifacts with the same rigor that software engineering applies to code. The organizations that do it poorly push untested detections to production at 4 PM on a Friday and spend the weekend explaining to customers why their alert queue exploded.

Detection as Deployable Artifact

The CI/CD Model for Detection Rules

A detection rule is a piece of logic that runs against a data stream and produces findings. It has authors, dependencies, expected behavior, edge cases, and failure modes. It can be correct or incorrect, performant or expensive, well-tested or untested. In every meaningful sense it is software, and the industry's refusal to treat it as software is the root cause of most detection quality problems.

Git-managed means every change is attributed, reviewable, and reversible. Tested means the detection has been validated against known-good and known-bad samples, and against representative production noise to establish a baseline false-positive rate. Measured means the detection carries metrics — true positive rate, noise ratio per customer profile, time-to-triage for generated findings, and whether those findings led to confirmed incidents. Rollbackable means that when a detection misbehaves in production, it can be reverted to its previous version in minutes, not hours.

Retired means there is a deliberate end-of-life process. Detections that are superseded by better logic, that target vulnerabilities which have been universally patched, or that consistently produce noise without actionable signal — they get decommissioned. The detection inventory does not grow without bound. Every active detection has a justification for being active.

Stage Gate Criteria Owner Typical Duration
Draft Logic authored, threat hypothesis documented, ATT&CK mapping assigned Detection engineer 1–3 days
Test Validated against lab replay data; FP rate below threshold on historical logs Detection engineer + QA 2–5 days
Canary Deployed to 5–10% of tenants; noise monitored for 48–72 hours Detection ops 3–7 days
Deploy Rolled to all applicable tenants based on customer profile matching Detection ops 1–2 days
Measure TP rate, noise ratio, triage time tracked per profile; SQR contribution calculated Detection engineering lead Ongoing (30/60/90-day reviews)
Retire Superseded, patched, or consistently below SQR threshold — decommissioned with documentation Detection engineering lead As triggered

Software engineers solved this problem twenty years ago. Version control, code review, automated testing, staged rollout, observability, rollback. Detection engineering is not a novel discipline — it is software engineering applied to security logic. The tooling exists. The methodology exists. The gap is organizational: most security teams have not adopted the practices that their engineering counterparts take for granted.

The Product Function Argument

Detection Engineering Is Not a SOC Function

In single-tenant security operations, detection engineering can live inside the SOC. The analysts who triage alerts also write and tune the detections. The feedback loop is short because the people producing signal are the same people consuming it. This works at small scale.

In multi-tenant operations — managed detection and response, managed SOC, security-as-a-service — every detection multiplies across hundreds of environments. A single poorly-tuned detection running across 300 tenants does not generate one noisy alert; it generates thousands. The cost of a false positive scales linearly with the number of tenants, while the cost of writing a good detection is fixed. This asymmetry means detection quality has direct economic leverage that no other SOC function has.

The economic question is not "how many detections do we have" but "which detections deliver the most value per unit of analyst capacity they consume?" A detection that fires 200 times per month across the tenant base and leads to 3 confirmed incidents has a very different unit economics profile than a detection that fires 5 times per month and leads to 4 confirmed incidents. Both are producing roughly the same number of true positives, but one of them is consuming 40 times more analyst capacity to do it.

This is product thinking, not SOC thinking. It requires roadmaps, prioritization frameworks, customer segmentation, unit economics analysis, and lifecycle management. Detection engineering in a multi-tenant environment is a product function that happens to produce security outcomes. Treating it as a SOC function — staffed by analysts on shift rotation, measured by ticket throughput, managed as operational work — is an organizational design failure that caps detection quality at whatever the shift team can maintain between incidents.

The Flywheel

Detection Value Compounds with Client Maturity

When a new client onboards to managed detection, the first 90 days are noisy. MFA enforcement gaps generate credential-abuse alerts. Stale service accounts trigger impossible-travel detections. Legacy authentication protocols create a baseline of suspicious activity that is entirely expected given the environment's current state. In a typical onboarding, 80% of alert volume comes from known-hygiene gaps, not actual threats.

As remediation progresses — MFA gets enforced, stale accounts get disabled, legacy auth gets blocked — the noise floor drops. Detections that were producing constant false positives against hygiene gaps start surfacing actual anomalies. The same detection set becomes more valuable without any changes to the detection logic, because the environment it is monitoring has improved.

This is a retention story, not just a security story. The longer a client stays, the more value the detection program delivers per unit of effort. Each remediation cycle makes the next one cheaper because the noise floor is lower and the detections are tuned to the actual environment instead of fighting against misconfiguration. Switching to a new provider means starting the noise calibration from zero.

The compounding effect: Year one, the detection program fights noise. Year two, it finds threats. Year three, it catches things that would have been invisible twelve months earlier. The flywheel does not spin faster because the technology improved — it spins faster because the environment matured and the detection program matured with it.