Analysts spend 20 minutes per case gathering context. What if the platform assembled it the moment the alert fired?
Why SOAR Fails at Scale
SOAR platforms were designed for a single organization with stable processes. A dedicated security team writes playbooks that encode their environment’s specific logic: “if this alert fires on a server in the finance VLAN, escalate immediately; if it fires on a dev box, enrich and queue.” For one organization, this works. The team knows their environment, maintains the branching logic, and updates playbooks when infrastructure changes.
The model breaks when applied across hundreds of tenants. Every customer is a snowflake in multi-tenant MDR. One runs a hybrid Azure AD environment with legacy on-prem Exchange; another is cloud-native with Okta and Google Workspace; a third has a flat network with no segmentation and a firewall from 2018. Each if-then branch in a playbook multiplies by the number of tenant configurations. What starts as elegant automation becomes a brittle rule forest that consumes engineering hours to maintain, eats margins, and still misses edge cases.
The problem is not automation itself. The problem is that playbook-based automation encodes context at write-time instead of resolving it at run-time. A playbook written for “tenant with Defender ATP” cannot handle a tenant that switched to CrowdStrike last Tuesday. A playbook that enriches via VirusTotal assumes the hash exists in VT’s corpus. These assumptions compile silently and fail silently—the worst kind of failure in security operations, because nobody notices until a real incident goes uncontained.
Where SOAR Still Belongs
None of this means SOAR has no place. The critique above targets one specific misuse—SOAR as the investigation engine, the thing that decides whether an alert is real. That is the job it was never built for and cannot scale to across hundreds of tenants. As a thin adapter layer, it remains the right tool.
The modern pattern is XDR-core with SOAR as a narrow orchestration layer: detection and high-confidence containment happen natively in the platform, and SOAR is invoked only when an action genuinely needs it. The branch test is simple—reach for a playbook when the step requires either stitching together multiple disjoint systems, or a human approval gate. Ticketing, Slack and Teams notifications, a firewall rule on a legacy appliance, customer-specific escalation steps, evidence capture for an audit trail: this is honest adapter work, and a deterministic playbook does it well.
SOAR still earns its place where the conditions favor it—a mature SOC with its own engineering capacity, a multi-vendor environment that needs vendor-agnostic orchestration, and heavy approval or compliance workflows that must be encoded and audited. The mistake is not using SOAR. It is asking brittle if-then trees to carry the contextual judgment that belongs in the triage layer above them.
The Three-Layer Framework
Investigation automation is not a single problem. It decomposes into three distinct layers, each with different build-vs-buy economics and different failure modes. Treating them as one system is how platforms end up doing all three poorly.
Layer 1: Symbolic / Deterministic
Buy
Sigma rules, YARA signatures, known-bad IOC matching. This is commodity capability. The patterns are published, the matching engines are mature, and the detection coverage improves with community contribution. There is no competitive advantage in building a custom engine to match a SHA-256 hash against a threat intel feed—dozens of vendors do this, and the open-source alternatives are production-grade.
The discipline here is standardization, not innovation. Pick a detection format (Sigma is the obvious choice for portability), normalize telemetry into a consistent schema so rules apply cleanly, and move on. Engineering time spent on deterministic matching is engineering time not spent on the layers where differentiation actually lives.
Layer 2: Contextual Triage
Build
“Is this alert real for THIS customer?” This is the question that separates investigation automation from alert enrichment. The same Entra ID impossible-travel alert means completely different things depending on whether the tenant has traveling sales staff, a VPN concentrator that egresses in another country, or a policy that explicitly allows specific geographic regions. The alert is identical. The context makes it a true positive or noise.
This layer requires tenant context: environment size, cloud configuration, security maturity level, industry vertical, known exceptions, historical baselines. This is where the intellectual property lives. It only works if telemetry is structured as a queryable context layer—not just raw SIEM data dumped into an index, but a data model that can answer “what is normal for this tenant” in under a second.
No vendor sells this out of the box because no vendor knows the specific combination of tenant attributes that make an alert actionable in a given environment. Building the contextual triage layer is building the core product.
Layer 3: Response / Action
Build Cautiously, Human-Gated
DARPA’s AIxCC competition found that 20–40% of AI-generated patches are semantically incorrect. They compile. They pass basic tests. They introduce subtle behavioral changes that only manifest under specific conditions. In software engineering, a semantically incorrect patch means a bug report next quarter. In security operations, a semantically incorrect response means a missed containment, an unnecessary outage, or a quarantined production server that takes the business offline.
Human-in-the-loop for destructive actions is not conservatism. It is engineering discipline applied to a domain where the cost of a wrong action exceeds the cost of a delayed action. Isolating a compromised endpoint is the right call—until it is a domain controller and the entire company loses authentication. Blocking an IP is the right call—until it is a shared NAT egress point for a major customer.
The automation goal for this layer is not autonomous action. It is presenting the analyst with a pre-assembled response plan, the relevant context for each action, and a single approval gate. Reduce decision time from minutes to seconds, but keep the human in the loop for anything that changes state in a tenant environment.
What Makes This Hard
The concept is straightforward: assemble the right context for the right analyst at the right time. The implementation breaks in specific, known ways. Each of these challenges has been solved in other domains, but the combination in multi-tenant security operations creates compounding difficulty.
Tenant Context Goes Stale
A tenant onboards with a hybrid Exchange environment. Six months later, they complete a migration to Exchange Online. Nobody notifies the SOC. The investigation playbooks still reference on-prem mailbox audit logs that no longer exist, and the enrichment pipeline queries a server that has been decommissioned. The investigation proceeds with incomplete context, and the analyst either wastes time chasing ghost data or—worse—closes the case as a false positive because the expected evidence is missing.
With thousands of tenants, configuration changes without notification are not the exception; they are the norm. The solution is automated drift detection: continuously comparing the tenant model against observable reality and flagging divergence. “Tenant reality no longer matches the model” must be a first-class operational signal, not a footnote in a quarterly review. This is fundamentally an operational discipline problem, not just an engineering one.
Per-Tenant Baselining at Scale
“Anomalous” means nothing without a definition of “normal.” A user logging in from three countries in one day is suspicious for a 50-person accounting firm. It is Tuesday for a global consulting company with VPN split-tunneling. The same telemetry, interpreted against different baselines, produces opposite conclusions.
Building per-tenant baselines raises immediate questions: how much history is needed before the baseline is reliable? What happens with newly onboarded tenants that have no history? What if “normal” is already compromised—does the baseline encode the attacker’s behavior as legitimate? What about seasonal variation, M&A activity that suddenly doubles the user population, or a tenant that rotates their entire IT staff? Each of these is a known failure mode for behavioral analytics, and each requires an explicit design decision rather than a default assumption.
Enrichment Latency
Context assembly for a single investigation might hit six external APIs in parallel: threat intel for IOC reputation, directory services for user attributes, asset inventory for endpoint context, WHOIS for domain registration, geolocation for IP mapping, and historical alert data for prior incidents involving the same entity. Five return in 200 milliseconds. One takes 8 seconds because the vendor’s API is rate-limited or experiencing degradation.
The design question is not “how do we make all enrichments fast”—external dependencies have external latency. The question is which enrichments are worth blocking on versus which can arrive asynchronously after the initial triage. Pre-computation helps for stable data (asset inventory, user attributes) but wastes resources for volatile data (threat intel scores that change hourly). Getting this wrong means either slow investigations that frustrate analysts or incomplete investigations that miss critical context. The tradeoff is per-enrichment-type, not global.
Multi-Tenant Data Isolation
Tenant A’s context must never leak into Tenant B’s investigation. This sounds obvious until the implementation details surface. A shared enrichment cache that stores VT lookup results by hash—is that cross-tenant? Technically the hash is not tenant-specific, but the fact that Tenant A queried it reveals something about their telemetry. A correlation engine that identifies a campaign across tenants—can it show Tenant A that Tenant B saw the same IOC, or does that violate isolation?
At scale, isolation must be baked into the data layer from day one. Retrofit isolation onto a system that was designed for convenience, and the attack surface grows with every new feature. Row-level security, tenant-scoped API tokens, query-path auditing, and blast-radius testing for every new data flow. This is not a feature to add later. Getting this wrong is a breach, not a bug—and the SOC that was supposed to protect tenants becomes the vector.
Schema Co-Ownership
Security operations knows what analysts need: “show me every authentication event for this user in the last 72 hours, correlated with their device trust score and any concurrent VPN sessions.” Engineering knows what is queryable and performant: “we can join auth events with device inventory if both use the same user identifier format, but VPN sessions are in a different schema with a 15-minute ingestion delay.”
Without joint ownership of the data model, the result is predictable: either an analyst wish list that engineering cannot build within latency requirements, or a technically elegant data model that analysts do not use because it does not map to how investigations actually proceed. The schema must be designed by people who understand both the investigation workflow and the query performance characteristics. This means security operations and engineering reviewing the same data model together, not throwing requirements over a wall.
The Build-vs-Buy Test
One question determines the build-or-buy decision for every component in the investigation stack: does building this make the platform harder to replicate? If the answer is yes, it touches the data model, the tenant context layer, or the analyst workflow—that is intellectual property. If the answer is no, it is a commodity capability that someone else maintains better and cheaper.
Build where it creates defensible differentiation. Buy where it creates operational leverage. The trap is building commodity capabilities because “it would be easy” or buying differentiating capabilities because “it would be faster.” Easy-to-build commodity work is engineering time that does not compound. Faster-to-buy differentiation is a dependency on a vendor who controls the roadmap.
| Build (Differentiating) | Buy (Commodity) |
|---|---|
| Tenant context model | Threat intel feeds |
| Investigation orchestration | Vulnerability scanning |
| Analyst dashboard | Sandboxing / detonation |
| Enrichment pipeline | EASM scanning |
The build column has a common thread: each item touches how the platform understands and serves individual tenants. The buy column has a different common thread: each item produces the same output regardless of who the tenant is. That distinction—tenant-specific versus tenant-agnostic—is the reliable heuristic for the build-or-buy decision.
Evidence — Applied Playbooks
These playbooks demonstrate the framework in practice. Each one shows how investigation automation requirements differ based on threat type, tenant configuration, and the specific context that determines whether an alert is signal or noise.