Supply Chain Incident Response

Vendor breach, compromised package or trojanised update — by scenario, with the reason for every step

Preparing, not responding? Readiness, detection and hardening are on Supply Chain Security. This page is for when it is happening.
T+0 — T+4 h  •  CONTAIN

Respond: The First Four Hours

A vendor or MSP disclosed a breach, or a package or update you install shipped malicious code. The flow is the same either way; By Scenario lists what changes. Goal in this phase: find every path from the compromised supplier into your environment, then cut every path and revoke every credential on them in one move.
First question: did their code run inside your environment? A vendor breach puts the attacker in their environment, holding your credentials. A compromised package or update means the attacker's code already ran on your laptops, CI runners or servers — install is execution. The second is an intrusion; scope it like one.
Contain at least as wide as the attacker could be. A supplier or MSP is the multi-path case: integrations, API keys, VPN, remote-management tools and whatever their code dropped on your hosts. Cutting one path now and revoking credentials hours later tells the attacker they've been found while the other paths still work. Cut and revoke in one coordinated action. If you can't yet list every path the supplier has, go broad: disable all of that supplier's accounts, integrations, VPN and RMM access and block its traffic, rather than choosing which ones to cut. Exception: if something destructive is already under way — encryption, backup deletion, mass data movement — contain immediately and scope afterwards.
T+0–15 minVerify the notice, check for impact, declare when you find it
Reach the vendor's security team on a contact you already had — never through a link or number in the notice. For a package, take the facts from the advisory: package name, malicious versions, when they were published and pulled, and the published IOCs. Bring in security, IT, the engineering lead, legal and communications. A notice alone is intelligence, not yet your incident: declare and name one incident lead as soon as you find the bad version installed, or the vendor's access used — or straight away if the supplier held privileged access (RMM, admin accounts, code running on your hosts) that you can't rule out. Why: breach notifications are also used as phishing lures, and the publication-to-removal window is your scope — any install in it counts, even if you are on a good version now.
T+15–60 minFind every place you are connected, or it was installed
Vendor: its accounts and guests, API keys, OAuth apps, VPN access, firewall rules and the software of theirs you run. MSP: the paths that never show as a vendor account — the GDAP or DAP relationship (Microsoft 365 admin center › Settings › Partner relationships), Azure Lighthouse delegations (az managedservices assignment list per subscription, or Azure portal › Service providers), its RMM agent on your hosts, and apps it authorised in Salesforce, Google Workspace, Slack or GitHub, which Entra never lists. Package: lockfiles across all repositories, the uses: lines of every workflow, CI build logs and caches, container images and their base-image digests, OS packages on hosts and images (xz-utils was a distro package, not a lockfile entry), developer machines, production hosts. Why: you can only cut and rotate what you have listed. Lockfiles show what should be installed; CI logs and SBOMs of built images show what actually was, and when; a pinned action's own dependencies are not pinned by your SHA.
# lockfiles and workflows in a checkout (repeat per repo, or use code search across the org)
grep -rn --include=package-lock.json --include=yarn.lock --include=pnpm-lock.yaml '<package>' .
grep -rn 'uses:' .github/workflows/            # every action, with its tag or SHA
npm ls <package> --all
pip show <package>
# OS packages and built images
dpkg -l '<package>*'; rpm -q <package>
syft <image> | grep -i <package>
docker image inspect <image> --format '{{index .RepoDigests 0}}'
T+15–60 min, in parallelExport first: preserve what ages out, and what the vendor can delete
Retention runs against you from the start: Entra sign-ins 30 days on P1/P2 and 7 free, the unified audit log 180 days Standard and 1 year Premium, CloudTrail event history 90 days, GitHub Actions run logs 90 days by default, the organisation audit log 180 days (Git events 7). The vendor's earliest access date is routinely months before its notice, so widen every window to that date and export from it. For a package, preserve the tarball and its registry metadata: the registry removes malicious versions, and your cache is then the only copy. Why: a hunt you cannot re-run is a finding you cannot defend to a regulator, an insurer or a court. Send the vendor a litigation-hold letter on day one and ask for its IR firm's report; a statement that there is “no evidence of” exfiltration describes their logging, not your data.
# GitHub: audit log (Enterprise Cloud API) and the run logs of every job that installed the package
gh api --paginate '/orgs/<org>/audit-log?phrase=created:>=<window start>' > audit-log.json
gh run list -R <org>/<repo> --created '>=<window start>' --json databaseId,name,createdAt > runs.json
gh api /repos/<org>/<repo>/actions/runs/<run id>/logs > run-<run id>.zip
# npm: publish history, who published, integrity and provenance of the bad version; then the cached tarball
npm view <package> time --json > time.json
npm view <package>@<bad version> _npmUser dist.integrity dist.attestations --json > version.json
grep -rl '<package>' ~/.npm/_cacache/index-v5 | xargs -I{} cp --parents {} ./evidence/
tar -czf npm-cache-<host>.tgz ~/.npm/_cacache          # or the registry proxy's storage for that package
npm audit signatures                                   # registry signatures and provenance of what is installed now
gh attestation verify <artifact> --owner <org>         # for artifacts with GitHub provenance
# AWS: event history (90 d) and the trail
aws cloudtrail lookup-events --start-time <window start> --lookup-attributes AttributeKey=EventName,AttributeValue=AssumeRoleWithWebIdentity > sts.json
# Entra and M365: see the Export act on the Defender tab. Hash everything.
sha256sum ./evidence/* audit-log.json runs.json *.zip *.json *.tgz > manifest.sha256
T+15–90 min, in parallelLook for the access being used
Sign-in, API and firewall logs for the vendor's IP ranges, service accounts, guests, GDAP technicians and integrations; cloud audit logs for role assumptions and API calls from unknown subjects and IPs; the GitHub audit log for repositories, workflows, tokens, secrets and pushes the package's stolen token could have made; your own packages' publish history, owners and dist-tags — back to the start of the window. Run this alongside the inventory, not after the cut: every new account, token, workflow, repository or host it turns up joins the list for the next step. Why: revoking stops future use; only the logs show what was already done, and a foothold you find after the cut is one the attacker kept. The 2025 npm worms used the stolen token within minutes to create repositories, add workflows and publish trojanised versions of the victim's own packages. This is where a vendor incident becomes yours — use the hunts under Hunt & Act by Platform.
# GitHub audit log: what a stolen token did (fine-grained PATs log request_created/access_granted; classic PAT creation is not logged, only its use)
gh api --paginate '/orgs/<org>/audit-log?phrase=created:>=<window start>' | jq -r '.[] | select(.action | test("^(repo\\.create|repo\\.access|workflows\\.created_workflow_run|git\\.push|personal_access_token\\.|org\\.(create|update)_actions_secret|repo\\.(create|update)_actions_secret|environment\\.(create|update)_actions_secret|hook\\.create|integration_installation\\.create|public_key\\.create|protected_branch\\.destroy|org\\.add_outside_collaborator|repo\\.add_member)")) | [.created_at, .action, .actor, .repo // .org, .user_agent] | @tsv'
# Workflows and branches the worm added; private repositories flipped public (repo.access) and "-migration" copies
gh search repos --owner <org> --created '>=<window start>' --json name,visibility,description,createdAt
# Registry: did our packages get republished, and by whom
npm view <our package> time --json | jq 'to_entries | map(select(.value >= "<window start>"))'
npm owner ls <our package>; npm dist-tag ls <our package>
npm view <our package>@<version> _npmUser dist.attestations
T+60–120 minCut every path and revoke every credential — one action
Paths. Vendor: suspend integrations, file transfers and syncs; disable its accounts and guests, VPN and RMM access; restrict the network segments and firewall rules that let its services in. MSP: each path has its own switch and Update-MgServicePrincipal flips none of them — end the GDAP or DAP relationship (admin center › Partner relationships › Remove roles), delete every Lighthouse assignment, block the partner tenant inbound in cross-tenant access settings (guests and direct connect only; it does not govern GDAP), disable the RMM agent service fleet-wide and revoke the RMM's API keys. SaaS-to-SaaS: revoke the vendor's connected app in each platform it was authorised in; the Entra cut does not reach a Salesforce or Workspace grant. Package: block the bad versions in your registry proxy and pause the CI workflows that pull it; network-contain the hosts and runners where it ran — they get rebuilt from clean later, not cleaned in place (see Rebuild from clean below). Trojanised update: block the vendor's update host until it confirms a clean release, and block the stolen certificate, not only the hashes. Credentials, in the same pass. Vendor: API keys, OAuth tokens and app consents, service accounts and VPN accounts it holds; reset any password shared with it. A multi-tenant vendor app's secret lives on the app object in the vendor's tenant: you cannot rotate it, so send the vendor the key id from the stage 3 hunt and get written confirmation. Package: everything readable from where it ran, listed below. Revoke first, then reissue. Pins. Pin to the last known-good version only after checking it: its hash or signature matches the registry's record, it was published before the compromise window, and the advisory doesn't list it. For a GitHub Action, pin to a full commit SHA and verify that SHA against the upstream commit history, not against the tag: in the tj-actions/changed-files incident (March 2025) every version tag was moved to the malicious commit. A SHA pin covers that action only; the actions it calls internally are pinned however its author pinned them. Why: a path left open while another is cut is a warning, not a containment; most malicious packages are credential stealers, and self-spreading ones use stolen publish tokens to trojanise the victim's own packages next. The next build reinstalls whatever the pin points at, and tags can be moved by whoever controls the repository.
# Package containment: no lifecycle scripts, nothing published after the window, a cooling-off period, an allowlist
npm ci --ignore-scripts                              # or ignore-scripts=true in .npmrc for CI
npm install --before=<date before the compromise>   # resolves only versions published before that date
pnpm config set minimumReleaseAge 10080              # minutes: no version younger than 7 days (pnpm 10.16+)
# registry proxy (Artifactory, Nexus, Verdaccio): block <package>@<bad versions>, or allowlist what CI may fetch

# Rotation scope for every host and runner where it ran, in this order
# 1. cloud: the IMDS role the runner could reach, then every long-lived key on it
aws iam put-role-policy --role-name <runner role> --policy-name AWSRevokeOlderSessions --policy-document \
  '{"Version":"2012-10-17","Statement":[{"Effect":"Deny","Action":"*","Resource":"*","Condition":{"DateLessThan":{"aws:TokenIssueTime":"<now, ISO 8601>"}}}]}'
aws iam update-access-key --user-name <user> --access-key-id <key> --status Inactive   # then delete after reissue
# review every OIDC trust policy: sub/aud conditions that let a pull_request or fork token assume a deploy role
aws iam get-role --role-name <runner role> --query Role.AssumeRolePolicyDocument
# 2. platform tokens the payload could read: ~/.npmrc, ~/.pypirc, ~/.docker/config.json, ~/.kube/config,
#    ~/.aws/credentials and ~/.aws/sso/cache, ~/.azure/msal_token_cache.json, ~/.config/gcloud/, ~/.config/gh/hosts.yml,
#    mounted Kubernetes service-account tokens, Vault tokens
npm token revoke <id>; npm profile enable-2fa auth-and-writes      # npm: revoke, then 2FA for publish; move CI to trusted publishing (OIDC)
vault token revoke -accessor <accessor>; vault lease revoke -prefix <mount>/
kubectl -n <ns> delete secret <sa token secret>; kubectl -n <ns> rollout restart deployment/<runner>
# 3. GitHub: org, repository, environment and Dependabot secrets; deploy keys; the App's private key and installation tokens;
#    runner registration tokens; self-hosted runner state; and the Actions cache, which is persistence if it was poisoned
gh secret list --org <org>; gh secret list -R <org>/<repo>; gh secret list -R <org>/<repo> -e <env>; gh secret list -R <org>/<repo> --app dependabot
gh api /repos/<org>/<repo>/keys; gh api -X DELETE /repos/<org>/<repo>/keys/<id>
gh cache delete --all -R <org>/<repo>
gh api /orgs/<org>/actions/runners; gh api -X DELETE /orgs/<org>/actions/runners/<id>    # re-register from a rebuilt image
# 4. registry side, if our own packages were touched
npm dist-tag add <our package>@<good version> latest; npm deprecate <our package>@<bad version> "compromised, see advisory"
# report to [email protected] (npm) or [email protected] (PyPI) with the version list; PyPI yanks via the project page
T+2–4 hVerify the cut held
Re-run the access hunt from the moment of the cut: any sign-in, API call, token use or connection from the supplier's accounts, ranges or integrations afterwards means a path was missed. Widen the cut before moving on. Why: the containment is proven by the logs after it, not by the list of actions you took.
Hour 4 — Day 7  •  SCOPE & REBUILD

Hour 4 – Day 7: Scope, Rebuild

Map the blast radius
Every internal system, data store and application connected to the vendor or built with the package, and what data each held (personal data, financial, credentials, intellectual property). Why: this list drives the notification decision, the rebuild list and the re-connection decision.
Rebuild from clean
Collect before you rebuild: memory first on any host still running the payload (Windows: Magnet RESPONSE 1.7.2 d315c63d1ad4b89e03c7688b29169e977c2abc4559c23b9f6b2282f3c91a6c7d, run elevated from removable media or pushed with your EDR), then a persistence snapshot to diff against a clean build (Windows: Get-PersistenceSnapshot.ps1 a9534865f9e8e7b1f5077b4e1cfd8c29e7e32f6e9e49a4918ce4f8fb713e64df, read-only, -Zip, -CompareTo; Linux runners and hosts: get-persistence-snapshot.sh 8fe386a7d6ebd50225b04ad83581ee3e5d504a10a4263253b5a11b8a61eeca67, --compare-to; both cover the places an install script writes to: cron, systemd, shell rc files, SSH keys, scheduled tasks, Run keys). Then purge package and build caches, delete the Actions cache (gh cache delete --all: a poisoned cache re-infects the next run of a clean workflow), rebuild images from pinned base-image digests with the verified pins, re-image CI runners and the developer machines where the payload executed, and re-register self-hosted runners from the rebuilt image. Hosts where the payload or the vendor's tooling ran are rebuilt, not cleaned — deleting the package folder removes what you know about, not what it dropped. Why: caches and long-lived runners reinstall the bad version or keep what it dropped, and the rebuild destroys the evidence the notification decision needs.
Gate: eradication confirmed before anything reconnects
Before rebuilt hosts rejoin the network, reissued credentials go back into CI, or paused workflows restart: every known foothold is removed or rebuilt, and a verification hunt shows no activity from the supplier's paths or the payload's indicators since the eviction time. Record who confirmed it and when. Why: recovery that starts before eradication is finished hands the attacker a clean, trusted foothold — this gate covers your own systems; the vendor decision below covers theirs.
Document root cause and the evidence chain
What the vendor or advisory confirmed, what you verified yourself, and what you did when. Why: regulators, insurers and your own contract negotiation will ask for it.
Day 1 — Day 30  •  NOTIFY & DECIDE

Day 1 – Day 30: Notify, Decide, Learn

Run the notification clocks from awareness
Several clocks start at once and none of them waits for the vendor: GDPR Art. 33, 72 hours from awareness to the supervisory authority; NIS2, early warning within 24 hours, incident notification within 72, final report within a month; DORA Art. 19 for financial entities, initial notification within 4 hours of classifying the incident as major and no later than 24 hours from awareness, intermediate report within 72 hours, final report within a month; the notice window in your cyber-insurance policy, often shorter than any of these; and the customer-notification SLAs in your own contracts, which make you the notifying supplier the moment your product or data is in scope. Secrets dumped into a public repository or a public-repository workflow log are a disclosure, not a leak risk: the 2025 npm worms published credential dumps to public GitHub repositories, and public Actions logs are readable by anyone until the run is deleted. Why: the clocks run from when you became aware, not from when the vendor finishes its investigation, and the vendor's “no evidence of” is a statement about their logging that a regulator will not accept as yours.
If you shipped it, you are now the supplier
If a trojanised build reached customers, or your own packages were republished with the payload: notify customers and downstream teams, report to the registry's security team, and yank the affected versions. Why: your customers now need this same playbook, and they need your version list and window to run it.
Decide on the relationship, then update the playbook
Decide on evidence, not on the vendor's summary: ask for its IR firm's report, the earliest date of attacker access, the list of your credentials and data in scope, and written confirmation of every rotation you could not do yourself; keep the litigation hold in place until you have them. Update the vendor's risk assessment; continue with conditions or exit. Then route what the incident taught you: how the package or vendor path got in goes to posture and third-party governance, what detection missed goes to the detection backlog, and the visibility or process gaps that made it slow go to structural work. Hand the confirmed indicators and TTPs to hunting to look for earlier exploitation elsewhere. Machines collect, humans curate: an indicator becomes intelligence — a permanent rule, a block, a line in a shared feed — only after an analyst checks it and gives it a date, a source, an expiry and a confidence level. Why: re-connecting before the vendor shows the gap is closed restores the attacker's path, and an incident that ends at "systems restored" gets paid for twice.

How Far Has It Got?

Ask these at triage, top to bottom, and again at every update. The furthest stage you can show evidence for is how far the incident has got — write it in the case's status line. A stage you cannot answer yet is a scoping gap: say when it will be known. Why: the stage reached decides which scenario card applies, and whether this is still the vendor's incident or has become yours.

Vendor or MSP breach

#StageATT&CKQuestionWhere to look
1DisclosureT1195What exactly does the vendor say happened, from what earliest date, and is the notice genuine?Vendor security contact, reached on a known number, not the notice; the vendor's IR firm's report
2Our connectionT1199What access does the vendor have to us: accounts and guests, API keys, OAuth apps, GDAP or DAP, Lighthouse, VPN, RMM, software?Vendor register; Entra sign-ins by home tenant and path; Partner relationships page; Lighthouse delegations; connected apps in each SaaS; firewall rules; software inventory
3CredentialsT1528 T1098.001Which credentials or tokens tied to us are in scope, which tenant holds them, and was one added?Vendor statement; Entra audit log; service-principal sign-ins (key id, owner tenant); our secret inventory
4SoftwareT1195.002 T1553.002Did we install an affected update or version, and what else is signed with the same certificate?Software inventory; SBOM; package versions; file certificate telemetry; module loads
5ActivityT1078.004Is there anomalous activity from vendor accounts, IPs, integrations or stolen CI tokens in our environment?Sign-in logs; CloudTrail and cloud audit logs; GitHub audit log; API logs; firewall logs
6Our spreadT1219 T1195.001Did the attacker move from the vendor's access, its RMM agent or the package's install script into our systems?EDR; RMM command history; lateral-movement hunts from vendor-connected hosts and runners
7Data heldWhich of our data does the vendor hold, and whose is it?Contract and DPA; data map
8Data exposedT1567Is our data confirmed in the stolen set, or dumped somewhere public?Vendor confirmation; Have I Been Pwned; leak-site monitoring; public repositories and workflow logs

Compromised package or software update

#StageATT&CKQuestionWhere to look
1ExposureT1195.001Is a bad version in any of our lockfiles, workflows, images, base images or OS packages?Lockfiles across repos; uses: lines; SBOMs of built images (Syft); image digests; dpkg -l / rpm -q; registry proxy logs
2InstalledT1195.001Was it actually installed — where and when: developer machines, CI runners, production?CI build logs; package-manager caches (~/.npm/_cacache); EDR file events; Velociraptor FileFinder
3ExecutedT1059.007Did its install script or code run?CI job logs; EDR process tree (node spawning cmd.exe /d /s /c or sh -c, trufflehog, gh, AI CLIs); node's DNS and connections
4Secrets reachableT1552.001 T1552.005Which secrets could it read where it ran?CI secret and variable inventory (org, repo, environment, Dependabot); IMDS roles and OIDC tokens on runners; mounted service-account and Vault tokens; ~/.npmrc, ~/.pypirc, ~/.docker/config.json, cloud CLI caches, .env and SSH keys on developer machines
5ExfiltrationT1567.001Did it send data out?Egress logs from runners and developer machines; the advisory's domains and IPs; new public repositories and workflow logs in our org
6Credential useT1550.001Have the stolen tokens been used?CloudTrail (AssumeRoleWithWebIdentity subjects, leaked key ids); GitHub audit log; registry publish history
7PropagationT1195.001Were our own packages, repositories, workflows or runners modified?Registry publish history, owners and dist-tags; GitHub audit log (repo.create, repo.access, workflows.*, git.push, secrets, runners); Actions cache
8ProductionT1195.002Did the compromised code reach production builds?Deploy history; image digests; release records; provenance attestations
9DownstreamT1195.002Did we ship it to customers or other teams?Release records; customer deliverables; contractual notification SLAs

Hunt & Act by Platform

What to look for in your EDR or SIEM for each stage of the vendor-breach questions, and the response actions in each tool. Pick your platform.
Test these before you need them. The queries follow each vendor's documented schema, but field names depend on your data sources, versions and ingestion. Run each one in your own tenant during peacetime and fix it there — not during an incident.
KQL. Identity hunts run in Sentinel on SigninLogs, AADNonInteractiveUserSignInLogs, AADServicePrincipalSignInLogs, AuditLogs and AzureActivity (Entra ID and subscription diagnostic settings); AWSCloudTrail needs the AWS connector. Endpoint hunts run in Defender XDR advanced hunting (30 d). The Entra portal itself keeps sign-ins 30 d on P1/P2 and 7 d free: Sentinel retention is the real window.

Hunt

Stage 2 · Our connectionEvery external tenant with a path into ours: guests, direct connect, GDAP and DAP

Why: The vendor register is usually incomplete. Sign-ins carry the home tenant and the path used: serviceProvider is an MSP acting through a Partner Center GDAP or DAP relationship, b2bCollaboration a guest, b2bDirectConnect a Teams shared channel. Each path has a different cut, so the path matters as much as the account.
// Which external tenants act in ours, and by which path (serviceProvider = GDAP/DAP)
union SigninLogs, AADNonInteractiveUserSignInLogs
| where TimeGenerated > ago(90d)
| where CrossTenantAccessType != "none" and HomeTenantId != AADTenantId
| summarize SignIns=count(), LastSeen=max(TimeGenerated), Users=dcount(UserPrincipalName),
            Apps=make_set(AppDisplayName, 20) by HomeTenantId, HomeTenantName, CrossTenantAccessType
// The named vendor: every account, what it reached, from where
union SigninLogs, AADNonInteractiveUserSignInLogs
| where TimeGenerated > ago(90d) and HomeTenantId == "<vendor tenant id>"
| summarize LastSeen=max(TimeGenerated), Apps=make_set(AppDisplayName), IPs=make_set(IPAddress, 20)
    by UserPrincipalName, UserType, CrossTenantAccessType
T1199 T1078.004 Tuning: Your own MSP, Microsoft support (microsoftSupport) and every Teams shared-channel partner appear here legitimately. The first query is the inventory; the vendor under investigation is the filter, and any serviceProvider tenant you cannot name is its own finding.

Stage 2 · Our connectionVendor-owned apps, their permissions, and Azure Lighthouse delegations

Why: A multi-tenant vendor app has its app object, and its secret, in the vendor's tenant; AppOwnerTenantId tells you which apps those are and therefore which credentials you cannot rotate yourself. Lighthouse delegations give the MSP ARM access to subscriptions without any sign-in to your tenant, so the sign-in logs never show them.

Non-Microsoft SaaS-to-SaaS grants (a vendor app authorised in Salesforce, Google Workspace, Slack or GitHub) do not appear in Entra at all. List them in each platform's connected-apps page: the Salesloft Drift tokens used against Salesforce in August 2025 were invisible to an Entra-only inventory.

// Service principals the vendor owns (app object, and its secret, live in their tenant)
AADServicePrincipalSignInLogs
| where TimeGenerated > ago(90d)
| where AppOwnerTenantId == "<vendor tenant id>" or AppId == "<vendor app id>"
| summarize LastSeen=max(TimeGenerated), Resources=make_set(ResourceDisplayName), IPs=make_set(IPAddress, 20)
    by ServicePrincipalName, AppId, AppOwnerTenantId
// App governance: what those apps are allowed to do, and whether an admin consented
OAuthAppInfo
| where AppOwnerTenantId == "<vendor tenant id>" or AppId == "<vendor app id>"
| project AppName, AppId, PrivilegeLevel, Permissions, IsAdminConsented, AssignedRoles, LastUsedTime
// Azure Lighthouse: delegations written in the window. The current list is
// `az managedservices assignment list` per subscription, or Azure portal > Service providers.
AzureActivity
| where TimeGenerated > ago(90d)
| where OperationNameValue =~ "Microsoft.ManagedServices/registrationAssignments/write"
| project TimeGenerated, Caller, CallerIpAddress, ResourceId, ActivityStatusValue

Stage 3 · CredentialsWhich vendor-app credential is in use, was one added, and whose tenant holds it

Why: Rotating the wrong secret leaves the attacker in. A credential, consent or role assignment added during the exposure window is the attacker's own foothold and is removed, not rotated. The key id the app signs in with identifies the exact secret to rotate; if the app's owner tenant is the vendor's, that secret is on their app object and only they can rotate it.
AuditLogs
| where TimeGenerated > ago(30d)
| where OperationName in ("Add service principal credentials", "Consent to application",
                          "Add app role assignment to service principal", "Add delegated permission grant")
     or OperationName has "Certificates and secrets management"
| where tostring(TargetResources) has "<vendor app name>"
| project TimeGenerated, OperationName, Result, InitiatedBy, TargetResources
// Anything the vendor's own identities changed in our directory (GDAP technicians, guests, the app)
AuditLogs
| where TimeGenerated > ago(30d)
| where tostring(InitiatedBy) has "<vendor upn domain>" or tostring(InitiatedBy) has "<vendor app id>"
| summarize Count=count(), Examples=make_set(tostring(TargetResources), 10) by OperationName, Category
// Which key the app actually signs in with: KeyId on OUR service principal only if
// AppOwnerTenantId is our tenant; otherwise it is the vendor's app-object key (send them the id)
AADServicePrincipalSignInLogs
| where TimeGenerated > ago(30d) and AppId == "<vendor app id>"
| summarize FirstSeen=min(TimeGenerated), LastSeen=max(TimeGenerated), IPs=make_set(IPAddress)
    by AppOwnerTenantId, ServicePrincipalCredentialKeyId, ServicePrincipalCredentialThumbprint
T1098.001 T1528 Tuning: A vendor's scheduled credential rollover fires 'Add service principal credentials' legitimately. The tell is a second credential added without the old one removed, a consent granted by a user rather than an admin, or a new key first used from an address the vendor does not own (stage 5).

Stage 4 · SoftwareWhere the package, the vendor binary or the sideloaded payload exists, ran or was loaded

Why: Every host with the bad version is in scope, including CI runners and developer laptops no asset register lists. A trojanised installer often carries the payload as a sideloaded library (3CX, 2023: ffmpeg.dll loaded by the signed desktop app), so the installer hash alone misses the hosts that matter.
// The trojanised binary by hash: written, executed, or loaded as a module
union DeviceFileEvents, DeviceProcessEvents, DeviceImageLoadEvents
| where Timestamp > ago(30d)
| where SHA256 in ("<sha256 of installer>", "<sha256 of payload dll>")
| summarize FirstSeen=min(Timestamp), LastSeen=max(Timestamp), Types=make_set(ActionType) by DeviceName, FolderPath, FileName
// The sideloaded payload by name and parent, hash-agnostic (catches variants the advisory missed)
DeviceImageLoadEvents
| where Timestamp > ago(30d)
| where FileName =~ "<sideloaded dll>" and InitiatingProcessFileName =~ "<vendor exe>"
| summarize Hosts=dcount(DeviceName), Hashes=make_set(SHA256) by FolderPath
// The npm package on disk: laptops and runner workspaces (PyPI: FolderPath has "site-packages")
DeviceFileEvents
| where Timestamp > ago(30d)
| where FolderPath has "node_modules" and FolderPath has "<package>" and FileName =~ "package.json"
| summarize FirstSeen=min(Timestamp) by DeviceName, FolderPath, InitiatingProcessFileName
// Vendor desktop and server software by version. DeviceTvmSoftwareInventory indexes installed
// programs only: it does not see npm, pip, container or base-image contents.
DeviceTvmSoftwareInventory
| where SoftwareVendor has "<vendor>" | project DeviceName, SoftwareName, SoftwareVersion
T1195.002 T1195.001 T1574.001 Tuning: DeviceFileEvents samples writes: an empty result is not absence. Containers on Linux runners are only seen when Defender runs on the host and the workspace is on the host filesystem; otherwise the CI job logs and the registry proxy are the source for stage 4.

Stage 4 · SoftwareEverything signed with the vendor's stolen certificate

Why: A stolen signing key makes a hash block insufficient: the attacker can sign a new payload. Hunt by signer and serial number across executions and module loads, then block the certificate, not just the files.
// Signer, serial and countersignature (timestamp) time come from DeviceFileCertificateInfo, joined on SHA1
DeviceFileCertificateInfo
| where Timestamp > ago(30d)
| where Signer has "<vendor signer name>" or CertificateSerialNumber =~ "<serial from the advisory>"
| join kind=inner (
    union DeviceProcessEvents, DeviceImageLoadEvents
    | where Timestamp > ago(30d)
    | project DeviceName, SHA1, FileName, FolderPath, Timestamp) on SHA1
| summarize Hosts=dcount(DeviceName), FirstSeen=min(Timestamp1), Files=make_set(FileName, 20)
    by Signer, CertificateSerialNumber, CertificateCountersignatureTime, IsTrusted
// IsTrusted stays true until the CA revokes and the device fetches the CRL. A countersignature
// dated before the revocation keeps the signature valid even then: trust the serial and the
// countersignature time inside the compromise window, not the trust flag.
T1553.002 T1588.003 Tuning: Every legitimate binary from the vendor matches the signer name. The discriminators are the serial of the leaked certificate (the vendor names it), a CertificateCountersignatureTime inside the compromise window, and files the vendor's release manifest does not list.

Stage 5 · ActivityVendor accounts, apps and MSP technicians signing in from outside the vendor's ranges

Why: A vendor identity used from an address the vendor does not own is the attacker using the stolen access, and moves the incident from theirs to ours. GDAP sessions come from the MSP's technicians' own egress, so ask the MSP for that list before reading the result.
union SigninLogs, AADNonInteractiveUserSignInLogs
| where TimeGenerated > ago(30d)
| where HomeTenantId == "<vendor tenant id>"
| where not(ipv4_is_in_any_range(IPAddress, "<vendor cidr>"))
| summarize SignIns=count(), Apps=make_set(AppDisplayName), FirstSeen=min(TimeGenerated)
    by UserPrincipalName, CrossTenantAccessType, IPAddress, ResultType
AADServicePrincipalSignInLogs
| where TimeGenerated > ago(30d) and AppId == "<vendor app id>"
| where not(ipv4_is_in_any_range(IPAddress, "<vendor cidr>"))
| summarize SignIns=count(), FirstSeen=min(TimeGenerated)
    by ServicePrincipalName, IPAddress, ResourceDisplayName, ServicePrincipalCredentialKeyId
T1078.004 T1199 Tuning: Vendors run from cloud egress that changes without notice, and technicians work from home. Treat a new ASN with a known credential key id as a question for the vendor; a new key id, a new user agent and a new ASN together as the attacker.

Stage 5 · ActivityConnections and logons into our hosts from the vendor's ranges

Why: These are the hosts the vendor (or whoever holds its access) actually reaches; they are where to look for spread and what to cut off first.
DeviceNetworkEvents
| where Timestamp > ago(30d)
| where ipv4_is_in_any_range(RemoteIP, "<vendor cidr>")
| summarize Connections=count(), LastSeen=max(Timestamp) by DeviceName, LocalPort, ActionType, InitiatingProcessFileName
DeviceLogonEvents
| where Timestamp > ago(30d) and ipv4_is_in_any_range(RemoteIP, "<vendor cidr>")
| summarize Logons=count() by DeviceName, AccountName, LogonType, ActionType

Stage 5 · ActivityStolen CI credentials used in the cloud: OIDC role assumption from an unexpected subject

Why: A package that ran on a runner had the runner's OIDC token and any long-lived keys in its environment. In CloudTrail, UserIdentityUserName on AssumeRoleWithWebIdentity is the token's sub claim (GitHub: repo:org/repo:ref:refs/heads/main); a subject or source address you cannot map to a workflow is the stolen token in use. CloudTrail event history keeps 90 d; a trail to S3 keeps what you configured.
AWSCloudTrail
| where TimeGenerated > ago(30d)
| where EventName == "AssumeRoleWithWebIdentity"
| extend RoleArn = tostring(parse_json(RequestParameters).roleArn)
| where UserIdentityUserName !startswith "repo:<org>/"
     or UserIdentityUserName has "pull_request"
     or not(ipv4_is_in_any_range(SourceIpAddress, dynamic(["<runner egress cidr>"])))
| summarize Count=count(), FirstSeen=min(TimeGenerated), IPs=make_set(SourceIpAddress, 20) by UserIdentityUserName, RoleArn
// Long-lived keys that lived on runners or laptops: everything they did from a new address
AWSCloudTrail
| where TimeGenerated > ago(30d) and UserIdentityAccessKeyId in ("<leaked access key ids>")
| summarize Calls=count(), Events=make_set(EventName, 30) by UserIdentityAccessKeyId, SourceIpAddress, UserAgent
T1550.001 T1078.004 Tuning: A subject you do not recognise is usually a renamed repository or a new workflow. A pull_request subject assuming a deploy role, or GitHub-hosted runner traffic from outside GitHub's published actions ranges (api.github.com/meta), is not.

Stage 6 · Our spreadShells and commands launched by the MSP's remote-management agent

Why: An MSP's RMM agent is command-and-control as SYSTEM on every managed host; in the Kaseya VSA incident (2021) the attacker pushed a 'hot-fix' procedure through the agent to deploy ransomware. Commands the agent launched during the window show whether the attacker used it inside our network.
DeviceProcessEvents
| where Timestamp > ago(30d)
// Kaseya AgentMon.exe, ScreenConnect.ClientService.exe, AteraAgent.exe, NinjaRMMAgent.exe, CagService.exe (Datto), BASupSrvc.exe (N-able)
| where InitiatingProcessFileName in~ ("<rmm agent exe>")
| where FileName in~ ("cmd.exe", "powershell.exe", "pwsh.exe", "rundll32.exe", "certutil.exe", "msiexec.exe", "sh", "bash")
| extend Suspicious = ProcessCommandLine has_any ("-EncodedCommand", "-enc ", "certutil", "-decode", "DisableRealtimeMonitoring", "Set-MpPreference", "vssadmin", "bcdedit")
| summarize Hosts=dcount(DeviceName), Examples=make_set(ProcessCommandLine, 25), FirstSeen=min(Timestamp)
    by InitiatingProcessFileName, FileName, AccountName, Suspicious
T1219 T1059.003 T1059.001 Tuning: Every legitimate RMM task (patching, inventory, scripted fixes) spawns cmd and powershell as SYSTEM. Baseline the week before the window and diff. Encoded commands, certutil decoding, Defender tampering and shells on hosts with no open ticket are the signal.

Stage 6 · Our spreadInstall-time execution: lifecycle scripts, in-process credential theft and the 2025 worm pattern

Why: npm, pnpm and yarn run lifecycle scripts through a shell: cmd.exe /d /s /c "<script>" on Windows, sh -c on Linux and macOS, with node as the parent. In-process payloads never spawn curl: Shai-Hulud (September and November 2025) ran inside node, fetched trufflehog, read ~/.npmrc and cloud credential files, talked to api.github.com, created repositories and workflows and published with the stolen npm token. Nx s1ngularity (August 2025) used installed claude, gemini and q CLIs as the secret finder. Hunt node's children, its files and its connections, not a download tool.
// 1. Lifecycle scripts in the window: every package with an install script matches, so filter to the package
DeviceProcessEvents
| where Timestamp > ago(30d)
| where InitiatingProcessFileName in~ ("node.exe", "node")
| where (FileName =~ "cmd.exe" and ProcessCommandLine has "/d /s /c")
     or (FileName in~ ("sh", "bash", "dash") and ProcessCommandLine has "-c")
| where ProcessCommandLine has_any ("<package>", "preinstall", "postinstall")
| summarize Examples=make_set(ProcessCommandLine, 25), FirstSeen=min(Timestamp) by DeviceName, AccountName
// 2. Node or bun spawning the tools the 2025 worms used
DeviceProcessEvents
| where Timestamp > ago(30d)
| where InitiatingProcessFileName in~ ("node.exe", "node", "bun", "bun.exe")
| where FileName in~ ("trufflehog", "trufflehog.exe", "gh", "gh.exe", "claude", "gemini", "q", "q.exe")
     or ProcessCommandLine has_any ("--dangerously-skip-permissions", "--yolo", "--trust-all-tools", "npm publish", "gh repo create", "gh auth", "/tmp/inventory.txt")
| project Timestamp, DeviceName, AccountName, FileName, ProcessCommandLine, InitiatingProcessCommandLine
// 3. Node, bun or trufflehog talking to GitHub, a webhook sink, the registry or the metadata service
DeviceNetworkEvents
| where Timestamp > ago(30d)
| where InitiatingProcessFileName in~ ("node.exe", "node", "bun", "bun.exe", "trufflehog", "trufflehog.exe")
| where RemoteUrl has_any ("api.github.com", "webhook.site", "registry.npmjs.org") or RemoteIP in ("169.254.169.254")
| summarize Count=count(), FirstSeen=min(Timestamp) by DeviceName, InitiatingProcessFileName, RemoteUrl, RemoteIP
// 4. Files the worms wrote (Shai-Hulud: bundle.js, setup_bun.js, bun_environment.js, the *.json dumps, migrate-repos.sh)
DeviceFileEvents
| where Timestamp > ago(30d)
| where FileName in~ ("setup_bun.js", "bun_environment.js", "cloud.json", "contents.json", "environment.json", "truffleSecrets.json", "actionsSecrets.json", "migrate-repos.sh", "shai-hulud-workflow.yml")
     or (FileName =~ "bundle.js" and FolderPath has "node_modules" and FileSize > 2000000)
| summarize Hosts=dcount(DeviceName), Paths=make_set(FolderPath, 20) by FileName
T1195.001 T1059.007 T1552.001 T1567.001 Tuning: Query 1 fires on node-gyp builds, husky, esbuild and every other legitimate install script; it is a scoping query, read the command lines. Query 2: developers run gh and AI CLIs from editors that are node processes (VS Code extensions); the parent command line distinguishes an editor from an npm install. Query 3: node talking to api.github.com is normal on a laptop with gh or Dependabot tooling, abnormal on a CI runner during an install step.
Stage 1 · DisclosureNo hunt in this tool. No telemetry answers this. The evidence is the vendor's statement and, later, their IR firm's report: ask for the earliest date of attacker access, the systems and data types, and set the hunt windows above from that date, not from the notice. Verify the notice through a contact you already had, and run the notice's links and attachments through Defender for Office 365 Explorer rather than opening them.
Stage 7 · Data heldNo hunt in this tool. Not a telemetry question. The contract, the DPA and your data map say which of your data the vendor holds; the integration's own export logs (what you sent them) are the technical cross-check.
Stage 8 · Data exposedNo hunt in this tool. No tool here sees the vendor's stolen set. The sources are the vendor's written confirmation, your leak-site monitoring and a Have I Been Pwned domain search. Read 'no evidence of exfiltration' as a statement about the vendor's logging coverage, not about the data.

Act

Broad first, surgical once the scope is known. Export before you cut; collect before you stop anything; compromised hosts and runners are rebuilt, not cleaned.

T+15–60 · ExportExport the identity evidence before it ages out

Why: the Entra portal keeps sign-ins 30 d on P1/P2 and 7 d free, the unified audit log 180 d on Standard and 1 y on Premium, Defender XDR device data 30 d; a vendor's 'earliest access date' is routinely months earlier than its notice, so the export is the only copy you will have.

Export to a case folder with a hash manifest. Ask the vendor's IR firm for its report and send the vendor a litigation-hold letter on day one so their logs are preserved too.

# Entra sign-ins and directory audit (Microsoft Graph PowerShell)
Get-MgAuditLogSignIn -Filter "createdDateTime ge <window start>T00:00:00Z" -All |
  Where-Object { $_.HomeTenantId -eq '<vendor tenant id>' -or $_.AppId -eq '<vendor app id>' } |
  Export-Csv vendor-signins.csv -NoTypeInformation
Get-MgAuditLogDirectoryAudit -Filter "activityDateTime ge <window start>T00:00:00Z" -All | Export-Csv directory-audit.csv -NoTypeInformation
# Unified audit log (Exchange Online PowerShell): what the vendor's accounts and app touched in M365
Search-UnifiedAuditLog -StartDate <window start> -EndDate (Get-Date) -UserIds <vendor accounts> -ResultSize 5000 -SessionCommand ReturnLargeSet | Export-Csv ual-vendor.csv -NoTypeInformation
# Device timeline for every host in scope (Defender XDR API, 30 d)
GET https://api.security.microsoft.com/api/machines/<machine id>/alerts
Get-ChildItem <case folder> -File | Get-FileHash -Algorithm SHA256 | Export-Csv manifest.csv -NoTypeInformation

T+60–120 · CutCut the vendor's cloud access — each path with its own primitive

Why: Update-MgServicePrincipal cuts the app and nothing else. An MSP reaches you through a GDAP or DAP relationship, Lighthouse delegations and guests, each with its own switch; a path left open while another is cut tells the attacker they have been found.

GDAP and DAP: Microsoft 365 admin center › Settings › Partner relationships › select the partner › Remove roles. That ends the relationship from your side; the partner sees it terminated in Partner Center. Cross-tenant access settings govern guests and direct connect only — they do not apply to GDAP sign-ins, which is why the relationship itself has to go. Non-Microsoft OAuth (a vendor app authorised in Salesforce, Google Workspace, Slack, GitHub): revoke it in that platform — Salesforce Setup › Connected Apps OAuth Usage › Block — the Entra cut does not reach it. Ban the vendor's app in Defender for Cloud Apps › OAuth apps as well, so re-consent is blocked.

# Vendor service principal: disable, then strip what it was granted. Issued access tokens stay valid up to ~1 h.
Update-MgServicePrincipal -ServicePrincipalId <sp object id> -AccountEnabled:$false
Get-MgServicePrincipalOauth2PermissionGrant -ServicePrincipalId <sp object id> |
  ForEach-Object { Remove-MgOauth2PermissionGrant -OAuth2PermissionGrantId $_.Id }
Get-MgServicePrincipalAppRoleAssignment -ServicePrincipalId <sp object id> |
  ForEach-Object { Remove-MgServicePrincipalAppRoleAssignment -ServicePrincipalId <sp object id> -AppRoleAssignmentId $_.Id }
# Credentials: only when AppOwnerTenantId (stage 2) is OUR tenant does the secret live on an object we control.
# A multi-tenant vendor app's secret is on the app object in the vendor's tenant: you cannot rotate it.
# Send the vendor the ServicePrincipalCredentialKeyId from stage 3 and get written confirmation it was rotated.
Remove-MgServicePrincipalPassword -ServicePrincipalId <sp object id> -KeyId <key id>    # our-tenant app only
# Guests: disable first, then revoke (revoke alone leaves access tokens valid until they expire)
Update-MgUser -UserId <guest object id> -AccountEnabled:$false
Revoke-MgUserSignInSession -UserId <guest object id>
# Cross-tenant access settings: block the partner tenant inbound (B2B collaboration; not GDAP)
New-MgPolicyCrossTenantAccessPolicyPartner -TenantId <vendor tenant id> -B2BCollaborationInbound @{
  usersAndGroups = @{ accessType = "blocked"; targets = @(@{ target = "AllUsers"; targetType = "user" }) }
  applications   = @{ accessType = "blocked"; targets = @(@{ target = "AllApplications"; targetType = "application" }) } }
# Azure Lighthouse: per delegated subscription
az account set -s <subscription id>
az managedservices assignment list
az managedservices assignment delete --assignment <assignment id>

T+60–120 · CutBlock the payload, the stolen certificate and the update channel fleet-wide

Why: a tenant-wide indicator stops it on every device, including the ones your hunt hasn't found yet. Block the sideloaded payload's hash, not only the installer's; block the certificate so a re-signed variant is stopped too; and block the vendor's update host until the vendor confirms a clean release, because the update channel is the delivery path.

Settings › Endpoints › Indicators: File hashes, Certificates (upload the .CER or .PEM, or the thumbprint via the API) and URLs/Domains. IP indicators take single IPs only: block the vendor's ranges at the firewall or VPN instead. Server-side update compromises (the vendor's build or update server) also need the vendor's service accounts rotated on their side; ask for it in writing.

POST https://api.security.microsoft.com/api/indicators    # one call per indicator
{"indicatorValue": "<sha256 of sideloaded dll>", "indicatorType": "FileSha256", "action": "BlockAndRemediate",
 "title": "IR <case>: <vendor> payload", "severity": "High", "description": "Supply-chain incident <case>"}
{"indicatorValue": "<certificate thumbprint>", "indicatorType": "CertificateThumbprint", "action": "Block",
 "title": "IR <case>: <vendor> stolen signing certificate", "description": "Supply-chain incident <case>"}
{"indicatorValue": "<update host>", "indicatorType": "DomainName", "action": "Block",
 "title": "IR <case>: <vendor> update channel until a clean release is confirmed", "description": "Supply-chain incident <case>"}

T+60–120 · CutIsolate every device where the package or vendor tooling ran — in one pass

Why: isolating one device at a time tells the attacker they're found while the others still work; the device list comes from the install, signing and access hunts above.

Live Response keeps working on an isolated device, so collection continues.

$ids | ForEach-Object {
  Invoke-RestMethod -Method Post -Headers $h -ContentType 'application/json' `
    -Uri "https://api.security.microsoft.com/api/machines/$_/isolate" `
    -Body '{"Comment": "IR <case>: supply chain", "IsolationType": "Full"}'
}

T+60–120 · CollectMemory and persistence first, then the file

Why: a host that is still running the payload holds its unpacked code, its C2 and the stolen tokens in memory; stopping it first destroys that, and the rebuild destroys everything else. Collect, then stop and quarantine, then rebuild.

Upload the tools to the Live Response library once (Settings › Endpoints › Response › Library). Windows memory: Magnet RESPONSE 1.7.2 d315c63d1ad4b89e03c7688b29169e977c2abc4559c23b9f6b2282f3c91a6c7d. Windows persistence: Get-PersistenceSnapshot.ps1 a9534865f9e8e7b1f5077b4e1cfd8c29e7e32f6e9e49a4918ce4f8fb713e64df (read-only; -Zip for chain of custody, -CompareTo against a clean build's snapshot). Linux runners and hosts: get-persistence-snapshot.sh 8fe386a7d6ebd50225b04ad83581ee3e5d504a10a4263253b5a11b8a61eeca67 (cron, systemd, SSH keys, PAM, ld.so.preload, SUID, package verification, --compare-to). Then Device page › Collect investigation package, and File page › Stop and quarantine file (up to 1000 devices).

# Live Response on the isolated device
putfile MagnetRESPONSEv172_Self_Extracting_Archive.exe
run MagnetRESPONSEv172_Self_Extracting_Archive.exe      # elevated; RAM, pagefile, running processes, triage files
putfile Get-PersistenceSnapshot.ps1
run Get-PersistenceSnapshot.ps1 -parameters "-Zip"
getfile "<snapshot zip path from the script output>"
getfile "<path to package folder, vendor binary or sideloaded dll>"
# Linux host or runner (Live Response on Linux, or your SSH): run read-only, hash, then pull
run get-persistence-snapshot.sh

After the gate · ScopedRelease a device only after it is rebuilt and the gate is met

Why: releasing is the one surgical, per-device action here — each device comes back once it is rebuilt and the eradication gate is confirmed.

Device page › Release from isolation.

POST https://api.security.microsoft.com/api/machines/{id}/unisolate
{"Comment": "IR <case>: rebuilt, eradication gate confirmed"}

By Scenario: What Changes

The first-four-hours flow above is the universal path. These are the deltas — what you do differently, and first, when the incident is shaped like this. Each has its own TheHive case template.

1. Vendor or MSP Breach

Indicators: a breach notice from a vendor, SaaS provider or MSP; press or leak-site reports naming them; anomalous activity from their accounts or IP ranges.

Variation: The attacker is in the vendor's environment, not yours — yet. Verify the notice first, then cut data flows and revoke every credential the vendor holds; the question is whether their access into you was used, and the answer is in your logs from the vendor's earliest access date, not from the notice. An MSP is a different incident: its RMM agent is command-and-control as SYSTEM on every host it manages, and in the Kaseya VSA incident (July 2021) the attacker pushed ransomware to downstream organisations (fewer than 1,500 by Kaseya's own count) as a “hot-fix” procedure through the agent. Disable the agent service fleet-wide, revoke the RMM's API keys, end the GDAP or DAP relationship and the Lighthouse delegations, then review every script, policy and procedure the RMM pushed in the window — those are the attacker's actions, not a patch run. The multi-tenant vendor app's secret lives in the vendor's tenant: you cannot rotate it, so get written confirmation that they did. A vendor app authorised in Salesforce, Google Workspace or Slack is cut in that platform; the Salesloft Drift tokens used against Salesforce tenants (August 2025) were untouched by any Entra action. Template: supply-chain-response.

2. Compromised Package or GitHub Action

Indicators: an advisory for an npm or PyPI package, GitHub Action, OS package or base image you depend on; node on a runner spawning trufflehog, gh or an AI CLI, or talking to api.github.com during an install; new repositories, workflows or runners in your org; unexpected publishes of your own packages. Examples: Shai-Hulud npm worm (September and November 2025), Nx s1ngularity (August 2025), tj-actions/changed-files (March 2025), xz-utils (2024).

Variation: Install is execution: a bad version that only reached a CI runner or a laptop still had that machine's secrets. Scope by “where was it installed?”, not “did it reach production?”. The 2025 worms ran inside node with no curl to catch: they harvested ~/.npmrc, cloud and GitHub credentials, published them to public repositories and workflow logs, and used the npm token to trojanise every package the victim maintained — so stage 7 (propagation) is minutes behind stage 3, and your own packages' publish history is a first-hour check, not a day-two one. xz-utils was a distro package carried into base images, not a lockfile entry: pin base images by digest and inventory OS packages in images with Syft, not only application dependencies. Do not pin Actions to tags (every tj-actions tag was moved), do not rebuild on the same caches or runners, and install with --ignore-scripts and a cooling-off period until the registry proxy allowlist is in place. Template: supply-chain-package-compromise.

3. Trojanised Vendor Update

Indicators: a vendor advisory that a signed release was malicious; EDR detections on a trusted application; the application reaching unknown domains. Example: the 3CX desktop app (2023).

Variation: Treat it as an endpoint compromise, not a patching task: the update ran with the application's privileges on every host that installed it. The signature is the dimension that changes the hunt: the installer is signed with the vendor's real certificate, so a hash block covers one file while the attacker holds the key. Hunt by signer and serial (DeviceFileCertificateInfo, Sysmon 7, code_signature.subject_name) for everything signed with that certificate, block the certificate itself, and read the trust flag with care — it stays valid until the CA revokes and the device fetches the CRL, and a countersignature dated before the revocation keeps the file valid afterwards. In the 3CX case the signed desktop app sideloaded a trojanised ffmpeg.dll: the main executable's hash never identified the payload, the module load did. Block the vendor's update host until it confirms a clean release, because the update channel is the delivery path; if the compromise was on the vendor's build or update server rather than a developer's key, its service accounts are what need rotating, on their side. Isolate the hosts, collect memory and a persistence snapshot, then uninstall or roll back — the pre-encryption situation. Template: supply-chain-package-compromise.