Cloud Incident Response

The first 72 hours in AWS, Azure and GCP — control plane first, with the reason for every step

Preparing, not responding? Guardrails, logging and the IR-ready account structure are on Cloud Security. Entra ID or Microsoft 365 account takeover? Start with Identity Breach Response. This page is for the cloud control plane — accounts, subscriptions, projects, keys, roles and workloads — when it is happening.
Keeps the timed steps and triage table; hides provider queries and scenarios.
T+0 — T+1 h  •  CONTAIN

Respond: The First Hour

The flow is the same for every cloud incident; By Scenario lists what changes when the incident is shaped differently. Goal in this phase: keep the logs running, take the attacker's credentials away — all of them — and preserve the evidence before anything is deleted.
Contain in the control plane, not the workload. In the cloud the attacker is usually an identity making API calls, not malware on a host. Stopping an instance does nothing to an access key being used from someone else's laptop. Take away the credential and what it can do; isolate workloads only where a workload is the foothold.
Contain at least as wide as the attacker could be. A cloud attacker with write access creates new users, keys and roles within minutes of getting in. Disabling the one key you found, while the three they made stay live, tells them they're found. Until the scope is known, go broad: every credential of the compromised principal, every principal it could have created or assumed, and — if you can't bound it — the whole account, subscription or project. Go surgical once the scope is known.
T+0–10 minConfirm the logs are still running
Check that the organisation trail is logging, GuardDuty / Defender for Cloud / Security Command Center are enabled, and the log sinks and diagnostic settings are intact. Turn back on anything that is off, and write down when it went off. Why: stopping the trail, deleting the detector and removing log sinks are among the first things a cloud attacker does. Every minute without logs is a minute you can't scope later.
aws cloudtrail get-trail-status --name <org-trail-arn>      # IsLogging must be true
aws cloudtrail get-event-selectors --trail-name <org-trail-arn>   # are S3 data events on? decides stages 8-9
aws guardduty list-detectors --region <region>               # empty = detector deleted
az monitor diagnostic-settings subscription list --subscription <sub-id>
gcloud logging sinks list --organization=<org-id>
gcloud logging buckets describe _Default --location=global --project=<project-id>   # retentionDays: 30 unless changed
T+0–15 minDeclare, name one incident lead, move to out-of-band comms
Why: if the attacker holds an admin identity they may also read the tenant's email and chat. One decision owner stops two engineers containing in conflicting ways — or one of them "fixing" the attacker's resources before they are preserved.
T+0–15 minClassify before acting
Which principal: a long-lived access key, a role or service-account session, a service-principal secret, a console user, the root / Global Admin? Which accounts, subscriptions and projects? What is happening: discovery, resource abuse (mining), data access, destruction? Why: the kind of credential decides how it is revoked, and the stage reached decides how wide to cut. Answer How Far Has It Got?, then pick the card in By Scenario.
T+15–30 minRevoke the credential — and the sessions it already minted
Deactivating a key or resetting a password does not end sessions already issued from it, and the three clouds differ in what does. AWS: AWSDenyAll on the user denies the user's own sessions on each call, but not sessions of roles the user assumed — each of those roles needs its own aws:TokenIssueTime deny, and that deny also breaks every workload sharing the role until it fetches a new session. Azure: disabling a service principal or running Revoke-MgUserSignInSession stops new tokens and kills refresh tokens; access tokens already issued stay valid until they expire (60–90 minutes by default), and Azure Resource Manager does not enforce the revocation early, so removing the role assignments is what actually ends ARM access. A managed identity's tokens are cached for about 24 hours and cannot be forced to refresh. GCP: disabling the service account rejects its existing access tokens (IAM API reference); disabling only a key does not say so, so disable the account. Why: a "disabled" key whose temporary credentials keep working for an hour is the most common way cloud containment silently fails. The full commands, per cloud, are under Hunt & Act.
# AWS: the user, then every role it assumed (one revoke per role; expect shared workloads to break)
aws iam update-access-key --user-name <user> --access-key-id <AKIA...> --status Inactive
aws iam attach-user-policy --user-name <user> --policy-arn arn:aws:iam::aws:policy/AWSDenyAll
aws iam put-role-policy --role-name <role> --policy-name AWSRevokeOlderSessions --policy-document file://revoke-older-sessions.json   # Deny * where aws:TokenIssueTime < now

# Azure: identities off, every credential gone (SP and app), roles stripped — roles are the real cut
Update-MgServicePrincipal -ServicePrincipalId <sp-objectId> -AccountEnabled:$false
az ad sp credential list --id <sp-objectId>; az ad app credential list --id <appId>; az ad app federated-credential list --id <appId>
Update-MgUser -UserId <upn> -AccountEnabled:$false; Revoke-MgUserSignInSession -UserId <upn>
az role assignment list --assignee <objectId> --all --include-inherited -o table   # then delete at every scope

# GCP: the account, not just the key
gcloud iam service-accounts disable <sa-email>
gcloud iam service-accounts keys disable <key-id> --iam-account=<sa-email>
T+15–30 minQuarantine the account, subscription or project if the scope is unknown
AWS: move the account into the quarantine OU, whose deny-all SCP exempts only your IR role. GCP: attach an IAM deny policy at the project or folder — it applies to tokens already minted, but only covers the permissions on Google's supported list, so when in doubt replace the project's IAM policy with an IR-only one (the T0 dump is what you restore from). Azure has no deny-all equivalent you can apply ad hoc: Azure Policy deny only blocks creating or changing resources through ARM, denyAction only blocks deletes, Deployment Stacks' deny settings only cover the stack's own resources, and none of them touch data-plane reads or Entra — so strip role assignments at subscription and management-group scope, disable the identities, and use a denyAction assignment to pin resources against deletion while you work. Why: this is the wide cut: one move that stops every credential in the boundary, including the ones you haven't found yet. Two gaps to know: SCPs never apply to the management account or to service-linked roles, and they don't stop principals from other organisations using access a resource policy already granted — that needs a resource control policy or the resource's own policy.
aws organizations move-account --account-id <id> \
  --source-parent-id <current-ou> --destination-parent-id <quarantine-ou>
gcloud iam policies create ir-quarantine --kind=denypolicies \
  --attachment-point=cloudresourcemanager.googleapis.com%2Fprojects%2F<project-id> \
  --policy-file=deny-all-except-ir.json
T+30–45 minProtect the way back
Confirm backups, snapshots, object versions and keys still exist and can't be deleted by the compromised principal: AWS Backup vault lock, S3 versioning and Object Lock, Azure Recovery Services soft delete (14 days) and multi-user authorisation through a Resource Guard, GCS soft delete (7 days by default, 0–90 configurable) and retention locks. Keys have windows too: a KMS key scheduled for deletion waits 7–30 days and can be cancelled, a Key Vault with soft delete keeps a deleted vault 7–90 days unless it was purged, a Cloud KMS version in DESTROY_SCHEDULED waits at least 24 hours and can be restored. Check each one now, then copy critical snapshots to an account the attacker has no path to. Why: deleting snapshots, backups and keys is how a cloud attacker turns a breach into an outage; once they see containment start, it is the next move, and a lifecycle rule that expires everything tomorrow is the quiet version.
aws s3api get-bucket-versioning --bucket <bucket>; aws s3api get-bucket-lifecycle-configuration --bucket <bucket>
aws s3api get-object-lock-configuration --bucket <bucket>
aws backup describe-backup-vault --backup-vault-name <vault>   # Locked: true
aws kms list-keys --query 'Keys[].KeyId' --output text | xargs -n1 aws kms describe-key --key-id --query 'KeyMetadata.[KeyId,KeyState,DeletionDate]' --output text | grep PendingDeletion
aws kms cancel-key-deletion --key-id <key-id>
az backup vault backup-properties show -g <rg> -n <vault>        # softDeleteFeatureState: Enabled
az keyvault list-deleted -o table; az keyvault show -n <vault> --query 'properties.enablePurgeProtection'
gcloud storage buckets describe gs://<bucket> --format='value(soft_delete_policy,retention_policy)'
gcloud kms keys versions list --key=<key> --keyring=<ring> --location=<loc> --filter='state=DESTROY_SCHEDULED'
T+30–60 minMemory, then snapshot, then isolate compromised workloads
Only where a VM, container host or function is the foothold, and in that order. Memory first, while the agent channel still works: SSM Run Command, az vm run-command or gcloud compute ssh running AVML (Linux) or a Windows memory tool staged in the evidence bucket, uploaded straight back to it. Then snapshot the disks; then cut the network and remove the workload's identity. Three traps: an isolation security group kills SSM unless the VPC has interface endpoints for ssm, ssmmessages and ec2messages and the group allows 443 to them; an EBS snapshot under the default aws/ebs key cannot be shared to the evidence account — copy it to a customer-managed key first; security-group and NSG rule changes don't cut connections already tracked — add a subnet NACL deny, or an Azure route table with next hop None, if the cut must be immediate. Don't terminate or deallocate. Why: terminating destroys disk and memory evidence and often the ephemeral logs on the host; detaching the instance profile stops new credentials, but ones already fetched from the metadata service stay valid until the role's sessions are revoked (step above). The per-cloud sequences are under Hunt & Act.
aws ssm send-command --instance-ids <i-id> --document-name AWS-RunShellScript \
  --parameters 'commands=["aws s3 cp s3://<evidence-bucket>/tools/avml /tmp/avml","chmod +x /tmp/avml","/tmp/avml --compress /tmp/<i-id>.lime.compressed","aws s3 cp /tmp/<i-id>.lime.compressed s3://<evidence-bucket>/<case>/"]'
aws ec2 create-snapshot --volume-id <vol-id> --description "IR case <n>"
aws ec2 copy-snapshot --source-region <region> --source-snapshot-id <snap-id> --encrypted --kms-key-id <evidence-cmk-arn>
aws ec2 modify-instance-attribute --instance-id <i-id> --groups <isolation-sg>
aws ec2 modify-instance-attribute --instance-id <i-id> --disable-api-termination
aws ec2 disassociate-iam-instance-profile --association-id <iip-assoc-id>
T+0–60 minCall the insurer, counsel — and open a case with the provider
Why: the insurer's rules on approved IR firms apply here too, and counsel keeps the findings privileged. The provider sees things you can't — abuse reports about your IP addresses, key-exposure notices — and a billing dispute over attacker-run compute starts from an early support case.
Hour 1 — 24  •  EVICT & SCOPE

Hour 1–24: Evict, Scope, Preserve

Hunt every foothold the attacker created
New users, access keys and console passwords; MFA devices added or deactivated and the password policy weakened; new roles and changed trust policies (especially trust to accounts you don't own); new policy versions and removed permissions boundaries; identity providers, Roles Anywhere trust anchors, Identity Center assignments; Lambda code, EventBridge rules, SSM associations and instance user-data that run code on a schedule. Azure: service-principal secrets, certificates and federated credentials (logged as a plain Update application), VM extensions and run commands, Automation runbooks, Logic Apps, federated credentials on managed identities, subscriptions created on your billing account. GCP: service-account keys and iam.serviceAccountTokenCreator grants, ssh-keys in instance or project metadata, OS Login turned off, Cloud Scheduler and Functions, workload identity pool providers without an attribute condition, projects created on your billing account. Why: removing the entry credential while a backdoor role trusts the attacker's own account is not eviction. The queries are under Hunt & Act.
Export first — the logs age out faster than the investigation runs
CloudTrail event history keeps 90 days of management events per region (IAM events land in us-east-1); the Azure Activity Log keeps 90 days; Entra sign-in and audit logs keep 7 days on the free tier and 30 on P1/P2; GCP Admin Activity logs keep 400 days but Data Access and flow logs sit in _Default for 30. Export into an evidence account with write-once storage, hash each export with SHA-256 and record who collected it; the T0 dumps (IAM authorization details and credential report, Resource Graph role assignments and app credentials, search-all-iam-policies) are in the Preserve step of each Hunt & Act tab. Why: every one of those windows is shorter than a typical investigation, and evidence you can't prove is unaltered is weak evidence for the insurer and the regulator.
aws cloudtrail validate-logs --trail-arn <org-trail-arn> --start-time <window-start>
aws cloudtrail lookup-events --region <region> --start-time <window-start> --max-results 50 --output json > events-<region>.json   # one region at a time, 90 d
Get-MgAuditLogSignIn -All -Filter "createdDateTime ge <window-start>" | Export-Csv signins.csv
Get-MgAuditLogDirectoryAudit -All -Filter "activityDateTime ge <window-start>" | Export-Csv directory-audit.csv
az monitor activity-log list --start-time <window-start> --subscription <sub-id> -o json > activity-<sub-id>.json
gcloud logging copy _Default gs://<evidence-bucket>/<case>/logs-<project-id> --location=global --project=<project-id>
sha256sum * | tee SHA256SUMS
Build the timeline — and let the providers' own history fill the gaps
Anchor on the attacker's first use (stage 1), then order every stage 3–9 hit by time and principal. CloudTrail Insights, if it was enabled, has already flagged the API-rate spikes; gcloud asset get-history gives the IAM policy of any resource at any point in the last 35 days without depending on logs; Resource Graph's createdOn on role assignments and timeCreated on resources date the things the attacker made. Why: the status line says how far it got (the ladder); the timeline says in what order, which is what decides scope, notification and whether the eviction is complete.
Scope data access — and say plainly where you can't
Look for storage made public, snapshots and images shared to other accounts, storage keys listed, disk SAS issued, secrets read, presigned URLs and SAS tokens issued. Object-level reads (S3 data events, StorageBlobLogs and AZKVAuditLogs, GCP Data Access logs) are off by default; BigQuery is the one exception. If they were off, you cannot show what was read — record that, and scope the notification on what the principal could read. For a workload foothold, VPC, VNet or VPC flow logs (also off by default) give the bytes out. Why: "no evidence of exfiltration" and "no logs that could show exfiltration" are different statements, and regulators treat them differently.
Find the entry point
The usual ones: a key committed to a repository or baked into an image; a credential stolen from a developer laptop; SSRF to the instance metadata service (IMDSv1); a CI/CD pipeline identity with a trust policy that is too broad; an OAuth app or third-party integration. Why: restoring before the entry point is closed means doing this again next week — with an attacker who now knows your environment.
trufflehog git https://github.com/<org>/<repo> --only-verified
aws ec2 describe-instances --query "Reservations[].Instances[?MetadataOptions.HttpTokens=='optional'].InstanceId"
Rotate every secret the principal could read
Not only its own keys: everything in Secrets Manager, Parameter Store, Key Vault and Secret Manager it had read access to, function environment variables, database passwords, and third-party API keys stored in the account. Why: an attacker who listed the secret store leaves with credentials to systems outside the cloud account, and those don't show up in any cloud log.
Day 1 — 7  •  NOTIFY & RECOVER

Day 1–7: Notify, Rebuild, Recover

Run the notification clocks from awareness, not from close
GDPR: supervisory authority within 72 h of becoming aware of a personal-data breach, data subjects without undue delay if the risk is high. NIS2: early warning within 24 h, notification within 72 h, final report within one month. DORA (financial entities): initial notification within 4 h of classifying the incident as major, and no later than 24 h after becoming aware. Why: the clocks start before scoping ends; notify with what you know and update.
Rebuild from code, don't clean in place
Redeploy compromised workloads from infrastructure-as-code and known-good images, and bring data back from snapshots or versions taken before the earliest attacker activity. Where the whole account is suspect, deploy into a new account and migrate data rather than scrubbing the old one. Gate: reconnect only after the incident lead declares eradication complete — the entry point closed, every attacker-created identity and resource removed, every readable secret rotated, and a verification hunt clean since the eviction time. Why: one missed role trust or scheduled function puts the attacker straight back in, and the cloud makes rebuilding cheaper than proving a resource is clean.
Put a number on the damage
Cost Explorer or billing export for the incident window, per region and per service, and the list of resources the attacker created. Why: attacker-run compute can reach five figures in days; the billing case with the provider and the insurance claim both need the figure and the evidence behind it.
After Day 7  •  LEARN

After Day 7: Feed the Slow Loops

The incident doesn't close at recovery. It is the most expensive data the program will collect — spend it. Machines collect, humans curate.
  • Entry route → Cloud Security posture: long-lived keys replaced by roles and workload identity, IMDSv2 enforced, pipeline trust policies narrowed — everywhere the same weakness exists.
  • What slowed containment → structure: no pre-built quarantine OU, no read-only IR role across the estate, a region lock without a responder exemption, logs you had to find before you could read them.
  • What detection missed → the detection backlog: each stage they reached without an alert (logging tampering, new keys, cross-account trust, launches in unused regions) becomes a tested detection with an owner.
  • Guardrail gaps they used → the management account, service-linked roles, external principals, self-granting admin roles: close the ones this incident proved real.
  • Indicators become intelligence only after an analyst checks them and gives each one a date, a source, an expiry and a confidence level. Attacker IPs and access-key IDs rotate; an unexpired block list becomes noise.

How Far Has It Got?

Ask these at triage, top to bottom, and again at every update. The furthest stage you can show evidence for is how far the incident has got — write it in the case's status line. A stage you cannot answer yet is a scoping gap: say when it will be known, and whether the logs to answer it exist at all.
#StageATT&CKQuestionWhere to look
1Initial accessT1078.004Which credential, from where, and when was its first use by the attacker?Per-key (IP, user agent) pairs against the previous 90 days in Athena; GuardDuty credential-exfiltration findings; the four Entra sign-in tables (user, non-interactive, service principal, managed identity); GCP callerIp, user agent and serviceAccountKeyName
2DiscoveryT1087.004 T1580Have they enumerated who they are and what exists?AWS: readonly = true bursts and AccessDenied counts per principal (GetCallerIdentity, ListBuckets, GetAccountAuthorizationDetails). Azure: MicrosoftGraphActivityLogs — the Activity Log never records reads. GCP: Data Access logs, off by default; AzureHound, ROADtools, Pacu user agents
3Privilege escalationT1098.003 T1548.005Did they give themselves more permissions?Policy versions, attachments, permissions boundaries, trust policies, iam:PassRole, Identity Center assignments; Azure role assignments and Entra elevateAccess (service “Azure RBAC (Elevated Access)” in AuditLogs); GCP SetIamPolicy with bindingDeltas
4PersistenceT1098.001 T1136.003Did they create a way back in?Users, keys, login profiles, MFA devices, password policy, roles, identity providers, Lambda code, EventBridge rules, SSM associations, Roles Anywhere; service-principal secrets and federated credentials, VM extensions, run commands, runbooks, new subscriptions; service-account keys, ssh-keys metadata, OS Login off, Scheduler and Functions, identity pools, new projects
5Defence evasionT1562.008Did they blind you?Trails stopped, narrowed or deleted, event selectors and Insights changed, DeleteEventDataStore, GuardDuty filters and IP sets, trail-bucket lifecycle and policy; diagnostic settings, workspace retention, Defender plans, Sentinel connectors, locks; sinks, exclusions, bucket retention, auditConfigDeltas, SCC mutes; activity in regions nobody uses
6Lateral movementT1550.001 T1021.007Did they reach other accounts, subscriptions or projects?AssumeRole chains via sessionContext.sessionIssuer and cross-account hops joined on sharedEventID; a principal in subscriptions it never used, Lighthouse and management-group writes, Entra elevateAccess; GCP GenerateAccessToken / SignBlob on iamcredentials (Data Access logs) with the delegation chain
7Resource abuseT1496.001 T1496.004Are they running compute on your bill?RunInstances, fleets, SageMaker, ECS/EKS, Lightsail, Batch, Bedrock entitlements, quota increases; scale sets, AKS, ML compute, Azure OpenAI deployments; GKE, Dataproc, Vertex AI, Cloud Run jobs — in every region and every project; cost anomaly alerts
8Data accessT1530 T1537What could they read, and what did they read?Could: storage made public, snapshots shared out, keys listed, SAS and beginGetAccess, GetSecretValue. Did: S3 data events, AZKVAuditLogs, StorageBlobLogs, GCP Data Access logs — all off by default; flow logs for a workload foothold
9ImpactT1485 T1486 T1490Are they deleting, re-encrypting or holding data hostage?DeleteObjects, bucket lifecycle and replication rules, SSE-C writes (SSEApplied), ScheduleKeyDeletion, DeleteBackupVault, CloseAccount; backupconfig/write, Resource Guard and vault purge; DestroyCryptoKeyVersion, billing detached, soft delete removed; ransom notes in buckets
10ExtortionT1657Have they made contact, or is data already published?Emails to staff, the root or security contact, support cases you did not open; leak-site monitoring. No cloud log records this stage

Hunt & Act by Cloud

Where to look for each stage of the cloud attack chain, and the containment actions per provider. Pick your cloud.
Test these before you need them. The queries follow each provider's documented log formats, but what you can see depends on which logs are enabled and where they are shipped. Run each one in your own environment during peacetime and fix it there — not during an incident.
Primary: Athena over the organisation trail's S3 bucket (table cloudtrail_logs from the CloudTrail docs' partition-projection DDL; column names are lower-case, useridentity is a struct, requestparameters and additionaleventdata are JSON strings). The same SQL runs in CloudTrail Lake with camel-case fields. Fallback: aws cloudtrail lookup-events — no setup, but 90 days of management events only, one region and one lookup attribute per call, 2 requests per second, and the IP and user agent are inside the CloudTrailEvent JSON string (hence the jq). IAM, STS, Organizations and Account Management events land in us-east-1. Data events (S3 objects, Lambda invokes, DynamoDB items, Bedrock invocations) exist only if the trail's event selectors include them; where a hunt needs them it says so.

Hunt

Stage 1 · Initial accessEverything one access key did, from where, with which client

Why: the key is the thread to pull. The first call from an IP or user agent the owner never uses is the attacker's first use; everything after it from that pair is in scope.

No trail at all? aws iam get-access-key-last-used --access-key-id <AKIA...> still gives the last service, region and time. A key AWS found in public is also announced as an AWS Health event AWS_RISK_CREDENTIALS_EXPOSED and a quarantine policy on the user.

-- first and last use per source: which (IP, user agent) pairs are the owner, which are the attacker
SELECT sourceipaddress, useragent, min(eventtime) AS first_seen, max(eventtime) AS last_seen,
       count(*) AS calls, count_if(errorcode IS NOT NULL) AS errors,
       array_join(array_agg(DISTINCT eventsource), ',') AS services
FROM cloudtrail_logs
WHERE useridentity.accesskeyid = '<AKIA...>'
  AND eventtime > '<window-start ISO8601>'
GROUP BY 1, 2 ORDER BY first_seen;

-- then the full timeline for the attacker's pairs
SELECT eventtime, awsregion, eventsource, eventname, sourceipaddress, useragent, errorcode, requestparameters
FROM cloudtrail_logs
WHERE useridentity.accesskeyid = '<AKIA...>' AND sourceipaddress IN ('<attacker-ip>')
ORDER BY eventtime;

# no Athena: one region, 90 days, IP and UA pulled out of the CloudTrailEvent string
aws cloudtrail lookup-events --region <region> --start-time <window-start> \
  --lookup-attributes AttributeKey=AccessKeyId,AttributeValue=<AKIA...> --output json \
  | jq -r '.Events[].CloudTrailEvent | fromjson
           | [.eventTime, .awsRegion, .eventSource, .eventName, .sourceIPAddress, .userAgent, (.errorCode // "")] | @tsv'
T1078.004 Tuning: the owner's CI runners, NAT gateways and office egress produce stable (IP, user agent) pairs; compare the incident window against the key's previous 90 days before calling a pair the attacker's. A user agent of aws-cli on a key that only ever showed an SDK string is the usual tell.

Stage 2 · DiscoveryEnumeration bursts and permission probing

Why: an attacker's first minutes are GetCallerIdentity, ListBuckets, GetAccountAuthorizationDetails, ListSecrets and a spray of Describe/List calls, many of them denied. A principal that touches forty distinct read APIs in an hour, or collects twenty AccessDenieds, is enumerating. GetCallerIdentity is never denied and never needs a permission, so it is the one call every stolen key makes.
SELECT useridentity.arn AS principal, sourceipaddress, useragent,
       date_trunc('hour', from_iso8601_timestamp(eventtime)) AS hour,
       count(DISTINCT eventname) AS distinct_calls,
       count_if(errorcode = 'AccessDenied') AS denied,
       count_if(eventname IN ('GetCallerIdentity','ListBuckets','GetAccountAuthorizationDetails','ListSecrets',
                              'ListUsers','ListRoles','ListAccessKeys','GetAccountSummary','DescribeRegions')) AS hallmark_calls,
       array_join(array_agg(DISTINCT eventsource), ',') AS services
FROM cloudtrail_logs
WHERE eventtime > '<window-start ISO8601>' AND readonly = 'true'
GROUP BY 1, 2, 3, 4
HAVING count(DISTINCT eventname) > 40 OR count_if(errorcode = 'AccessDenied') > 20
ORDER BY denied DESC, distinct_calls DESC;
T1087.004 T1580 T1526 Tuning: Prowler, ScoutSuite, Steampipe, AWS Config, Security Hub and the SIEM collector role do exactly this on a schedule; exclude them by useridentity.arn and keep the thresholds above their usual counts. A human on a new laptop running aws configure and clicking through the console also trips the forty-call line once.

Stage 3 · Privilege escalationPolicies, boundaries and trust changed; a role handed to compute

Why: AWS escalation is a policy version set as default, a policy attached, a permissions boundary removed, a trust policy widened, a login profile added to someone else's user, or iam:PassRole handing an admin role to an instance or function the attacker controls. Identity Center assignments are the same thing one level up.

The before/after of a policy version: aws iam get-policy-version --policy-arn <arn> --version-id v<n> for the attacker's version and the previous one. The T0 dump from the Preserve step is the authoritative “before”.

SELECT eventtime, eventsource, eventname, useridentity.arn AS actor, sourceipaddress, requestparameters, errorcode
FROM cloudtrail_logs
WHERE eventtime > '<window-start ISO8601>' AND (
      (eventsource = 'iam.amazonaws.com' AND eventname IN (
         'CreatePolicyVersion','SetDefaultPolicyVersion','AttachUserPolicy','AttachRolePolicy','AttachGroupPolicy',
         'PutUserPolicy','PutRolePolicy','PutGroupPolicy','AddUserToGroup','UpdateAssumeRolePolicy',
         'PutUserPermissionsBoundary','DeleteUserPermissionsBoundary','PutRolePermissionsBoundary','DeleteRolePermissionsBoundary',
         'CreateLoginProfile','UpdateLoginProfile'))
   OR (eventsource = 'sso.amazonaws.com' AND eventname IN (
         'CreateAccountAssignment','CreatePermissionSet','AttachManagedPolicyToPermissionSet',
         'PutInlinePolicyToPermissionSet','AttachCustomerManagedPolicyReferenceToPermissionSet'))
   -- iam:PassRole: a privileged role attached to compute at launch
   OR (eventname IN ('RunInstances','CreateFunction20150331','CreateStack','CreateAutoScalingGroup','RegisterTaskDefinition')
       AND regexp_like(requestparameters, 'iamInstanceProfile|"role"|roleArn|RoleARN')))
ORDER BY eventtime;
T1098.003 T1548.005 Tuning: landing-zone and Terraform pipelines write policies and assignments on every apply, and RunInstances with an instance profile is normal. Filter on useridentity.arn not being the deploy role, then read the policy document in requestparameters: "Action":"*", AdministratorAccess, or a trust principal in an account you don't own is the finding.

Stage 4 · PersistenceUsers, keys, MFA, identity providers, functions, rules, SSM, Roles Anywhere

Why: IAM is global and lands in us-east-1; the rest is per region. Persistence is not only a new user: a deactivated MFA device, a weakened password policy, an OIDC provider you did not create, a Lambda whose code was swapped, an EventBridge rule that invokes it, an SSM association that runs a command on every instance, or a Roles Anywhere trust anchor that turns the attacker's own CA into AWS credentials.

Lambda event names carry an API-version suffix (CreateFunction20150331, UpdateFunctionCode…, AddPermission20150331v2), hence the LIKE. For every hit, the T0 dump answers whether the object existed before the incident.

SELECT eventtime, awsregion, eventsource, eventname, useridentity.arn AS actor, sourceipaddress, requestparameters
FROM cloudtrail_logs
WHERE eventtime > '<window-start ISO8601>' AND errorcode IS NULL AND (
      (eventsource = 'iam.amazonaws.com' AND eventname IN (
         'CreateUser','CreateAccessKey','CreateLoginProfile','UpdateLoginProfile','CreateRole','UpdateAssumeRolePolicy',
         'CreateSAMLProvider','UpdateSAMLProvider','CreateOpenIDConnectProvider','AddClientIDToOpenIDConnectProvider',
         'UpdateOpenIDConnectProviderThumbprint','CreateVirtualMFADevice','EnableMFADevice','DeactivateMFADevice',
         'DeleteVirtualMFADevice','UpdateAccountPasswordPolicy','DeleteAccountPasswordPolicy',
         'CreateServiceSpecificCredential','UploadSSHPublicKey','UploadSigningCertificate'))
   OR (eventsource = 'lambda.amazonaws.com' AND (eventname LIKE 'CreateFunction%' OR eventname LIKE 'UpdateFunctionCode%'
         OR eventname LIKE 'UpdateFunctionConfiguration%' OR eventname LIKE 'AddPermission%' OR eventname LIKE 'CreateEventSourceMapping%'
         OR eventname = 'CreateFunctionUrlConfig'))
   OR (eventsource = 'events.amazonaws.com' AND eventname IN ('PutRule','PutTargets'))
   OR (eventsource = 'ssm.amazonaws.com' AND eventname IN ('CreateAssociation','UpdateAssociation','SendCommand','StartSession','CreateDocument','UpdateDocument'))
   OR (eventsource = 'rolesanywhere.amazonaws.com' AND eventname IN ('CreateTrustAnchor','UpdateTrustAnchor','CreateProfile','UpdateProfile'))
   OR (eventsource = 'sso.amazonaws.com' AND eventname IN ('CreateAccountAssignment','CreatePermissionSet','CreateInstance'))
   OR (eventsource = 'ec2.amazonaws.com' AND eventname IN ('ModifyInstanceAttribute','CreateLaunchTemplateVersion') AND requestparameters LIKE '%userData%'))
ORDER BY eventtime;
T1098.001 T1136.003 T1556.006 Tuning: CloudFormation, Terraform and Serverless deploys create roles, functions, rules and associations all day, usually with useridentity.invokedby set or from a known deploy role; exclude those first. What is left that was made by a human, an access key, or a role the attacker held is the finding; so is any MFA deactivation or password-policy change at all.

Stage 5 · Defence evasionLogging and detection switched off, narrowed or starved

Why: the time of the first hit is where your visibility ends and the scoping gap begins. The loud version is StopLogging or DeleteDetector; the quiet ones are UpdateTrail dropping multi-region or global events, PutEventSelectors removing data events, a GuardDuty filter that archives findings, an IP set that trusts the attacker's address, a lifecycle rule or policy change on the trail bucket, and DeleteEventDataStore on CloudTrail Lake.
SELECT eventtime, awsregion, eventsource, eventname, useridentity.arn AS actor, sourceipaddress, requestparameters
FROM cloudtrail_logs
WHERE eventtime > '<window-start ISO8601>' AND errorcode IS NULL AND (
      (eventsource = 'cloudtrail.amazonaws.com' AND eventname IN (
         'StopLogging','DeleteTrail','UpdateTrail','PutEventSelectors','PutInsightSelectors',
         'DeleteEventDataStore','UpdateEventDataStore','StopEventDataStoreIngestion'))
   OR (eventsource = 'guardduty.amazonaws.com' AND eventname IN (
         'DeleteDetector','UpdateDetector','CreateFilter','UpdateFilter','CreateIPSet','UpdateIPSet',
         'CreateThreatIntelSet','UpdateThreatIntelSet','DeleteMembers','DisassociateFromAdministratorAccount',
         'DeletePublishingDestination','UpdateMalwareScanSettings'))
   OR (eventsource = 'config.amazonaws.com' AND eventname IN ('StopConfigurationRecorder','DeleteConfigurationRecorder','DeleteDeliveryChannel','DeleteConfigRule'))
   OR (eventsource = 'securityhub.amazonaws.com' AND eventname IN ('DisableSecurityHub','BatchDisableStandards','DisableImportFindingsForProduct'))
   OR (eventsource = 'ec2.amazonaws.com' AND eventname = 'DeleteFlowLogs')
   OR (eventsource = 'macie2.amazonaws.com' AND eventname = 'DisableMacie')
   OR (eventsource = 'detective.amazonaws.com' AND eventname = 'DeleteGraph')
   OR (eventsource = 'logs.amazonaws.com' AND eventname IN ('DeleteLogGroup','PutRetentionPolicy','DeleteSubscriptionFilter','DeleteResourcePolicy'))
   OR (eventsource = 's3.amazonaws.com'
       AND eventname IN ('PutBucketLifecycle','PutBucketPolicy','DeleteBucketPolicy','PutBucketVersioning','DeleteBucket','PutBucketAcl')
       AND json_extract_scalar(requestparameters, '$.bucketName') IN ('<trail-bucket>', '<flow-log-bucket>')))
ORDER BY eventtime;
T1562.008 T1562.001 Tuning: the central security account's automation legitimately calls UpdateDetector, UpdateFilter and UpdateTrail; a log-pipeline change ticket explains PutRetentionPolicy and trail-bucket lifecycle edits. Match each hit to a ticket; anything from the compromised principal, or without one, is the finding.

Stage 5 · Defence evasionGuardDuty findings you haven't read, log-file integrity, Insights

Why: GuardDuty writes the evasion findings for you if it was running; validate-logs proves the trail files in S3 were not altered or deleted after delivery (needs log-file validation on the trail); CloudTrail Insights, if enabled, has already flagged the API-rate spike that the attacker's enumeration or mass-delete produced.
aws guardduty list-findings --region <region> --detector-id <detector-id> --finding-criteria '{"Criterion":{"type":{"Eq":[
  "Stealth:IAMUser/CloudTrailLoggingDisabled","Stealth:IAMUser/PasswordPolicyChange","Policy:IAMUser/RootCredentialUsage",
  "UnauthorizedAccess:IAMUser/TorIPCaller","UnauthorizedAccess:IAMUser/InstanceCredentialExfiltration.OutsideAWS",
  "Discovery:IAMUser/AnomalousBehavior","Persistence:IAMUser/AnomalousBehavior","Exfiltration:IAMUser/AnomalousBehavior",
  "CredentialAccess:IAMUser/AnomalousBehavior","DefenseEvasion:IAMUser/AnomalousBehavior"]}}}' \
  | jq -r '.FindingIds[]' | xargs -r aws guardduty get-findings --region <region> --detector-id <detector-id> --finding-ids \
  | jq -r '.Findings[] | [.UpdatedAt, .Type, .Severity, .Resource.AccessKeyDetails.UserName // "", .Title] | @tsv'

# were the delivered log files tampered with after the fact?
aws cloudtrail validate-logs --region <trail-home-region> --trail-arn <org-trail-arn> --start-time <window-start>

# Insights (only if PutInsightSelectors was ever applied to the trail)
aws cloudtrail lookup-events --region <region> --event-category insight --start-time <window-start> \
  --query 'Events[].[EventTime,EventName,Resources[0].ResourceName]' --output table
T1562.008 Tuning: GuardDuty's AnomalousBehavior families fire on new admins and new automation; read the finding's “unusual” section before treating it as the attacker. A validate-logs gap that lines up with a bucket lifecycle rule is retention, not tampering.

Stage 6 · Lateral movementRole chains and cross-account assumption

Why: cloud lateral movement is AssumeRole. A session assuming another role is a chain; a target account that differs from the caller's account is a hop into another account. useridentity.sessioncontext.sessionissuer.arn names the role the caller already held, requestparameters.roleArn the one it took, and sharedeventid ties the caller-side record to the record in the target account on the organisation trail.

Console role switches are SwitchRole under signin.amazonaws.com. The same principal or IP in a partner's or customer's trail is the next question — ask them.

SELECT eventtime, useridentity.type AS caller_type, useridentity.arn AS caller,
       useridentity.sessioncontext.sessionissuer.arn AS caller_role,
       json_extract_scalar(requestparameters, '$.roleArn') AS target_role,
       json_extract_scalar(requestparameters, '$.roleSessionName') AS session_name,
       useridentity.accountid AS from_account, recipientaccountid AS to_account,
       sharedeventid, sourceipaddress, useragent, errorcode
FROM cloudtrail_logs
WHERE eventsource = 'sts.amazonaws.com'
  AND eventname IN ('AssumeRole','AssumeRoleWithSAML','AssumeRoleWithWebIdentity','GetFederationToken','AssumeRoot')
  AND eventtime > '<window-start ISO8601>'
  AND useridentity.invokedby IS NULL                      -- drop AWS services assuming roles on your behalf
  AND (useridentity.accountid <> recipientaccountid        -- cross-account hop
       OR useridentity.type = 'AssumedRole')               -- role assuming a role: a chain
ORDER BY eventtime;
T1550.001 T1021.007 Tuning: every workload assumes roles constantly and pipelines chain them by design. Exclude the known pipeline and instance roles, then compare each (caller_role, target_role) pair against the previous 30 days: a pair that never occurred before, a target in an account you don't own, or AssumeRoot at all, is the finding.

Stage 7 · Resource abuseCompute, fleets, GPUs, model access and quota increases in every region

Why: miners launch where you don't look and in the shapes you don't watch: spot fleets, SageMaker notebooks and training jobs, ECS tasks, EKS node groups, Lightsail, Batch. A RequestServiceQuotaIncrease for GPU families is the attacker preparing. Since 2024 the same stolen key is used to enable and invoke Bedrock models and resell the access.

What is live right now: loop aws ec2 describe-instances --region $r --filters Name=instance-state-name,Values=running --query 'Reservations[].Instances[].[InstanceId,InstanceType,LaunchTime]' over describe-regions. Bedrock InvokeModel is a data event; the entitlement calls below are management events.

SELECT eventtime, awsregion, eventsource, eventname, useridentity.arn AS actor, sourceipaddress,
       json_extract_scalar(requestparameters, '$.instanceType') AS instance_type, errorcode
FROM cloudtrail_logs
WHERE eventtime > '<window-start ISO8601>' AND useridentity.invokedby IS NULL AND (
      (eventsource = 'ec2.amazonaws.com' AND eventname IN ('RunInstances','CreateFleet','RequestSpotFleet','RequestSpotInstances','CreateLaunchTemplate'))
   OR (eventsource = 'sagemaker.amazonaws.com' AND eventname IN ('CreateNotebookInstance','CreateTrainingJob','CreateProcessingJob','CreateEndpoint'))
   OR (eventsource = 'ecs.amazonaws.com' AND eventname IN ('RunTask','CreateService','RegisterTaskDefinition'))
   OR (eventsource = 'eks.amazonaws.com' AND eventname IN ('CreateCluster','CreateNodegroup','CreateFargateProfile'))
   OR (eventsource = 'lightsail.amazonaws.com' AND eventname IN ('CreateInstances','CreateContainerService'))
   OR (eventsource = 'batch.amazonaws.com' AND eventname IN ('CreateComputeEnvironment','SubmitJob'))
   OR (eventsource = 'servicequotas.amazonaws.com' AND eventname = 'RequestServiceQuotaIncrease')
   OR (eventsource = 'bedrock.amazonaws.com' AND eventname IN ('PutFoundationModelEntitlement','PutUseCaseForModelAccess','CreateModelInvocationJob','InvokeModel')))
ORDER BY eventtime;
T1496.001 T1496.004 Tuning: Auto Scaling, spot fleet and Batch launch instances on your behalf with useridentity.invokedby set, which the query already drops. CI runners and data-science teams launch GPU instances legitimately; a launch in a region you never use, a p/g family you never ordered, or InstanceLimitExceeded errors followed by a quota request is the finding. Cost Explorer by region and service for the same window confirms it.

Stage 8 · Data accessStorage opened, snapshots shared, secrets read, disks exported

Why: sharing a snapshot or AMI with an outside account copies a whole disk out without a single data-plane read; a public bucket policy or a removed public-access block does the same for objects; GetSecretValue and a decrypted GetParameter are management events, so the secret reads are visible even without data events.
SELECT eventtime, awsregion, eventsource, eventname, useridentity.arn AS actor, sourceipaddress, requestparameters
FROM cloudtrail_logs
WHERE eventtime > '<window-start ISO8601>' AND errorcode IS NULL AND (
      (eventsource = 's3.amazonaws.com' AND (eventname IN ('PutBucketPolicy','PutBucketAcl','PutAccessPointPolicy','PutBucketCors')
                                             OR eventname LIKE '%PublicAccessBlock%'))
   OR (eventsource = 'ec2.amazonaws.com' AND eventname IN ('ModifySnapshotAttribute','ModifyImageAttribute','CopySnapshot','CreateSnapshot','CreateSnapshots','CreateImage'))
   OR (eventsource = 'rds.amazonaws.com' AND eventname IN ('ModifyDBSnapshotAttribute','ModifyDBClusterSnapshotAttribute','CopyDBSnapshot','CreateDBSnapshot','StartExportTask'))
   OR (eventsource = 'secretsmanager.amazonaws.com' AND eventname IN ('GetSecretValue','BatchGetSecretValue','ListSecrets','PutResourcePolicy'))
   OR (eventsource = 'ssm.amazonaws.com' AND eventname IN ('GetParameter','GetParameters','GetParametersByPath') AND requestparameters LIKE '%"withDecryption":true%')
   OR (eventsource = 'kms.amazonaws.com' AND eventname IN ('PutKeyPolicy','CreateGrant','ReplicateKey'))
   OR (eventsource = 'dynamodb.amazonaws.com' AND eventname = 'ExportTableToPointInTime')
   OR (eventsource = 'sts.amazonaws.com' AND eventname = 'GetFederationToken'))
ORDER BY eventtime;
T1530 T1537 T1555.006 Tuning: workloads call GetSecretValue on every start and AWS Backup copies snapshots on a schedule (invokedby backup.amazonaws.com). Group by actor: a human, an access key, or a role the attacker held reading secrets it never read before, or a snapshot shared to an account not in your organisation, is the finding.

Stage 8 · Data accessWhat was read — object-level S3 data events

Why: the only record of an object read is an S3 data event, and the trail only has them if its event selectors included S3 objects before the incident. If they were off, write that down — “no evidence of exfiltration” and “no logs that could show exfiltration” are different statements — and scope on what the principal could read. GuardDuty S3 Protection (Exfiltration:S3/AnomalousBehavior) and bucket server-access logs are the only other sources.
-- requires S3 data events on the trail (eventCategory = Data); bytesTransferredOut is in additionaleventdata
SELECT useridentity.arn AS actor, sourceipaddress, useragent,
       json_extract_scalar(requestparameters, '$.bucketName') AS bucket,
       count(*) AS reads,
       sum(cast(json_extract_scalar(additionaleventdata, '$.bytesTransferredOut') AS bigint)) AS bytes_out,
       min(eventtime) AS first_read, max(eventtime) AS last_read
FROM cloudtrail_logs
WHERE eventsource = 's3.amazonaws.com'
  AND eventname IN ('GetObject','HeadObject','ListObjects','ListObjectsV2','CopyObject','SelectObjectContent','GetObjectTorrent')
  AND eventtime > '<window-start ISO8601>'
  AND (useridentity.accesskeyid = '<AKIA... or ASIA...>' OR sourceipaddress IN ('<attacker-ip>'))
GROUP BY 1, 2, 3, 4
ORDER BY bytes_out DESC;
T1530 Tuning: replication, backup, CDN origin fetches and analytics jobs read entire buckets; exclude the replication and backup roles and anything with invokedby. Order by bytes, not by count: a ListObjects storm is reconnaissance, a GetObject sweep with the bytes to match is exfiltration.

Stage 8 · Data accessBytes out of a compromised instance — VPC flow logs

Why: for a workload foothold the control plane does not show the exfiltration or the C2; the ENI's flow logs do, as 5-tuples with byte counts. They are off by default and the default format has no direction flag, so read from the instance's point of view. GuardDuty reads flow logs independently of whether you enabled them, so its Backdoor:EC2/C&CActivity.B, CryptoCurrency:EC2/BitcoinTool.B and Trojan:EC2/DNSDataExfiltration findings exist even when your own flow logs don't.
-- table from the VPC flow logs Athena DDL (default v2 fields); "end" is reserved, hence the quotes
SELECT dstaddr, dstport, protocol, sum(bytes) AS bytes, count(*) AS flows,
       from_unixtime(min(start)) AS first_seen, from_unixtime(max("end")) AS last_seen
FROM vpc_flow_logs
WHERE interface_id = '<eni-of-compromised-instance>' AND action = 'ACCEPT'
  AND start > to_unixtime(timestamp '<window-start, e.g. 2026-09-20 00:00:00>')
GROUP BY 1, 2, 3
ORDER BY bytes DESC
LIMIT 50;
T1041 T1048 Tuning: backups and artefact uploads to S3 through a gateway endpoint, package mirrors and the SSM/CloudWatch agents dominate bytes. Drop the S3 and DynamoDB prefix-list ranges (aws ec2 describe-prefix-lists) and your known egress first; what is left on 443 to an address with no DNS history, or anything on 4444, 8333 or 3333, is the finding.

Stage 9 · ImpactDeletion, re-encryption, lifecycle wipes, backup and key destruction

Why: cloud ransomware is a bucket lifecycle rule that expires everything tomorrow, a replication rule to the attacker's bucket, objects copied over themselves with a customer-provided key (SSE-C — AWS never holds the key, SSEApplied = SSE_C in the data event), a KMS key scheduled for deletion (7–30 day window, nothing readable after), snapshots and recovery points deleted, and CloseAccount or LeaveOrganization to take the account out of your reach.

A DeleteObject whose requestparameters carries a versionId is a permanent delete of that version, not a delete marker. DeleteObjects and SSE-C writes are data events.

SELECT eventtime, awsregion, eventsource, eventname, useridentity.arn AS actor, sourceipaddress,
       json_extract_scalar(additionaleventdata, '$.SSEApplied') AS sse, requestparameters
FROM cloudtrail_logs
WHERE eventtime > '<window-start ISO8601>' AND errorcode IS NULL AND (
      (eventsource = 's3.amazonaws.com' AND eventname IN ('DeleteBucket','PutBucketLifecycle','PutBucketReplication','DeleteBucketReplication',
                                                          'PutBucketVersioning','PutBucketEncryption','PutObjectLockConfiguration'))
   OR (eventsource = 's3.amazonaws.com' AND eventname IN ('DeleteObjects','DeleteObject','PutObject','CopyObject')     -- data events
       AND (json_extract_scalar(additionaleventdata, '$.SSEApplied') = 'SSE_C' OR eventname LIKE 'Delete%'))
   OR (eventsource = 'ec2.amazonaws.com' AND eventname IN ('DeleteSnapshot','DeregisterImage','DeleteVolume','TerminateInstances'))
   OR (eventsource = 'rds.amazonaws.com' AND eventname IN ('DeleteDBSnapshot','DeleteDBClusterSnapshot','DeleteDBInstance','DeleteDBCluster'))
   OR (eventsource = 'backup.amazonaws.com' AND eventname IN ('DeleteBackupVault','DeleteRecoveryPoint','DeleteBackupPlan','DeleteBackupSelection',
                                                              'PutBackupVaultAccessPolicy','DeleteBackupVaultLockConfiguration'))
   OR (eventsource = 'kms.amazonaws.com' AND eventname IN ('ScheduleKeyDeletion','DisableKey','DeleteAlias','DeleteImportedKeyMaterial','PutKeyPolicy'))
   OR (eventsource = 'dynamodb.amazonaws.com' AND eventname IN ('DeleteTable','DeleteBackup','UpdateContinuousBackups'))
   OR (eventsource = 'glacier.amazonaws.com' AND eventname IN ('DeleteVault','DeleteArchive'))
   OR (eventsource = 'organizations.amazonaws.com' AND eventname IN ('CloseAccount','LeaveOrganization','RemoveAccountFromOrganization')))
ORDER BY eventtime;
T1485 T1486 T1490 Tuning: lifecycle expirations and your own cleanup jobs delete objects in bulk, and snapshot lifecycle policies delete snapshots daily under invokedby. Compare each actor's delete count per hour with its previous 30 days; any SSE-C at all in an estate that never uses it, any ScheduleKeyDeletion, and any backup-vault change outside a ticket is the finding.
Stage 10 · ExtortionNo hunt in this tool. No CloudTrail event records the extortion contact. It arrives by email to the root and alternate security contacts, by a Support case you did not open (aws support describe-cases --include-resolved-cases, Business plan or higher), or through the provider's abuse channel; the ransom-note object itself is a stage 9 PutObject that needs S3 data events to be visible. The mail tenant and leak-site monitoring are the sources.

Act

Preserve first, then cut broad, then surgical once the scope is known. Containment rewrites IAM, so the T0 dump comes before any of it; memory comes before isolation; compromised workloads are rebuilt from infrastructure-as-code, not cleaned.

T+0–15 · PreserveValidate the trail and dump the T0 state before you change anything

Why: every revoke and quarantine below rewrites IAM; the authorization-details dump and the credential report are the only record of what the attacker's principals held, and which keys existed, before you cut them. validate-logs needs log-file validation enabled on the trail and proves the files in S3 are the ones CloudTrail delivered. The hash list is the chain of custody.
aws cloudtrail validate-logs --region <trail-home-region> --trail-arn <org-trail-arn> --start-time <window-start>
aws iam get-account-authorization-details > t0-iam-$(date -u +%Y%m%dT%H%MZ).json
aws iam generate-credential-report && sleep 5 && \
  aws iam get-credential-report --query Content --output text | base64 -d > t0-credential-report.csv
aws organizations list-accounts > t0-accounts.json          # from the management or a delegated-admin account
aws sso-admin list-instances > t0-identity-center.json       # then list-permission-sets / list-account-assignments per instance
sha256sum t0-* | tee t0-SHA256SUMS
aws s3 cp . s3://<evidence-bucket>/<case>/t0/ --recursive --exclude '*' --include 't0-*'   # bucket with Object Lock, in the evidence account

T+15–30 · QuarantineMove the account into the quarantine OU

Why: the OU's deny-all SCP stops every credential in the account at once, including the ones you haven't found; only your IR role is exempt. Three gaps to know — SCPs never apply to the management account or to service-linked roles, and they don't stop principals in other organisations using access a resource policy already granted them (that needs a resource control policy or the resource's own policy). The OU and SCP must exist before the incident; building them now takes longer than the attacker needs.
aws organizations move-account --account-id <id> \
  --source-parent-id <current-ou> --destination-parent-id <quarantine-ou>
aws organizations list-policies-for-target --target-id <id> --filter SERVICE_CONTROL_POLICY --output table   # confirm it applied

T+15–30 · RevokeDeactivate the user's key, deny the user, and find the roles it assumed

Why: an inactive key stops new calls, and AWSDenyAll on the user denies the user's own sessions (console, GetSessionToken) on each call. It does not touch sessions of roles the user assumed — those carry the role's permissions, not the user's, and keep working until the role itself is revoked (next step). Deleting the login profile removes the console password.
aws iam update-access-key --user-name <user> --access-key-id <AKIA...> --status Inactive
aws iam attach-user-policy --user-name <user> --policy-arn arn:aws:iam::aws:policy/AWSDenyAll
aws iam delete-login-profile --user-name <user>          # NoSuchEntity means there was no console password
# every role this user assumed in the window needs its own session revoke (next step)
aws cloudtrail lookup-events --region us-east-1 --start-time <window-start> \
  --lookup-attributes AttributeKey=Username,AttributeValue=<user> --output json \
  | jq -r '.Events[] | select(.EventName == "AssumeRole") | .CloudTrailEvent | fromjson | .requestParameters.roleArn' | sort -u

T+15–30 · RevokeRevoke the role's active sessions — and expect shared workloads to break

Why: the inline deny on aws:TokenIssueTime rejects every session minted before the timestamp, including the attacker's, on each call. It also rejects every legitimate session of that role: instance profiles, Lambda, EKS pods and pipelines using it fail until they fetch a new session, which SDKs only do at expiry — restart them. New sessions still work until the trust policy is fixed. The console's “Revoke active sessions” writes exactly this policy.
NOW=$(date -u +%Y-%m-%dT%H:%M:%SZ)
cat > revoke-older-sessions.json <<JSON
{"Version":"2012-10-17","Statement":[{"Effect":"Deny","Action":"*","Resource":"*",
  "Condition":{"DateLessThan":{"aws:TokenIssueTime":"$NOW"}}}]}
JSON
aws iam put-role-policy --role-name <role> --policy-name AWSRevokeOlderSessions --policy-document file://revoke-older-sessions.json
aws iam get-role --role-name <role> --query Role.AssumeRolePolicyDocument       # then narrow the trust before anyone re-assumes it

T+15–30 · RevokeRoot compromised: there is no session revoke — replace the credential

Why: nothing revokes a root console session; it lives until it expires. The containment is a new password and a new MFA device, then watching for userIdentity.type = 'Root' after that time. With centralised root access enabled on the organisation, the management account can do it without the root password by assuming root with the IAMDeleteRootUserCredentials task policy (15-minute session) and deleting the login profile and MFA. Otherwise it is the root email's recovery flow — and if the attacker changed the email, AWS Support with proof of ownership.
# from the management account, if enable-organizations-root-sessions was done in peacetime
aws sts assume-root --target-principal <member-account-id> --duration-seconds 900 \
  --task-policy-arn arn=arn:aws:iam::aws:policy/root-task/IAMDeleteRootUserCredentials
# with those credentials:
aws iam delete-login-profile
aws iam list-mfa-devices && aws iam deactivate-mfa-device --serial-number <arn>
aws iam list-access-keys && aws iam delete-access-key --access-key-id <AKIA...>
# then, as root of the member account, set a new password and MFA; verify who is root now:
aws account get-contact-information && aws account get-alternate-contact --alternate-contact-type SECURITY

T+30–60 · IsolateMemory, then snapshot, then isolate — in that order, and don't terminate

Why: the isolation security group cuts SSM too unless the VPC has interface endpoints for ssm, ssmmessages and ec2messages and the isolation group allows 443 to them, so take memory while the agent can still talk. AVML writes LiME format; stage it in the evidence bucket (write-only for the instance role) rather than fetching it from the internet. Snapshots encrypted with the default aws/ebs key cannot be shared to another account: copy to a customer-managed key whose policy includes the evidence account, then share. Security-group changes don't cut connections already tracked — add a subnet NACL deny if the cut must be immediate. Keep the instance profile attached until the memory upload has finished; credentials already fetched from the metadata service stay valid until the role's sessions are revoked.
# 1. memory, while the network and the instance role are still there (Windows: AWS-RunPowerShellScript with a memory tool staged the same way)
aws ssm send-command --region <region> --instance-ids <i-id> --document-name AWS-RunShellScript --comment "IR <case> memory" \
  --parameters 'commands=["aws s3 cp s3://<evidence-bucket>/tools/avml /tmp/avml","chmod +x /tmp/avml","/tmp/avml --compress /tmp/<i-id>.lime.compressed","aws s3 cp /tmp/<i-id>.lime.compressed s3://<evidence-bucket>/<case>/","sha256sum /tmp/<i-id>.lime.compressed"]'
# 2. disk
aws ec2 create-snapshot --volume-id <vol-id> --description "IR <case> <i-id>" \
  --tag-specifications 'ResourceType=snapshot,Tags=[{Key=ir-case,Value=<case>}]'
aws ec2 copy-snapshot --source-region <region> --source-snapshot-id <snap-id> --encrypted --kms-key-id <evidence-cmk-arn>   # aws/ebs-encrypted snapshots cannot be shared
aws ec2 modify-snapshot-attribute --snapshot-id <copied-snap-id> --attribute createVolumePermission --operation-type add --user-ids <evidence-account>
# 3. isolate
aws ec2 modify-instance-attribute --instance-id <i-id> --groups <isolation-sg>
aws ec2 modify-instance-attribute --instance-id <i-id> --disable-api-termination
aws ec2 disassociate-iam-instance-profile --association-id <iip-assoc-id>

By Scenario: What Changes

The first-72-hours flow above is the universal path. These are the deltas — what you do differently, and first, when the incident is shaped like this.

1. Leaked Access Key or Service Credential

Indicators: a key found in a public repository, image or paste; a provider notice about an exposed key; API calls from an IP or user agent the owner never uses.

Variation: Assume it was used — public keys are harvested by scanners within minutes; aws iam get-access-key-last-used says when and from which service even before the trail is queried. For keys exposed on GitHub, AWS typically attaches its AWSCompromisedKeyQuarantineV3 managed policy within seconds, opens a support case and raises an AWS Health event of type AWS_RISK_CREDENTIALS_EXPOSED. Leave that policy in place — AWS says not to remove it — but it only denies a set of high-risk actions: it does not revoke the key, so deactivate it and deny the user anyway. GCP disables a leaked service-account key itself when the organisation policy iam.serviceAccountKeyExposureResponse is DISABLE_KEY (the default since June 2024) and writes an audit-log event plus an email to the project owners and security contacts; under WAIT_FOR_ABUSE it only notifies — check which one you run. Then run the full persistence hunt for everything the key could have created, and remove the key from the repository's history — deleting the file leaves it in every clone.

2. Cryptomining / Resource Abuse

Indicators: a cost anomaly alert, GPU or large instances in regions you don't use, service-quota increase requests you didn't make, GuardDuty or Defender crypto-currency findings, a subscription or project on your billing account that nobody created.

Variation: It looks like a billing problem and it is an access problem: the miner is only the visible part. Check every region and every shape — fleets, SageMaker, ECS/EKS, scale sets, AKS, ML compute, GKE, Dataproc — and since 2024 the model-access variant: Bedrock entitlements, Azure OpenAI deployments and Vertex AI jobs resold as LLM access. Look where the bill hides: an attacker with billing rights creates their own account, subscription (Microsoft.Subscription/aliases/write) or project (UpdateProjectBillingInfo seen with --billing-account) and runs the miners there, outside every folder you watch. Snapshot one instance for evidence, then stop the rest — cost is the harm here. The credential that launched them still gets the full hunt; miners are often not the only thing a stolen key was used for. Open the billing case with the provider early.

3. Data Exposed or Taken from Storage

Indicators: a bucket or container made public (allUsers, a removed public-access block), a snapshot or image shared with an unknown account, storage keys listed, a disk SAS from beginGetAccess, an extortion email with sample files.

Variation: Close the exposure first, then work out what could be read. Azure SAS tokens signed with an account key can't be revoked one by one — rotating both account keys is the only way to kill them; user-delegation SAS die when the delegation keys are revoked. Remove external snapshot permissions and look at whether the other account already copied them. If object-level logging was off (S3 data events, StorageBlobLogs, GCP Data Access), notification is scoped on what was reachable, not what was proven read — and the status line says which.

aws ec2 modify-snapshot-attribute --snapshot-id <snap-id> \
  --attribute createVolumePermission --operation-type remove --user-ids <external-account>
gcloud storage buckets remove-iam-policy-binding gs://<bucket> --member=allUsers --role=roles/storage.objectViewer

4. Storage Held Hostage or Destroyed

Indicators: mass object deletion, objects rewritten with a customer-provided key (SSE-C), a new lifecycle or replication rule, snapshots and backups deleted, a key scheduled for deletion, a ransom note object in a bucket.

Variation: With SSE-C the provider never holds the key — there is nothing to recover it from, so recovery comes from object versions, Object Lock, or backups the attacker couldn't reach. The quiet versions are deletion by rule and deletion by key: a lifecycle rule that expires everything (remove the rule — objects already expired are only recoverable if versioning kept noncurrent versions), a replication rule to the attacker's bucket, and a KMS key scheduled for deletion, which makes every object it wrapped unreadable when the 7–30 day window ends — aws kms cancel-key-deletion while it is pending. Azure: Recovery Services soft delete keeps deleted backup data 14 days and a Resource Guard makes turning it off a multi-user operation; Key Vault purge protection stops a purge inside the 7–90 day soft-delete window. GCP: soft delete keeps deleted objects 7 days by default (gcloud storage ls --soft-deleted, gcloud storage restore) and a Cloud KMS version in DESTROY_SCHEDULED can be restored for at least 24 hours. Stop the principal before protecting the backups; then follow Ransomware Response, Scenario 6 for the extortion and board decisions.

aws s3api list-object-versions --bucket <bucket> --prefix <path> --max-items 20
aws s3api get-bucket-lifecycle-configuration --bucket <bucket>; aws s3api delete-bucket-lifecycle --bucket <bucket>
aws kms cancel-key-deletion --key-id <key-id>
gcloud kms keys versions restore <version> --key=<key> --keyring=<ring> --location=<loc>

5. Compromised Workload

Indicators: an instance or container running unexpected processes, SSRF against the metadata service, GuardDuty InstanceCredentialExfiltration findings (the instance role's credentials used from outside AWS).

Variation: Two containments, not one. The host: memory first, through SSM, az vm run-command or gcloud compute ssh while the agent channel still works, then snapshot, then isolate — and keep it running. The identity: the role's credentials may already be in use from elsewhere, so revoke the role's sessions and hunt what it did, even after the host is isolated. Flow logs are the only record of what left the host. Enforce IMDSv2 before anything is redeployed.

6. Compromised CI/CD or Deployment Identity

Indicators: infrastructure changed outside a merged change, a pipeline role used from an unknown runner, an OIDC trust policy that accepts any repository or branch, a workflow on pull_request_target that runs fork code with secrets.

Variation: Where humans are read-only, the pipeline identity is the most privileged principal you have — treat it like a domain admin. Pause the pipelines, revoke the role's sessions, and narrow the trust before re-enabling: on AWS the OIDC condition on token.actions.githubusercontent.com:sub must name the exact repository and branch or environment and :aud must be sts.amazonaws.com; on GCP the workload identity pool provider needs an attribute condition such as attribute.repository == "org/repo"; on Entra a federated credential's subject and issuer are the trust, and its addition is logged only as Update application with FederatedIdentityCredentials in the modified properties. Treat the Terraform or Pulumi state as a secret store that was read: it holds database passwords, keys and tokens in plain text, so rotate everything in it, not only the pipeline's own credential. Then diff deployed state against the code: anything the attacker deployed through the pipeline looks like a legitimate change.

7. Root, Management Account or Tenant Admin Compromised

Indicators: root user sign-in or AssumeRoot, root email, billing contact, alternate contact or payment method changed, MFA devices changed, accounts closed or leaving the organisation, SCPs detached, Global Administrator activity nobody owns, an elevateAccess entry, a partner or delegation nobody set up.

Variation: This is the one incident where your own guardrails don't help: SCPs never apply to the management account, and a Global Admin can remove every other control. AWS: there is no session revoke for root — replace the password and MFA, then watch for userIdentity.type = 'Root' after that time. With centralised root access enabled, the management account recovers a member account's root without its password (aws sts assume-root with the IAMDeleteRootUserCredentials task policy). Hunt the account itself: StartPrimaryEmailUpdate / AcceptPrimaryEmailUpdate, PutContactInformation and PutAlternateContact on account.amazonaws.com, SetAccountPreferences and SetAdditionalContacts on billingconsole.amazonaws.com, payment instruments on payments.amazonaws.com, and in the organisation CloseAccount, LeaveOrganization, RemoveAccountFromOrganization, DetachPolicy and DisablePolicyType. The attacker's other route is a support case: a convincing “lost MFA” account-recovery request to the provider — keep the alternate security contact current and tell Support the account is under incident response so recovery requests are challenged. Azure: an elevateAccess entry (Entra AuditLogs, service “Azure RBAC (Elevated Access)”) makes a Global Admin owner of every subscription; check it, then the delegations that reach your tenant from outside — GDAP and CSP relationships (Get-MgTenantRelationshipDelegatedAdminRelationship) and Lighthouse (az managedservices assignment list). Use the break-glass accounts, which are excluded from Conditional Access and monitored, not the compromised admin. For Entra ID admin compromise, follow Identity Breach Response in parallel. GCP: the organisation administrator and billing-account roles are the equivalent; search-all-iam-policies at organisation scope for roles/resourcemanager.organizationAdmin and roles/billing.admin says who holds them now.