Why This Is Loop 0
| Threat | Without Immutable Backups | With Immutable Backups |
|---|---|---|
| Ransomware encrypts everything | Pay $613K or lose data | Restore from immutable copy. Don't pay. |
| BEC — attacker deletes data | Data gone | Restore from immutable copy |
| Failed patch breaks production | Scramble to fix or roll back manually | Restore pre-patch state in hours |
| Supply chain compromise | Compromised state is the only state | Restore to pre-compromise baseline |
| Config drift — need to roll back | Hope someone remembers the old config | Restore known-good config from backup |
| Insider threat — data destruction | Permanent loss | Immutable = even admin can't delete within retention |
| Cyber insurance claim | Denied — 82% of denied claims lack proper backup | Evidence of immutable, tested backups = approved |
How Attackers Kill Your Backups
| Method | How It Works | What Stops It |
|---|---|---|
| Domain-joined backup server | Backup server on same AD domain. Domain admin = backup admin. One compromise owns everything. | Separate backup credentials. No domain join for backup infra. |
| Shadow copy deletion | vssadmin Delete Shadows /all /quiet — near-universal in ransomware, runs in seconds | Off-host backups. VSS alone is not a backup strategy. |
| Credential theft | Backup admin creds found in NTDS.dit, Kerberoasting of service accounts, plaintext on staging servers | Separate, unique backup admin accounts. MFA. No password reuse. |
| API-based deletion | Stolen cloud API keys used to modify retention policies or delete snapshots | S3 Object Lock Compliance mode — even root can't delete within retention. |
| Backup software takeover | Attacker controls Veeam/Commvault console, deletes everything through the UI | Immutability at storage layer, not application layer. App compromise can't bypass WORM. |
| Kill backup services | Disable backup agents/services before encryption. Standard operating procedure. | Backups already written to immutable storage. Killing the agent doesn't delete what's already there. |
Entry Point: Backup Resilience Assessment
Assessment Checklist
- Backup existence: What's backed up? What isn't? Is M365 data backed up or just relying on retention policies?
- Immutability check: Can a domain admin delete your backups right now? If yes, they're not immutable.
- Credential separation: Are backup admin credentials the same as domain admin? Same password? Same MFA?
- Network isolation: Is backup infrastructure on the same VLAN/subnet as production? Can ransomware on a workstation reach the backup server?
- Restore test: Pick a critical system. Actually restore it. Time it. Compare against your stated RTO. Any RTO claim that hasn't been validated by a timed restore is fiction.
- Retention depth: How far back can you restore? Dwell time can be 14+ days — if your oldest backup is 7 days, you're restoring to a compromised state.
- M365 backup: Are Exchange, SharePoint, Teams, OneDrive backed up independently? Native retention policies are NOT backup.
Output
Backup resilience scorecard: what's protected, what's immutable, what's tested, what's the actual (not theoretical) RTO. Insurance gap analysis for backup controls specifically.
- Change backup admin passwords to something unique — not shared with domain admin, not in a password vault that's on the domain
- Enable MFA on backup admin accounts — separate MFA from production MFA (different authenticator app or hardware key)
- If using cloud backup, check that API keys have minimum permissions — write-only for backup agents, separate admin key for restore/delete
- Disable the ability to delete backup jobs from the backup management console without separate authorization
- Create a manual offline copy of the most critical data right now — external USB drive, locked in a safe. Ugly but air-gapped.
- Dump databases before backing up — backing up running DB volumes risks corruption. Add
pg_dump/mysqldumppre-backup hooks.
- Credential audit: Map every account that can access, modify, or delete backup data. Ensure zero overlap with domain admin accounts.
- Immediate S3 Object Lock: If backups go to S3-compatible storage, enable Object Lock in Governance mode today (can upgrade to Compliance mode after validation).
- Network isolation: Move backup server to a separate VLAN if possible. At minimum, restrict inbound access to only the backup agents.
Zero shared credentials between production and backup. Backup infrastructure unreachable from compromised workstations. Immediate offline copy exists for worst-case scenario.
- Alert on VSS shadow copy deletion — Windows Event ID 8225 or Sysmon process create for
vssadmin.exe delete - Alert on backup service stop/disable — Windows Event ID 7036 for backup agent service state changes
- Alert on backup job failures — most backup tools can email on failure. Enable it. Check the email account is monitored.
- Monitor backup storage capacity — sudden drops in backup size may indicate deletion. Sudden spikes may indicate encryption of source data.
- Wazuh rules for backup integrity: FIM on backup config files, alerts on backup credential access, process monitoring for shadow copy deletion tools
- Canary files in backup repos — if a canary file disappears from backup, someone is tampering
- Backup completion dashboard — single pane showing: last successful backup per system, backup size trend, any failed jobs
Any attempt to delete, modify, or disable backups generates an immediate alert. Backup health visible at a glance. No silent failures.
- Install Kopia or Restic — free, encrypted, deduplicated backups with S3 Object Lock support
- Set up BorgBackup in append-only mode on a Linux server — clients can create archives, cannot delete
- Configure backup credentials as write-only / append-only — the backup agent can push data but cannot delete or modify existing backups
- Primary backup: Kopia or Restic to on-prem storage (MinIO with Object Lock, or NAS with WORM snapshots). Fast restores.
- Secondary backup: Replicate to cloud with S3 Object Lock Compliance mode (Wasabi $6.99/TB or Backblaze B2 $6/TB). Geographic redundancy.
- Tertiary backup (optional): Quarterly LTO tape or disconnected USB rotation to physical safe. True air gap for catastrophic scenarios.
- Backup VLAN: Dedicated network segment for backup infrastructure. No inbound access from workstations. Firewall rules restrict traffic to backup agents only.
- Separate credentials: Backup admin accounts not in AD. Different passwords, different MFA. Stored in a separate password manager or physical safe.
- Retention policy: Minimum 30 days for daily backups (covers typical dwell time). 90 days for weekly. 1 year for monthly. Retention MUST exceed typical attacker dwell time (14+ days).
- Pre-backup hooks: Database dumps (
pg_dump,mysqldump) before filesystem backup. Never back up running database files directly.
Reference Architecture (200 seats)
[200 Endpoints + Servers]
|
| (kopia/restic client, write-only credentials)
v
[Backup Server: kopia server / restic-rest-server --append-only]
(Dedicated VLAN, no domain join, separate admin creds)
|
+--> [Primary: MinIO on-prem, Object Lock Compliance]
| (fast restore, on-site)
|
+--> [Secondary: Wasabi/B2, Object Lock Compliance]
| (geographic redundancy, offsite)
|
+--> [Tertiary: LTO tape / USB in safe]
(true air gap, quarterly rotation)
3-2-1-1 fully implemented. Primary on-prem for fast restore. Secondary cloud for geographic redundancy. Both with Object Lock Compliance — even root can't delete within retention. Backup infrastructure on isolated VLAN with separate credentials.
- Set a monthly calendar reminder: restore one random file per backed-up system. Verify checksum.
- Check backup storage capacity trend — is it growing as expected or has it plateaued (indicating silent failure)?
- Verify backup admin credentials still work — and that they're still separate from domain admin
| Cadence | Test | What It Validates |
|---|---|---|
| Monthly | File-level restore — 10 random files, verify checksums | Data integrity, backup completeness |
| Quarterly | Application restore — full app stack (DB + app + config) to isolated environment, verify it works | RTO claim, application recoverability |
| Quarterly | Credential review — backup admin accounts, API keys, permissions audit | Credential separation, least privilege |
| Annually | Full DR simulation — simulate total infrastructure loss, restore from immutable only, time entire process | Organizational readiness, actual RTO |
| Annually | Adversarial test — can a compromised domain admin delete backups? Can stolen API keys modify retention? | Immutability guarantee under attack |
Monthly file restores passing. Quarterly full-app restores validating RTOs. Annual DR simulation documenting actual recovery time. Adversarial testing confirming immutability holds under attack.
- Write the Recovery Priority List: Tier 0 (AD/identity — nothing works without this), Tier 1 (payroll — almost always #1 for employees), Tier 2 (ERP/core business), Tier 3 (everything else). Email it to IT and management.
- Print the backup admin credentials and store in a physical safe (not in the password manager that might be encrypted)
- DR runbook: Step-by-step recovery per tier. Named roles. Tested quarterly. Action cards on laminated paper, not a 60-page PDF on the encrypted file server.
- M365 backup solution: Native retention is NOT backup. Deploy commercial M365 backup (~$2–4/user/month) for Exchange, SharePoint, Teams, OneDrive with independent retention.
- Immutable infrastructure: Where possible, rebuild from code/image rather than restore from backup. Infrastructure-as-code for servers means faster reprovisioning.
- Insurance evidence package: Document immutable backup architecture, retention policies, restore test results, credential separation. This is the evidence that gets claims approved.
- NIS2 Article 21 alignment: Backup management and disaster recovery are explicitly mandated. Documented 3-2-1-1 + tested DR plan = compliance artifact.
- Board reporting: Quarterly backup health report: systems covered, last successful backup, last tested restore, RTO status, immutability status. One page.
- Separate management AD (Red Forest): All privileged accounts (backup admin, hypervisor admin, network admin) live in a separate AD forest. Production domain trusts management domain (one-way) — management admins can manage production resources, but production cannot query the management domain. Attackers who compromise production AD cannot see who the admins are, where they live, or target them.
net group "Domain Admins" /domainon production returns nothing useful — the real admins don't exist in that directory. BloodHound against production can't map the attack path. Kerberoasting and NTDS.dit extraction yield zero management credentials. The admins are invisible to reconnaissance. - Log retention architecture: 365 days on every server. Increase Windows Event Log to 1GB+ per channel (Security, System, PowerShell). Linux syslog: 12 months. Forward to central logging on isolated infrastructure AND keep on-host copies — if central logging is compromised, on-host logs are your forensic backup. 1TB disk costs less than a coffee per day.
Recovery is a tested, documented, practiced process — not a hope. Admin accounts invisible to attackers via management AD. Logs retained 365 days. Any scenario (ransomware, insider, failed patch, supply chain compromise) has a recovery path with a validated RTO. Insurance approved. NIS2 compliant. Board informed.
Program Economics (200-seat reference, 50TB data)
| Engagement | Duration | Investment |
|---|---|---|
| Assessment + immediate hardening | 1 day | €1,250 |
| 1 Containment | 1 day | €1,250 |
| 2 Detection | 1 day | €1,250 |
| 3 Posture (3-2-1-1 build) | 3–4 days | €3,750 – 5,000 |
| 4 Vuln Mgmt setup | 1 day + cadence | €1,250 + €2,500/yr |
| 5 Structural | 2–3 days | €2,500 – 3,750 |
| Full program | 9–11 days | €11,250 – 13,750 + retainer |
Ongoing Storage Costs
| Component | Option | Monthly Cost |
|---|---|---|
| Backup software | Kopia / Restic / Borg (FOSS) | €0 |
| On-prem immutable storage | MinIO on existing hardware | €0 (amortized HW) |
| Cloud immutable (50TB) | Wasabi ($6.99/TB) | ~€350/mo |
| Backblaze B2 ($6/TB) | ~€300/mo | |
| Hetzner Storage Box (no WORM) | ~€36/mo | |
| M365 backup (200 users) | Commercial ($3/user/mo) | ~€600/mo |
What "Tested" Actually Means
| Level | Test | Cadence | Time | Validates |
|---|---|---|---|---|
| 1 | File restore — 10 random files, verify checksums | Monthly | ~1 hour | Data integrity, backup completeness |
| 2 | App restore — full stack to isolated env, verify it works | Quarterly | 4–8 hours | RTO claim, application recoverability |
| 3 | Full DR simulation — total loss, restore everything, time it | Annually | 1–3 days | Organizational readiness, actual RTO |
| 4 | Adversarial test — attempt to delete backups as compromised admin | Annually | Half day | Immutability holds under attack |
Free Backup Tools
- Kopia — Best all-around: GUI, server mode, S3 Object Lock, multi-platform
- Restic — Proven, broadest backend support, largest community
- BorgBackup — Best dedup/compression, append-only, Linux/SSH focus
- MinIO — Self-hosted S3 with Object Lock (Compliance mode)
Immutable Storage Providers
- Wasabi — $6.99/TB, free egress, S3 Object Lock
- Backblaze B2 — $6/TB, S3 Object Lock, no min retention
- Hetzner — ~EUR 6/TB, EU data residency, WORM
- MinIO (self-hosted) — Free, Object Lock Compliance, you own the hardware
Tool Comparison
| Feature | Kopia | Restic | Borg |
|---|---|---|---|
| GUI | Built-in web + desktop | 3rd-party | Vorta (3rd-party) |
| Windows | Native | Native | WSL only |
| S3 Object Lock | Yes | Yes | No (SSH only) |
| Server mode | Built-in multi-user | REST server | SSH per-host |
| Deduplication | Content-addressed | Content-addressed | Best-in-class |
| Compression | Multiple algos | None built-in | 4 algorithms |
| Maturity | Good | Excellent | Excellent |
| Best for | 200-seat mixed env | Cloud-first | Linux-only |