The Anti-Vendor Philosophy
Centralized logging exists for one purpose: active detection and investigation. That needs 30–90 days of normalized, searchable data. Not 365 days. Not 7 years. If you need to investigate something from 8 months ago, you pull the raw logs from the machine or restore them from backup. You don't need them sitting in Elasticsearch burning RAM 24/7.
The Three Copies You Already Have
| Copy | Where | Retention | Purpose | Cost |
|---|---|---|---|---|
| 1. On-host | The machine that generated the log | 365 days (compressed) | Forensic raw data. If central logging is compromised, this is your backup. If you need to investigate something older than 90 days, you come here. | 1TB disk per server = ~€3/month. Compressed logs take a fraction of that. |
| 2. Cloud provider native | AWS CloudTrail, Azure Activity Log, Entra sign-in logs, Google Cloud Audit Logs | 30–90 days free | Cloud providers already store your logs for free. CloudTrail = 90d management events. Azure Activity Log = 90d. Entra = 30d. You only pay when you export or query. Leave them where they are. | €0 — already included in your cloud subscription. Don't export unless you need to. |
| 3. Immutable backup | Your backup infrastructure (Borg/Restic/Kopia → S3 Object Lock) | As long as your backup retention (30d–1yr+) | Disaster recovery. If the machine is destroyed (ransomware, hardware failure), the logs survive in backup. You're already paying for this. | €0 incremental — logs are already in the machine backup. |
| 4. Central/SIEM | Wazuh, Elastic, Splunk, Grafana Loki, etc. | 30–90 days (normalized, indexed, searchable) | Active detection. Alert rules run here. Investigation happens here. This is the only copy that needs to be fast and searchable. | Self-hosted Wazuh: €0 software + server hardware. Cloud SIEM: €€€. |
What You Already Have for Free vs The Gap
Already Covered (Free)
| Source | Where It Lives | Retention (Free) | What It Covers |
|---|---|---|---|
| Windows Event Logs | Every Windows machine | 365d if you increase channel size (1 GPO change) | Authentication, process creation, service changes, PowerShell, group membership |
| Linux syslog | Every Linux server | 365d with logrotate config | Auth, sudo, cron, kernel, service state |
| AWS CloudTrail | AWS native | 90d management events free | Every API call: IAM changes, S3 access, EC2 lifecycle, security group changes |
| Azure Activity Log | Azure native | 90d free | Control plane operations: role assignments, policy changes, resource modifications |
| Entra ID Sign-in Logs | Microsoft 365 | 30d free (7d for free tier) | Every authentication: who, from where, MFA status, CA policy evaluation, risk level |
| Entra ID Audit Logs | Microsoft 365 | 30d free | Admin actions: user/group/app/policy changes |
| Google Cloud Audit Logs | GCP native | 400d admin activity (free), 30d data access (paid) | API calls, IAM changes, resource access |
| Google Workspace Audit | Google Workspace | 180d (6 months) | Admin actions, login events, Drive activity, Gmail events |
| Exchange Online Audit | M365 | 180d (E5) / 90d (E3) | Mailbox access, send-as, forwarding rules, admin changes |
| Firewall / VPN logs | On the appliance | Varies (typically 7–90d on disk) | Connection logs, auth events, rule changes, admin access |
| Immutable backup | Borg/Restic/Kopia | Your backup retention (30d–1yr) | ALL of the above — if the machine is backed up, the logs are backed up |
The Actual Gap
| Gap | What's Missing | Why It Matters | How to Fill It |
|---|---|---|---|
| Cross-source correlation | Logs exist in 10+ places but can't be queried together. "Show me all activity from this IP across AD, email, VPN, and cloud" is impossible. | Attackers move across systems. You need to follow them across log sources to understand the full attack chain. | Central SIEM (Wazuh) with 30–90d retention. Forward only the high-value sources. |
| Real-time alerting | On-host and cloud native logs are stored but not monitored. Nobody gets an alert when a shadow copy is deleted at 2 AM. | 88% of ransomware deploys outside business hours. Detection without alerting is just evidence collection. | 15–30 detection rules in Wazuh/Loki with phone/SMS/Signal alert routing. Not email. |
| Normalization | Windows events, Linux syslog, cloud JSON, and firewall logs all use different formats. Can't compare apples to apples. | Investigation speed. When you're in an incident, you don't have time to manually parse three different log formats. | Wazuh normalizes automatically. Or use Elastic Common Schema / OpenTelemetry for DIY. |
| On-host defaults are too small | Windows Event Log at 20MB overwrites in hours. Linux default logrotate keeps 4 weeks. Firewall disk fills up. | When you need forensic data from 3 months ago, it's already gone. The attacker's dwell time was longer than your retention. | 1 GPO change (Windows 1GB+), 1 config edit (Linux 365d), external syslog for firewalls. All free. |
| No coverage verification | Nobody checks if all machines are actually logging, if agents are alive, if retention is holding. | You think you have logs. During the breach investigation, you discover the agent stopped 3 months ago. | Monthly coverage check. Wazuh agent heartbeat monitoring. 30-minute investment per month. |
Beyond Machines: Containers, Kubernetes, Serverless, SaaS
Containers & Docker
| What You Already Have | Where It Lives | Retention | The Gap |
|---|---|---|---|
| Container stdout/stderr | Docker json-file log driver on the host at /var/lib/docker/containers/ |
Default: unlimited (fills disk). Fix: "log-opts": {"max-size":"100m","max-file":"10"} in daemon.json = ~1GB per container on host. |
Logs are per-container, not searchable across containers. Need Loki/Wazuh for cross-container correlation. |
| Docker events | docker events stream on the host |
Ephemeral (not stored by default). Forward to syslog or Wazuh. | Container start/stop/kill/die events = security-relevant. Who started that container at 3 AM? |
| Compose/stack logs | docker compose logs aggregates all services |
Same as container log driver settings | Only accessible on the host. No remote search without central forwarding. |
{"log-driver":"json-file","log-opts":{"max-size":"100m","max-file":"10"}} in /etc/docker/daemon.json on every host. This keeps ~1GB of logs per container on-host for free. The immutable backup of the host captures these.Kubernetes
| What You Already Have | Where It Lives | Retention | The Gap |
|---|---|---|---|
| Pod logs (stdout/stderr) | Node filesystem at /var/log/containers/ and /var/log/pods/ |
Until kubelet rotates (default 10MB per container, 5 files). Configurable via --container-log-max-size and --container-log-max-files. |
Pods are ephemeral. When a pod dies, the node-level logs eventually get cleaned. Need to forward before that happens. |
| Kubernetes audit logs | API server (self-managed) or cloud provider (managed K8s) | EKS: 90d in CloudWatch (if enabled). AKS: 90d in Azure Monitor. GKE: 400d admin activity free. | Critical for security: who created that privileged pod? Who modified RBAC? Cloud-managed K8s stores this for free — enable it. |
| Node-level system logs | Same as any Linux server — journald/syslog on the node | Standard Linux retention (365d with logrotate) | Nodes are the "on-host" layer for Kubernetes. Same principle applies. |
Serverless & Managed Services
| Service | Free Logging | Retention | What You Get for Free |
|---|---|---|---|
| AWS Lambda | CloudWatch Logs (automatic) | Indefinite (you pay for storage after free tier: 5GB) | Every invocation, duration, errors, stdout. Set a retention policy or it grows forever. |
| Azure Functions | Application Insights (built-in) | 90d default (configurable) | Invocations, dependencies, exceptions, traces. |
| GCP Cloud Functions | Cloud Logging (automatic) | 30d default (_Default bucket, configurable) | Execution logs, errors, cold starts. |
| RDS / CloudSQL / Azure SQL | Audit logs + slow query logs | Provider-dependent (typically 7–90d free) | Query logging, connection events, schema changes. Enable it — usually off by default. |
| S3 / Azure Blob / GCS | Access logging | Stored in a target bucket (you pay for storage) | Every read/write/delete on your storage. Critical for data exfiltration detection. |
| API Gateway / CloudFront / Azure Front Door | Access logs | Provider-dependent | Every API call: source IP, path, status, latency. Enable it. |
SaaS Applications
| SaaS | Free Audit Log Retention | Access Method | What to Watch For |
|---|---|---|---|
| Microsoft 365 | 90d (E3) / 180d (E5) | Unified Audit Log, Graph API | Mailbox forwarding rules, OAuth consents, admin role changes, file sharing |
| Google Workspace | 180d (6 months) | Admin Console > Reports, Reports API | Admin actions, login events, Drive sharing, OAuth tokens |
| Salesforce | Event Monitoring (add-on) or Login History (free, 6 months) | EventLogFile API, Setup Audit Trail | Login anomalies, data exports, report runs, API access |
| Slack | Enterprise Grid: audit logs via API. Free/Pro: no audit API. | Audit Logs API (Enterprise only) | File uploads, app installations, channel creation, guest access |
| GitHub | Enterprise: 180d audit log. Org: limited. | Audit log UI + REST/GraphQL API | Repository access, branch protection changes, SSH key additions, Actions secrets |
| Okta / Auth0 | 90d System Log | System Log API | Auth failures, MFA bypasses, app assignments, admin changes |
The Pattern Is the Same
- Increase Windows Event Log sizes — defaults are 20MB per channel. At 20MB, Security logs overwrite in hours on a busy server. Set to 1GB minimum per channel (Security, System, PowerShell/Operational) via GPO.
- Linux syslog retention — configure logrotate to keep 365 days compressed. Default is usually 4 weeks.
- Enable PowerShell Script Block Logging (GPO) — ransomware operators use PowerShell extensively. Without this, you're blind.
- Enable Windows command-line process auditing — shows exactly what commands attackers ran. GPO: Audit Process Creation → Include command line.
- IR log cheat sheet: Laminated card for each system type showing exactly which logs to pull and where they live during an incident. Not a 40-page document — a card.
- Pre-built log collection scripts: PowerShell/Bash scripts that grab the critical logs from a compromised machine, compress them, and copy to a safe location. Ready before you need them.
- Ensure on-host logs survive reboot and clearing: Forward critical events to a secondary local log file (Windows Event Forwarding to local collector, or rsyslog to a second file) so even if an attacker clears Event Viewer, the copy persists.
365 days of raw logs on every server. IR team can pull any log from any machine within minutes. Logs survive attacker clearing attempts via secondary local copy.
- Deploy Wazuh (free, open-source) as your SIEM. Agents on all servers and critical workstations. Receives logs, normalizes, indexes, alerts.
- Or deploy Grafana Loki — lightweight log aggregation. Doesn't index log content (cheaper), queries via labels. Good for mid-market that doesn't need full SIEM.
- Set retention to 90 days in your SIEM. Not 365. Not "as much as we can fit." 90 days covers the detection window for most attacks (median dwell time is 3–14 days). Older data lives on-host and in backup.
- Alert on the things that matter (from the Config Drift playbook's 15 settings):
• Legacy auth attempts (Event ID 4776 with legacy protocol)
• Shadow copy deletion (Event ID 8225 or vssadmin process creation)
• Backup service stopped (Event ID 7036)
• RDP from unexpected sources
• New admin group membership (Event ID 4732)
• Off-hours admin activity (88% of ransomware deploys at night)
- Detection rule library: Tuned to the 15 breach-forensic settings. Not 500 generic rules — 15–30 high-signal rules that detect the specific attack patterns from real incidents.
- Alert routing: Critical alerts go to phone/SMS/Signal 24/7. Not just an email that nobody reads on Saturday night.
- Log source prioritization: Not all sources are equal. Prioritize: domain controllers, VPN/firewall, email gateway, backup servers, file servers. Workstations are secondary.
- Retention tier configuration: Hot (30d, fully indexed), Warm (60d, compressed index), Delete at 90d. Older data is on-host and in backup.
30–90 days of searchable, normalized logs in SIEM. 15–30 high-signal detection rules running 24/7. Alerts reach humans at any hour. The SIEM is sized for detection, not storage.
- Alert on log source silence — if a machine that normally sends 1,000 events/day suddenly sends 0, something is wrong. Wazuh has built-in agent heartbeat monitoring.
- Monthly log coverage check — how many machines should be logging vs how many actually are? The delta is your blind spot.
- Verify on-host log retention is actually at 365 days — GPO settings can be overridden locally, logrotate configs can change.
- Log coverage dashboard: Single view showing every machine, its last log event, agent health, and on-host retention status.
- Automated enrollment: New machines auto-deploy the Wazuh agent via Intune/GPO/Ansible. No manual step = no machines missed.
- Detection rule health: Are your alert rules actually firing? A rule that never triggers is either perfect coverage or a broken rule. Test monthly.
100% log coverage verified continuously. Zero silent agent failures. Detection rules tested monthly. New machines automatically enrolled.
| Cadence | Action |
|---|---|
| Continuous | Agent heartbeat monitoring (silent source detection) |
| Monthly | Log coverage check (expected vs actual sources) |
| Monthly | Storage usage trending (is retention holding at 90d or growing unbounded?) |
| Quarterly | Detection rule testing (do rules actually fire when they should?) |
| Quarterly | On-host retention verification (is 365d GPO still applied?) |
| Annually | SIEM right-sizing review (are you over-collecting? under-detecting?) |
The Right Architecture for Mid-Market
[Every Server + Critical Workstation]
|
|-- On-host: 365d compressed logs (1GB+ Event Log channels, logrotate 365d)
| Cost: ~EUR 3/month per server (disk)
|
|-- Immutable backup: Logs included in machine backup (Borg/Restic/Kopia)
| Cost: EUR 0 incremental (already backing up the machine)
|
+-- Forward to central: Wazuh agent / rsyslog / Filebeat
|
v
[Central SIEM: Wazuh / Grafana Loki]
Retention: 30-90 days normalized
Purpose: Detection rules + investigation
NOT for: Long-term storage (that's what copies 1+2 are for)
Cost: 1 server, 4-8 cores, 16-32GB RAM, 500GB-1TB SSD
= EUR 50-150/month self-hosted or existing hardware
- SIEM deployment: Single Wazuh server sized for your environment. Not a cluster. Not a managed cloud SIEM at €10K/month. One server that does the job.
- Log source isolation: SIEM on its own VLAN with separate admin credentials. If an attacker compromises production, they can't delete the SIEM's copy of the evidence. But remember — you also have on-host and backup copies.
- Tiered retention: SIEM at 30–90d, on-host at 365d, backup retention per your DR policy. Each tier serves a different purpose. Don't pay SIEM prices for backup-tier storage.
- NIS2 alignment: Article 21 requires incident handling capability. A functioning SIEM with detection rules + 365d on-host logs + immutable backups = compliance. You don't need Splunk Enterprise to satisfy the auditor.
- Insurance evidence: "Do you have logging and monitoring?" Yes — centralized SIEM with detection rules, on-host retention, and backup. Here's the dashboard. Here are the alert rules. Here's the last incident we detected and investigated.
Simple, right-sized logging infrastructure. One SIEM server. 30–90 days searchable. 365 days on-host. Backup covers everything. Detection rules tuned to real threats. No over-engineering. No vendor lock-in. No surprise bills.
What to Actually Log (and What Not To)
Must Centralize (High Detection Value)
| Source | Why | Key Events |
|---|---|---|
| Domain Controllers | Every authentication, every group change, every policy modification passes through here. Attackers reach AD in 3.4 hours average. | 4624/4625 (logon), 4720 (account created), 4732 (group membership), 4768/4769 (Kerberos), 1102 (audit log cleared) |
| VPN / Firewall | 71% of initial access is via external remote services. This is where attacks start. | Connection logs, auth failures, rule changes, admin access |
| Email Gateway | BEC is 81% of investigated incidents. Phishing is the primary delivery mechanism. | Inbound message logs, attachment scanning results, URL click tracking, quarantine events |
| Backup Servers | 94% of ransomware targets backups. You need to know the moment someone touches backup infrastructure. | Backup job status, admin access, configuration changes, deletion attempts |
| DNS | C2 traffic, data exfiltration, and malware callbacks all use DNS. If you log nothing else, log DNS. | Query logs (all queries or at least unique external domains) |
| Cloud Identity (Entra ID / Google Workspace) | 64% of cloud breaches involve identity misuse. OAuth consent, admin role changes, conditional access modifications. | Sign-in logs, audit logs, app consent grants, role assignments |
Keep On-Host Only (Low Detection Value per GB)
| Source | Why On-Host Only |
|---|---|
| File server access logs | Massive volume, low signal. Useful for forensics ("what did the attacker access?") but not for real-time detection. On-host 365d + backup is enough. |
| Print server logs | Compliance use only. No detection value. |
| Application debug logs | Developer forensics, not security detection. Keep on the app server. |
| Web server access logs (internal apps) | Useful for incident scoping but generates enormous volume. Centralize only if you have WAF detection rules. |
| Workstation event logs (general) | 200 workstations x 10,000 events/day = noise. Centralize only the Wazuh agent alerts, not the raw events. |
What Vendors Tell You
- "Centralize everything" — so they can charge per GB ingested
- "Store for 7 years" — for compliance that doesn't apply to you
- "You need a 12-node cluster" — for 200 employees?
- "Our SIEM detects threats with AI" — it detects what rules you write
- "Cloud SIEM eliminates management overhead" — at €10K–50K/month
- "You need SOAR for automation" — you need 15 good detection rules
What Actually Works
- 365 days on-host — 1GB+ Event Log channels, logrotate 365d. ~€3/month/server.
- Immutable backup — logs are already in your machine backup. €0 incremental.
- 30–90 days in Wazuh — one server, self-hosted, free. Sized for detection, not storage.
- 15–30 detection rules — tuned to real breach patterns, not vendor defaults. Alert to phone, not email.
- 6 log sources centralized — DCs, VPN/FW, email, backup, DNS, cloud identity. Everything else stays on-host.
- Monthly coverage check — are all agents alive? Are rules firing?
Economics: Right-Sized vs Over-Engineered
| Approach | Monthly Cost (200 seats) | What You Get |
|---|---|---|
| Right-sized (recommended) | €50–150 | Wazuh self-hosted (1 server, existing or €50–150/month VPS). 365d on-host. Backup copies. 30–90d searchable. 15–30 detection rules. |
| Elastic Cloud (managed) | €500–2,000 | Managed Elasticsearch. More features than you'll use. Scales well but costs scale too. |
| Splunk Cloud | €2,000–10,000 | Enterprise SIEM. Per-GB pricing punishes you for logging more. Mid-market rarely needs this. |
| Microsoft Sentinel | €500–5,000 | Cloud-native SIEM. Good M365 integration. Per-GB ingestion pricing. Costs grow with data. |
| Managed SIEM/MDR | €2,000–6,000 | Someone else runs the SIEM and alerts you. Valuable if you have zero security staff. Expensive otherwise. |
Consulting Economics
| Engagement | Duration | Investment |
|---|---|---|
| Log coverage assessment | 1 day | €1,250 |
| On-host retention hardening | 1 day | €1,250 |
| SIEM deployment (Wazuh) | 2–3 days | €2,500 – 3,750 |
| Detection rule library | 1–2 days | €1,250 – 2,500 |
| Alert routing + testing | 1 day | €1,250 |
| Full program | 6–8 days | €7,500 – 10,000 |
Free / Open-Source Tools
- Wazuh — Full SIEM/XDR, agents + manager + dashboard
- Grafana Loki — Lightweight log aggregation (no full-text index)
- Graylog Open — Open-source log management
- CrowdSec — Collaborative log-based intrusion prevention
- Elasticsearch OSS — Search engine (self-managed)
- NXLog CE — Cross-platform log collector
Quick Setup Commands
# Windows: Increase Security log to 1GB
wevtutil sl Security /ms:1073741824
# Windows: Enable PowerShell Script Block Logging
# GPO: Computer > Admin Templates > Windows
# Components > PowerShell > Script Block Logging
# Windows: Enable command-line auditing
# GPO: Computer > Windows Settings > Security
# Settings > Advanced Audit > Detailed Tracking
# > Audit Process Creation > Enable
# Then: Include command line in process
# creation events > Enabled
# Linux: 365 day retention in logrotate
# /etc/logrotate.d/rsyslog
/var/log/syslog {
rotate 365
daily
compress
delaycompress
missingok
notifempty
}