AI Agent Security

Govern the Machines Before They Govern You — Data Safety Levels for AI in Mid-Market (50–500 seats)

40%
of Block's 2025 detections were AI-assisted
8%
of MCP repos flagged as potentially malicious
0
mid-market companies with AI governance tiers
6mo–2yr
to autonomous D&R agents (Block CISO estimate)
10/7k
autonomous exploits by AI (Anthropic Mythos, 2026)
Your company is already using AI agents. Developers have Copilot. Marketing uses ChatGPT. Finance pastes data into Claude. The question isn't whether AI is in your environment — it's whether you know what data it's touching, who's responsible for its output, and what happens when it gets tricked. Most mid-market companies have zero visibility into this.
The attack surface is real: 8% of 17,845 MCP server repositories on GitHub were flagged as potentially malicious (VirusTotal, 2025). Prompt injection attacks in agent recipes have been demonstrated end-to-end by Block's red team. Anthropic's threat research shows AI-assisted attacks from reconnaissance through post-exploitation.

Entry Point: AI Agent Readiness Assessment

Inventory what's deployed. Classify by data sensitivity. Identify the three biggest governance gaps. Score each OODA loop. Determine if the organization can survive an AI-related incident today.
1–2 days Entry point / loss leader

Discovery Checklist

  • Interview IT, security, and 3–5 business unit leads on AI tool usage (Copilot, ChatGPT, Claude, Gemini, custom agents, MCP servers)
    2 hoursFree
  • Search SSO/OAuth logs for AI service provider authorizations (openai.com, anthropic.com, google.com AI scopes)
    30 minFree
  • Grep internal repos for API key patterns: sk-ant-, sk-proj-, AIza, Hugging Face tokens
    15 minFree
  • Review network egress logs for traffic to known AI API endpoints
    30 minFree
  • Catalog any MCP servers, custom agents, or automation recipes in use
    1 hourFree

Immediate Risk Check

# Find AI API keys in environment variables (Linux/macOS)
env | grep -iE "(OPENAI|ANTHROPIC|CLAUDE|GEMINI|HUGGING)" 2>/dev/null

# Search git repos for committed API keys
git log --all -p | grep -iE "(sk-ant-|sk-proj-|AIza|hf_)" | head -20

# Check OAuth app registrations in Entra ID
# (Azure AD PowerShell)
Get-AzureADServicePrincipal | Where-Object {
  $_.DisplayName -match "OpenAI|Anthropic|Claude|Copilot|ChatGPT"
} | Select DisplayName, AppId, ReplyUrls

# List browser extensions with AI capabilities (Chrome, enterprise)
# Check Chrome management console for extension inventory

Output

One-page AI agent inventory with data safety level classification per use case. Top 3 governance gaps prioritized by risk. 90-day roadmap. Maps to NIS2 Article 21 asset management requirements.

1 Inventory — What AI Is Running?
You can't govern what you can't see. Build a living inventory of every AI agent, copilot, MCP server, and model provider in your environment. Shadow AI is the new shadow IT.
Quick Wins (Day 0, Free)
  • Send a 5-question survey to all team leads: what AI tools does your team use, what data goes in, what actions can it take?
    30 minFree
  • Search SSO logs for OAuth grants to AI providers (OpenAI, Anthropic, Google AI, Microsoft Copilot). Export the list.
    30 minFree
  • Run TruffleHog across all internal repos to find committed AI API keys
    15 minFree
  • Check DNS/proxy logs for traffic to api.openai.com, api.anthropic.com, generativelanguage.googleapis.com
    15 minFree
  • Create a shared spreadsheet: AI Tool | Owner | Data Type | Access Level | Approved (Y/N)
    15 minFree
Core Engagement (2–3 days)
  • Comprehensive AI asset discovery: SSO audit + network analysis + endpoint scanning + team interviews across all business units
  • Map every AI integration's data flow: what goes in, what comes out, where it's stored, who has access
  • Identify unsanctioned ("shadow AI") usage — personal API keys, browser-based tools, unofficial MCP servers
  • Classify each integration by data sensitivity tier (see Loop 2)
  • Deliver: AI asset register with risk scoring, suitable for NIS2 Article 21 evidence
Target State
  • Automated continuous discovery: new AI OAuth grants trigger approval workflow
  • Network-level monitoring for AI API traffic with DLP rules for sensitive data patterns
  • MCP server registry with allow/deny lists per team role
  • Drift detection: alert when new AI integrations appear outside the approved register
2 Classify — Data Safety Levels
Block (the company behind Square/Cash App) models their AI governance on CDC biocontamination levels: the more sensitive the data an agent touches, the higher the security constraints. This is the framework adapted for mid-market.
Quick Wins (Day 0, Free)
  • Define four data sensitivity tiers for your organization. Start simple:
    1 hourFree
TierData TypesAI Allowed?Controls Required
Public Marketing content, public docs, open-source code Yes, any approved tool Approved tool list only. No credentials in prompts.
Internal Internal comms, non-sensitive business data, internal code Yes, with enterprise agreement Enterprise AI account (not personal). Data retention policy. No copy-paste of credentials or tokens.
Confidential Customer PII, financial data, HR records, proprietary algorithms Only with specific approval Enterprise AI with data processing agreement. Audit logging. Human review of all AI outputs before action. No training on company data.
Regulated Payment data (PCI), health records (medical), legal privilege, authentication secrets Only self-hosted or explicitly certified Self-hosted models or contractually guaranteed no-retention. Full audit trail. Human-in-the-loop mandatory. Separate network segment. Annual review.
  • Map each AI tool from your inventory (Loop 1) to its highest data tier
    1 hourFree
  • Publish a one-page "AI acceptable use" policy: what's allowed at each tier
    2 hoursFree
Core Engagement (2–3 days)
  • Formal data safety level framework with security controls matrix per tier
  • Map each tier to NIS2 Article 21 risk-based controls
  • Review all AI vendor DPAs (Data Processing Agreements) against tier requirements
  • Build approval workflow: new AI tool requests go through classification before deployment
  • Deliver: AI governance framework document + tier classification for all current tools + NIS2 compliance mapping
Target State
  • Policy-as-code: tier classification enforced at API gateway / proxy level
  • DLP integration: block sensitive data patterns from reaching unapproved AI endpoints
  • Automatic re-classification when tool capabilities change (new MCP server added, new model connected)
  • Quarterly review cycle aligned with NIS2 reporting requirements
3 Identity — Who's Responsible?
When an AI agent writes code, creates a detection rule, or sends a message — who's responsible? Block ties every agent action to a human identity. Most mid-market companies can't trace AI-generated output to the person who approved it.
Quick Wins (Day 0, Free)
  • Establish the rule: the human who runs the agent is responsible for its output. Write it down. Communicate it.
    15 minFree
  • Require AI-generated code to go through the same PR review process as human code. No exceptions.
    15 minFree
  • Audit: can you trace any AI-generated change in your codebase back to the human who approved it? Try it now with git log --author
    30 minFree
  • Ensure AI tools authenticate with individual user accounts, not shared service accounts
    30 minFree
  • Document which AI agents have write access to production systems (even indirectly via CI/CD)
    1 hourFree
Core Engagement (2–3 days)
  • Agent identity attribution audit: map every AI agent to its human owner, data access, and action permissions
  • Review CI/CD pipelines for AI-generated code paths — ensure human approval gates exist
  • Design approval workflow for agent-generated actions: analysis (auto) vs. response (human gate) vs. destructive (multi-approval)
  • Build accountability matrix: for each AI use case, document who approves, who reviews, who is liable
  • Deliver: Agent identity chain documentation + approval workflow design + accountability matrix
Target State
  • Automated provenance tagging on all agent output (commit metadata, audit logs, decision trails)
  • Identity-bound agent tokens: agent actions inherit the permissions and audit trail of the invoking user
  • Separation of duties enforcement: agents that analyze cannot also take action without human gate
  • Quarterly access review includes AI agent permissions alongside human permissions
4 Hardening — Prompt Injection & Data Exfiltration
LLMs can't reliably distinguish between data and commands. Prompt injection is trivial. Block red-teamed their own Goose agent with hidden prompt injection in recipes and built dual-layer defenses. Your agents need the same treatment.
Quick Wins (Day 0, Free)
  • Review agent tool permissions: are any agents running with unrestricted access when they only need read-only?
    30 minFree
  • Apply least privilege: separate READ_ONLY tools from WRITE/DESTRUCTIVE tools. Agents should only get what they need.
    1 hourFree
  • Disable unused MCP servers and tool integrations. Every connected tool is attack surface.
    30 minFree
  • Add input validation: block common prompt injection patterns in agent inputs (system prompt leaks, instruction overrides)
    2 hoursFree
  • Review agent output for data leakage: do any agents include sensitive context in their responses to users?
    1 hourFree
The other direction: Loop 4 focuses on agents being tricked. But AI also makes your infrastructure more vulnerable to exploitation. Anthropic's Mythos model (April 2026) autonomously discovered zero-days in operating systems and network services, wrote working ROP chains, and chained vulnerabilities for privilege escalation — capabilities previous models lacked entirely. Systems your agents connect to via MCP need aggressive patching because the attack surface is now discoverable at machine speed. Hardening agents is necessary but not sufficient — harden the infrastructure they touch.
Core Engagement (3–5 days)
  • Prompt injection red team: test all agent workflows with adversarial inputs (hidden instructions in documents, emails, tool results)
  • Implement dual-layer defense: deterministic input validation + LLM-as-judge reviewing agent actions before execution
  • Design tool access tiers modeled on Block's approach: analysis tools (auto) → enrichment tools (auto) → response tools (human gate) → destructive tools (multi-approval)
  • Data exfiltration testing: can an agent be tricked into sending sensitive data to an external endpoint?
  • Deliver: Red team report + remediation plan + tool access tier design + hardening implementation
Target State
  • Adversarial AI: secondary LLM reviews all agent tool calls and context for injection attempts before execution
  • Tool result sanitization: strip potential injection payloads from external data before it enters agent context
  • Constitutional guardrails: agents have hard-coded boundaries (no-go zones, maximum action severity, mandatory escalation triggers)
  • Regular red team cadence: quarterly prompt injection testing as part of penetration testing program
5 Detection — Monitor Agent Behavior
You monitor human user behavior for anomalies. You need to monitor agent behavior the same way. Agents that suddenly access new data sources, make unusual tool calls, or operate outside their normal patterns are either compromised or misconfigured.
Quick Wins (Day 0, Free)
  • Enable logging on all AI agent tool calls. At minimum: timestamp, agent ID, tool name, parameters, result status.
    1 hourFree
  • Set up basic alerts: agent accessing a data source it hasn't used before, agent making more than N tool calls per hour
    1 hourFree
  • Review agent logs weekly for unexpected patterns (off-hours activity, new tool calls, error spikes)
    30 min/weekFree
  • Log all data sent to external AI APIs. Even if you can't inspect it yet, having the logs means you can investigate later.
    1 hourFree
Core Engagement (2–3 days)
  • Design agent behavioral baseline: normal operating patterns per agent (tools used, data accessed, action frequency, timing)
  • Build detection rules for agent anomalies: deviation from baseline, privilege escalation attempts, data exfiltration patterns
  • Integrate agent audit logs into existing SIEM/detection pipeline (Wazuh, Panther, Sentinel)
  • Create incident response playbook for compromised agent scenarios
  • Deliver: Agent monitoring architecture + detection rules + IR playbook for AI-specific incidents
Target State
  • Vector database of historical agent actions for semantic anomaly detection (Block's "Binary Intelligent Triage" pattern)
  • MCP audit events feeding into centralized detection pipeline with correlation to human user activity
  • Autonomous agent monitoring: a dedicated monitoring agent reviews other agents' behavior with human-gated response
  • Full provenance chain: from external trigger → agent decision → tool call → action → outcome, all searchable and auditable

Economics: What This Costs vs. What It Prevents

InvestmentDurationCost (200 seats)What It Prevents
Assessment
Entry point
1–2 days €1,250 – 2,500 Visibility into shadow AI. Identifies top 3 risks before they materialize.
Loop 1: Inventory 2–3 days €2,500 – 3,750 NIS2 Article 21 asset management compliance. Eliminates unknown AI exposure.
Loop 2: Classify 2–3 days €2,500 – 3,750 Data leakage via inappropriate AI usage. Regulatory non-compliance.
Loop 3: Identity 2–3 days €2,500 – 3,750 Accountability gaps. Untraceable AI-generated changes in production.
Loop 4: Hardening 3–5 days €3,750 – 6,250 Prompt injection attacks. Agent compromise. Data exfiltration via AI.
Loop 5: Detection 2–3 days €2,500 – 3,750 Compromised agents operating undetected. Missed compliance evidence.
Full Program 12–19 days €15,000 – 23,750 Complete AI governance framework. NIS2-ready. Insurance-defensible.
The cost of NOT governing AI: A single data breach involving customer PII costs €4.3M on average (IBM, 2024). If an AI agent with access to customer data gets prompt-injected and exfiltrates records, you face GDPR penalties (up to 4% of global turnover), NIS2 penalties (up to €10M or 2% of turnover), plus the breach itself. The full governance program costs less than a single day of incident response.

NIS2 Compliance Mapping

Every loop in this playbook maps directly to NIS2 Article 21 requirements. AI governance isn't a separate compliance track — it's an extension of what you already need to do. Deadline: October 2026.
LoopNIS2 Article 21 RequirementHow This Playbook Addresses It
1. Inventory (a) Risk analysis and information system security policies AI asset register identifies all AI systems as information assets subject to risk analysis
2. Classify (a) Risk analysis + (e) Security in acquisition, development, and maintenance Data safety levels enforce risk-proportionate controls. Vendor DPA review covers acquisition security.
3. Identity (i) Human resources security, access control, and asset management Agent identity attribution ensures accountability. Access reviews include AI permissions.
4. Hardening (e) Security in acquisition, development, and maintenance + (j) Use of cryptography and encryption Prompt injection defense, tool access tiers, data exfiltration prevention, secure API key management
5. Detection (b) Incident handling + (g) Assessment of risk management measures Agent behavioral monitoring feeds into incident handling. Continuous assessment of AI security controls.
Insurance angle: Cyber insurers are beginning to ask about AI governance in underwriting questionnaires. Having documented data safety levels, agent identity chains, and prompt injection defenses positions you ahead of questions that will become standard within 12 months. Build the evidence now — it's cheaper than retrofitting under time pressure.

Free Afternoon Deploy: AI Governance Quick Start

Four things you can do this afternoon with zero budget that immediately reduce your AI risk exposure.
#ActionTimeWhat It Fixes
1 Send the AI usage survey to 5 team leads. Five questions: what tools, what data, what actions, how often, who approved? 30 min Shadow AI visibility. You'll be surprised.
2 Publish the accountability rule: "You are responsible for what your AI produces." Email it. Put it in the wiki. Done. 15 min Accountability gap. Establishes the norm before an incident forces it.
3 Run TruffleHog on your main repos: trufflehog git file://./your-repo --only-verified 15 min Committed AI API keys. Free, instant, no install needed (Docker).
4 Check SSO OAuth grants for AI provider authorizations. Revoke any that aren't sanctioned. 30 min Unsanctioned AI access to company data via OAuth. Immediate risk reduction.

Key References

  • Block GenAI Security Principles — engineering.block.xyz
  • OWASP Top 10 for LLM Applications — owasp.org/www-project-top-10-for-large-language-model-applications
  • NIST AI Risk Management Framework — nist.gov/artificial-intelligence
  • EU AI Act — high-risk AI system requirements applicable to agent deployments
  • VirusTotal MCP Repository Audit — 8% of 17,845 repos flagged (2025)
  • Anthropic Mythos Preview Assessment — AI-driven zero-day discovery: 10/7,000 autonomous exploits, CVE-2026-4747, friction-based defenses weakening (April 2026)

Free Tool Stack

  • Secret scanning: TruffleHog, GitLeaks
  • AI security: OWASP LLM Top 10 checklist, Rebuff (prompt injection detection)
  • MCP governance: Agent asset inventory (manual), OAuth audit via SSO logs
  • Monitoring: Wazuh (agent log ingestion), Prometheus (API call metrics)
  • DLP: Microsoft Purview (M365 E5), network proxy rules for AI endpoints
  • Compliance: NIS2 self-assessment templates (ENISA)