⛨Agent Harness Security
Hardening checklist

Controls for a secure harness

The catalog's mitigations, distilled into groups for design and review. Suitable as acceptance criteria for an agent or harness project.

Governing principle: a model that both reads untrusted data and holds tool privileges is structurally unsafe. Most items below exist to break that link.

1

Architecture (design it in, don't bolt it on)

  • ▢Adopt an instruction-source boundary: user chat = instructions; everything from tools/files/web = data.
  • ▢Consider a dual-LLM / CaMeL split: a privileged planner that never sees raw untrusted data, a quarantined reader with no tool access.
  • ▢Tag provenance on every context segment and carry it through to the action gate.
  • ▢Default the whole system to least privilege and reversible/dry-run actions.
2

Tools & MCP

  • ▢Allowlist MCP servers and pin exact versions; vendor them instead of `npx -y` from the internet.
  • ▢Hash-pin tool definitions; re-prompt on any description/schema change (rug-pull defense).
  • ▢Namespace tools per server; detect and block name collisions (shadowing).
  • ▢Render full tool descriptions to operators; strip instruction-like text before it reaches the model.
  • ▢Run each MCP server as an isolated, least-privileged process (container / restricted user).
3

Execution & secrets

  • ▢Sandbox all code/shell execution (container/VM), no host mounts, no ambient credentials.
  • ▢Prefer allowlists over denylists for commands; deny-by-default.
  • ▢Keep secrets out of model context; inject at the execution boundary; short-lived, per-tool-scoped tokens.
  • ▢Redact secrets in logs/traces; scan outputs for secret patterns before display or send.
4

Egress & actions

  • ▢Egress allowlist: only user/operator-approved domains; block destinations from observed content.
  • ▢Classify actions by reversibility/impact; require explicit human approval for irreversible ones.
  • ▢Show concrete effects (files, recipients, amounts, destinations) in every confirmation.
  • ▢Gate creation of standing rules, schedules, webhooks, and integrations as high-privilege.
5

Serving & model supply chain

  • ▢Bind inference servers to 127.0.0.1; never 0.0.0.0; put an authenticating proxy in front for remote use.
  • ▢Firewall inference ports; monitor for accidental exposure; patch the runtime promptly.
  • ▢Pull models only from trusted publishers; verify checksums; prefer safetensors over pickle.
6

Operations

  • ▢Log every prompt, tool call, argument, result (with provenance) to an append-only store.
  • ▢Alert on sensitive actions and anomalous volume; add circuit breakers between automated steps.
  • ▢Track your harness/runtime security advisories and keep dependencies patched.
  • ▢Red-team with adversarial prompt-injection and command-fuzzing before shipping.
The three highest-leverage controls
  1. Instruction-source boundary — tool and retrieved content is data, never commands.
  2. Least privilege and sandboxing — so a successful injection reaches very little.
  3. Egress allowlist plus a human gate on irreversible actions — to cap the blast radius.