Hardening checklist
Controls for a secure harness
The catalog's mitigations, distilled into groups for design and review. Suitable as acceptance criteria for an agent or harness project.
Governing principle: a model that both reads untrusted data and holds tool privileges is structurally unsafe. Most items below exist to break that link.
1
Architecture (design it in, don't bolt it on)
- ▢Adopt an instruction-source boundary: user chat = instructions; everything from tools/files/web = data.
- ▢Consider a dual-LLM / CaMeL split: a privileged planner that never sees raw untrusted data, a quarantined reader with no tool access.
- ▢Tag provenance on every context segment and carry it through to the action gate.
- ▢Default the whole system to least privilege and reversible/dry-run actions.
2
Tools & MCP
- ▢Allowlist MCP servers and pin exact versions; vendor them instead of `npx -y` from the internet.
- ▢Hash-pin tool definitions; re-prompt on any description/schema change (rug-pull defense).
- ▢Namespace tools per server; detect and block name collisions (shadowing).
- ▢Render full tool descriptions to operators; strip instruction-like text before it reaches the model.
- ▢Run each MCP server as an isolated, least-privileged process (container / restricted user).
3
Execution & secrets
- ▢Sandbox all code/shell execution (container/VM), no host mounts, no ambient credentials.
- ▢Prefer allowlists over denylists for commands; deny-by-default.
- ▢Keep secrets out of model context; inject at the execution boundary; short-lived, per-tool-scoped tokens.
- ▢Redact secrets in logs/traces; scan outputs for secret patterns before display or send.
4
Egress & actions
- ▢Egress allowlist: only user/operator-approved domains; block destinations from observed content.
- ▢Classify actions by reversibility/impact; require explicit human approval for irreversible ones.
- ▢Show concrete effects (files, recipients, amounts, destinations) in every confirmation.
- ▢Gate creation of standing rules, schedules, webhooks, and integrations as high-privilege.
5
Serving & model supply chain
- ▢Bind inference servers to 127.0.0.1; never 0.0.0.0; put an authenticating proxy in front for remote use.
- ▢Firewall inference ports; monitor for accidental exposure; patch the runtime promptly.
- ▢Pull models only from trusted publishers; verify checksums; prefer safetensors over pickle.
6
Operations
- ▢Log every prompt, tool call, argument, result (with provenance) to an append-only store.
- ▢Alert on sensitive actions and anomalous volume; add circuit breakers between automated steps.
- ▢Track your harness/runtime security advisories and keep dependencies patched.
- ▢Red-team with adversarial prompt-injection and command-fuzzing before shipping.
The three highest-leverage controls
- Instruction-source boundary — tool and retrieved content is data, never commands.
- Least privilege and sandboxing — so a successful injection reaches very little.
- Egress allowlist plus a human gate on irreversible actions — to cap the blast radius.