⛨Agent Harness Security
Engineering security reference

Securing an LLM agent harness, layer by layer

A structured catalog of the threats that arise when model output is turned into tool calls and actions — each with concrete mitigations, the limits of configuration, and how four common runtimes handle it by default. Mapped to the OWASP Top 10 for Agentic Applications (2026).

25 threats11 layers10 OWASP categories4 runtimesReviewed 2026-09-28

How to use this reference

01
Scoping a design
Walk the 11 layers; for each, decide which control the harness owns versus the surrounding infrastructure.
02
Choosing a runtime
Use the comparison matrix to see what each tool enforces by default and what you must add yourself.
03
Reviewing an implementation
Filter the threat catalog by layer and check each mitigation against the code under review.
04
Writing requirements
Turn the hardening checklist into acceptance criteria for the agent or harness.

The layered model

Agent security failures rarely stay in one place. A threat enters at one layer and causes damage at another — a poisoned document at the retrieval layer drives a tool call that exfiltrates data at the egress layer. Each layer is treated as an independent control point, on the assumption that the ones above it may be breached.

L013 threats
Prompt & Input
L022 threats
Retrieval & Memory
L033 threats
Tools & Function Calling
L044 threats
MCP & Connector Supply Chain
L052 threats
Code Execution & Sandbox
L062 threats
Identity & Secrets
L073 threats
Output & Egress
L081 threat
Model Supply Chain
L091 threat
Serving & Network
L103 threats
Orchestration & Multi-Agent
L111 threat
Observability & Audit

OWASP Top 10 for Agentic Applications (2026)

Published December 2025, these ten categories describe the risks most likely to compromise autonomous agents. Every threat in the catalog is tagged with the categories it maps to.

ASI01Agent Goal HijackObjective redirected via retrieved content, not code
ASI02Tool Misuse & ExploitationLegit tools bent to bad outcomes / chaining
ASI03Identity & Privilege AbuseBorrowed credentials with over-broad scopes
ASI04Agentic Supply ChainCompromised frameworks, connectors, MCP servers
ASI05Unexpected Code ExecutionNatural language becomes running code
ASI06Memory & Context PoisoningPlanted data shapes future behavior
ASI07Insecure Inter-Agent CommsSpoofing / trust inheritance across agents
ASI08Cascading FailuresOne compromise propagates through workflows
ASI09Human-Agent Trust ExploitationAgent manipulates humans into unsafe actions
ASI10Rogue AgentsOff-policy agents persisting across sessions

Runtime landscape

The four runtimes compared here occupy different points on a single axis: the more capable the built-in agent loop, the more safety controls it ships — and the barer the runtime, the more of the defense the surrounding harness must provide.

Claude CodeManaged harness
Ships the deepest default safety layer: instruction-source boundary, tiered action permissions, sandboxing, egress rules, MCP trust prompts.
LM StudioLocal GUI + MCP
OpenAI-compatible endpoint with a built-in MCP client and per-server trust prompts. The MCP path is unfiltered and unsandboxed beyond that.
mlx-lmBare inference
Apple-silicon inference library and server. No tools, no MCP, no harness — the smallest attack surface, and no guardrails.
OllamaLocal daemon
Inference daemon with function calling and no built-in MCP client. Notable for a history of unauthenticated API exposure and registry-pull CVEs.

Ollama and LM Studio are the same class of tool — local model runners exposing an OpenAI-compatible API with KV caching and function calling — but differ in posture: LM Studio defaults to a localhost-only server and a built-in MCP client with trust prompts, while Ollama is server-oriented, has no built-in MCP, and carries the ecosystem's main real-world exposure risk. See the full comparison.