⛨Agent Harness Security
Methodology & sources

How this reference is built

Version 1.0 · last reviewed 2026-09-28.

Scope

Security risks that apply when composing an LLM agent harness — the loop that turns model output into tool calls and actions — with local or hosted models. Focused on the design decisions an engineering team controls, not on individual model weights.

Severity scale

Each threat carries a severity that reflects its realistic blast radius, not merely its likelihood.

CriticalRealistically leads to full data exfiltration, RCE, or credential compromise with little or no user interaction. Documented exploits or CVEs exist.
HighEnables significant unauthorized action or data loss, but typically needs a precondition (a granted tool, a chained step, or user approval of a disguised action).
MediumMeaningful weakness that degrades safety or aids a larger attack, but limited direct blast radius on its own.
LowHygiene / defense-in-depth gap; harmful mainly in combination with other failures.

How ratings are assigned

In the comparison matrix and on each threat page, every runtime is rated against each control on a four-level scale describing default, out-of-the-box behavior.

Built-inThe tool ships an enforced control for this risk by default — not merely documentation or a setting that must be found and enabled.
PartialSome coverage exists but is incomplete, config-dependent, model-dependent, or advisory rather than enforced.
NoneNo control for this risk in the tool. If one is needed, it must be built in the harness or surrounding infrastructure.
N/AThe risk doesn't apply to this tool's role (e.g. tool-layer risks for a bare inference server that executes no tools).

Versions assessed

Claude CodePublic security docs, Sept 2026 (post-v2.1.90 deny-rule fix)
LM Studio0.3.x release line (built-in MCP client)
mlx-lmCurrent main; mlx_lm.server OpenAI-compatible endpoint
Ollama0.17.1+ (post-CVE-2026-7482); historical CVEs noted where relevant

Key terms

Agent harness
The software loop around a model that turns its output into tool calls and actions, manages context, and enforces (or fails to enforce) policy. The harness — not the model — is where most of these controls live.
Instruction-source boundary
A design rule that treats input from the user's own channel as instructions and everything obtained through tools (files, web pages, tool output) as data that can never issue commands.
Indirect prompt injection
An attack where malicious instructions are placed in content the agent will later read via a tool, so they execute without the user ever typing them.
Tool poisoning
Hiding adversarial instructions inside an MCP tool's description, parameter schema, or response, which the model reads even though the operator sees only the tool name.
Line jumping
A tool-poisoning variant where the payload sits in the tool list itself, influencing the model at load time — before any tool is invoked or approved.
Rug pull
A server silently changing a tool's behavior or description after the client approved it, because base MCP has no content-addressing or version pinning.
Tool shadowing
Registering a tool whose name collides with a trusted one so that calls are silently redirected to the attacker's implementation.
Excessive agency
Granting an agent more tools, permissions, or autonomy than its task requires, which maximizes the damage a single injection can do.
Egress allowlist
A control that restricts which outbound destinations a network-capable tool may reach, blocking exfiltration to attacker-supplied URLs.
CaMeL / dual-LLM pattern
An architecture that separates a privileged planner (sees only trusted input, can call tools) from a quarantined reader (processes untrusted data, cannot call tools), with data-flow tracking between them.
MCP (Model Context Protocol)
An open protocol for exposing tools and data to an agent through servers. Convenient, but its trust model is where several of the supply-chain risks here originate.
KV cache
The cached attention key/value state that lets a model reuse computation across turns. A performance feature, but cached prompts and process memory are also a data-exposure surface.

Sources

↗ OWASP Top 10 for Agentic Applications 2026↗ OWASP Agentic Top 10 — breakdown (Cycode)↗ Claude Code — Security docs↗ Mitigate jailbreaks & prompt injection (Claude Platform)↗ Claude Code prompt-injection risks (TrueFoundry)↗ MCP tool poisoning notification (Invariant Labs)↗ MCP has prompt injection problems (Simon Willison)↗ MCP tool poisoning, rug pulls, line jumping (LensHQ)↗ Agentic MCP Security Best Practices (Cloud Security Alliance)↗ Understanding & securing exposed Ollama instances (UpGuard)↗ Ollama registry injection + GGUF memory safety (GitHub #16236)↗ CaMeL: type-directed privilege separation (arXiv)

Disclaimer

This is an independent engineering reference, not affiliated with or endorsed by any vendor listed. Tool behavior, versions, and CVE status change frequently. Ratings reflect default, out-of-the-box behavior at the review date and are a starting map — verify against each tool's current release and your own threat model before making architecture decisions. Nothing here is a security guarantee.