Runtime comparison
Features and security controls
Each cell records default, out-of-the-box behavior at the review date. Ratings follow a fixed rubric — how ratings are assigned. The consistent pattern: capability and built-in safety rise together; the barer the runtime, the more defense the surrounding harness must supply.
Claude Code
Managed agent harness. Ships the deepest built-in safety layer: instruction-source boundary, permission model, sandboxing, MCP trust prompts.
LM Studio
Local inference GUI + OpenAI API with a built-in MCP client and per-server trust prompts. Thin permission model beyond that; the agent loop is the integrator's responsibility.
mlx-lm
Bare Apple-silicon inference library + server. No tools, no MCP, no harness. Tiny attack surface, but zero safety controls — every one must be added by the integrator.
Ollama
Local model daemon + OpenAI API with function calling. Server-oriented; no built-in MCP client. History of exposure CVEs and unauthenticated API.
Built-in— Real control shipped by the toolPartial— Some coverage, config-dependent or incompleteNone— No control — must be added in the harnessN/A— Not applicable to this tool's role
| Control | Claude Code | LM Studio | mlx-lm | Ollama |
|---|---|---|---|---|
Role What the tool fundamentally is | Built-in Full managed agent harness | Partial Inference GUI + thin agent loop | None Bare inference library/server | Partial Inference daemon + API |
Instruction-source boundary Treats tool/retrieved content as data, not commands | Built-in Explicit, enforced | None Model-dependent only | None Integrator-built | None Model-dependent only |
Permission / action model Tiered approval for sensitive actions | Built-in Prohibited/ask/regular + modes | Partial Per-call confirm prompt | None None | None None |
Execution sandbox Isolates code/tool execution from host | Partial Sandbox mode + folder grants | None Tools run on host | N/A No execution | None Client-side, unsandboxed |
Built-in MCP client Native MCP tool integration | Built-in Yes, managed + local | Built-in Yes (mcp.json) | None No | None Via external clients only |
MCP trust / rug-pull defense Vets servers and re-checks on change | Partial Managed set + trust step | Partial Trust on add, no re-check | N/A N/A | N/A N/A |
Egress / exfil control Restricts outbound destinations | Built-in Privacy rules + confirm | None None | N/A No net tools | None None |
KV cache / prompt reuse Caches attention state across turns | Built-in Server-side prompt caching | Built-in llama.cpp/MLX cache + prefix reuse | Built-in MLX cache, prompt-cache API | Built-in ggml/llama.cpp cache |
API authentication Auth on the inference endpoint | N/A Not a self-hosted server | Partial localhost default, no token | Partial localhost default, no auth | None No auth; often exposed |
Model supply-chain control Provenance/integrity of weights | N/A Hosted models | Partial Curated HF + checksums | Partial HF pull, safetensors default | None Arbitrary registries, GGUF CVEs |
Audit / observability Logs of prompts, calls, actions | Partial Transcripts + tool logs | Partial App/server logs | None Minimal | Partial Request logs |
Source availability Is the tool itself auditable? | Partial Closed harness, open SDK/docs | None Closed source app | Built-in Open source (MIT) | Built-in Open source (MIT) |
Positioning for a production harness
None of these is a complete security solution on its own. The practical pattern is to select an inference runtime for its performance and licensing, then implement the missing controls in a harness in front of it.
Claude Code
The reference point for built-in safety. A custom harness should aim to reproduce its instruction-source boundary, tiered action permissions, and egress rules — that is the baseline to match.
LM Studio
A capable local endpoint with easy MCP, but the MCP path is unfiltered and unsandboxed. Every added server should be treated as untrusted code running with the operator's privileges.
mlx-lm
The smallest attack surface — no tools, no MCP, open source — and no guardrails. Best used as a pure inference core wrapped in a controlled, purpose-built harness.
Ollama
Same inference class as LM Studio, but requires deliberate lockdown: bind to localhost, patch promptly, and restrict model sources to trusted registries.