Agents and MCP · 7 MIN
What are the security risks of MCP, and how do you reduce them?
The real security risks of MCP servers and AI agents, mapped to the MCP spec and OWASP LLM Top 10, with a practical control checklist for CTOs.
The main security risks of MCP are prompt injection through tool output, tools that can do more than they should, token mishandling in remote servers, and untrusted local servers that run code on a user's machine. The MCP specification says plainly that it cannot enforce its own safety principles at the protocol level, so the controls have to be built into your host, your servers, and your approval flow.
- MCP tools represent arbitrary code execution, and the spec says hosts must obtain user consent before invoking any tool.
- Prompt injection is the central risk. OWASP lists it as LLM01:2025 and says foolproof prevention is unclear, so design for containment.
- Giving a model too much functionality, permission, or autonomy (OWASP LLM06:2025) turns a manipulated output into a damaging action.
- Remote servers must not accept or forward tokens that were not issued for them, which the spec calls out as token passthrough.
- Local servers run with the user's privileges, so treat install commands as untrusted code and sandbox where possible.
- Nactore designs agent and MCP systems with least privilege, approval gates, audit logs, and evals before anything touches production data.
What does the MCP spec say about security?
The specification lists three key principles. Users must consent to and understand data access and operations. Hosts must obtain explicit consent before exposing user data to servers. And tools should be treated as arbitrary code execution, with tool descriptions and annotations considered untrusted unless they come from a trusted server.
It then adds the caveat that matters most for planning: MCP cannot enforce these principles itself. Implementers are expected to build consent and authorization flows, document security implications, and apply access controls. Buying or adopting MCP does not give you a security posture. You still have to design one.
How does prompt injection reach an MCP-connected agent?
Prompt injection is when content the model reads changes what the model does. OWASP's description separates direct injection, where the user's own input alters behavior, from indirect injection, where external content such as a website or file carries instructions the model then follows.
MCP widens the indirect path. Every tool result is text that goes back into the model's context. A support ticket, a web page, a shared document, or a row in a database can contain instructions. If the same agent can also send email or write records, the injected text has a route to an action.
OWASP's mitigations are practical.
- Constrain the model's role and limits in the system prompt.
- Validate expected output formats.
- Filter inputs and outputs.
- Enforce least privilege on API access.
- Require human approval for high-risk actions.
- Segregate and clearly mark untrusted external content.
- Run adversarial testing regularly.
OWASP also states that foolproof prevention methods remain unclear because of how these models work. So plan on a model that will occasionally be fooled, and limit what a fooled model can do.
Why does giving a model too much power matter for tools?
LLM06:2025 describes systems that take damaging actions based on unexpected or manipulated model output when they have been given tools. OWASP names three root causes.
| Root cause | What it looks like in an MCP setup | Control |
|---|---|---|
| Excessive functionality | A server exposes run_sql when the task needs get_order_status | Expose narrow, task-shaped tools |
| Excessive permissions | The server's database account can write and delete everywhere | Scope credentials to the minimum needed, per tool |
| Excessive autonomy | The agent sends refunds or emails without review | Human approval for consequential actions |
This is the section to read with your security team. Most real incidents with agents are not clever exploits of the protocol. They are a model with broad permissions following a bad instruction.
What are the MCP-specific attack classes?
The spec's security best practices page documents several. The ones most relevant to a mid-size company building servers are below.
- Confused deputy. An MCP proxy server that fronts a third-party API with a static client ID can be tricked into skipping user consent. The spec requires per-client consent before forwarding to the third-party authorization flow.
- Token passthrough. A server accepts a token from a client without checking it was issued to the server, then forwards it downstream. The spec forbids this. Servers must not accept tokens that were not explicitly issued for them.
- Server-side request forgery. A malicious server can point a client at internal URLs, including cloud metadata endpoints. Clients should enforce HTTPS, block private and link-local ranges, and consider an egress proxy.
- State handle hijacking. If your tools return handles, such as a cart or workflow ID, a guessed handle must not grant access. Bind handles to the authenticated user and use unpredictable values.
- Local server compromise. A local server is code running with the client's privileges. The spec requires clients to show the exact command and get approval before one-click installs, and recommends sandboxing.
- Scope inflation. Broad scopes granted up front enlarge the damage of a stolen token. Start minimal and elevate on demand.
What controls should be in place before launch?
Use this as a pre-launch checklist.
- Inventory the tools. For each one, record whether it reads, writes, or is irreversible.
- Apply least privilege. Give each server its own credentials with the narrowest access that works.
- Gate writes. Require approval for irreversible or external-facing actions.
- Validate inputs and sanitize outputs. The spec says servers must validate all tool inputs, apply access controls, rate limit, and sanitize outputs.
- Log usage. Keep an audit trail of tool calls on the server and the client.
- Vet third-party servers. Read the code, pin versions, and run them sandboxed.
- Test adversarially. Include prompt-injection cases in your evals, such as a document that tells the agent to email a file elsewhere.
Separate what an agent can read from what it can do. A common safe pattern is to let an agent that reads untrusted content hold no write tools at all, and pass its summary to a second step that has limited, approved actions.
How do you test for these risks?
Treat security cases as part of your eval set, not a separate exercise. Write tasks where the data contains an instruction the agent should ignore, where a tool returns an unexpectedly large or malformed result, and where the agent is asked to do something outside its remit. Score whether it refused, asked for approval, or acted. Run the set on every change to prompts or tools. Our post on evals before production explains how to build one, and tool calling design for agents covers the design side.
Frequently asked questions
Is MCP less secure than a normal API integration?
Not inherently. The difference is that a model decides when to call the tool, and it reads untrusted text while deciding. That adds prompt injection and over-permissioned tools to the usual API concerns.
Can we trust a public MCP server from a registry?
Treat it as untrusted code. The spec notes tool annotations should be considered untrusted unless they come from a trusted server. Review the code, pin the version, limit its access, and prefer running it in a sandbox.
Does human approval solve prompt injection?
It reduces the impact of the highest-risk actions, but it does not stop injection. Approval fatigue is real, so reserve it for consequential steps and combine it with least privilege.
Who should own agent security in a mid-size company?
Shared ownership works best. Engineering owns the controls, security reviews the tool inventory and permissions, and the business owner decides which actions require approval.
Takeaway
Assume the model can be misled, then make sure a misled model cannot do much harm. Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.