Agents and MCP · 7 MIN
Building an MCP server for production: lessons from the build
What to decide before you build an MCP server: tool shape, error handling, auth, logging, and evals. Engineering lessons for teams shipping to production.
A production-ready MCP server is mostly not about the protocol. The protocol part is a few hundred lines. The work is deciding which tools to expose, how they fail, who is allowed to call them, what gets logged, and how you prove the model uses them correctly. This post covers those decisions in the order we make them when we build one.
- Design tools around tasks, not around your API endpoints. Fewer, sharper tools beat a wrapper for every route.
- Report recoverable failures as tool execution errors so the model can correct itself, and keep protocol errors for malformed requests.
- On stdio, write only valid MCP messages to stdout and send logs to stderr. A stray print statement breaks the connection.
- Remote servers need real authorization: validate that tokens were issued for your server, and request the smallest scopes that work.
- Evals on real tasks tell you whether the server works. A passing unit test does not.
- Nactore builds MCP servers and the agents that use them, with evals, scoped to each team.
What should an MCP server expose first?
Start with the user's job, not your database schema. Ask what task a person is currently doing by hand across two systems, then write the smallest set of tools that does it end to end.
Anthropic's guidance on writing tools for agents makes the same point. It argues for building tools aimed at specific, high-impact workflows rather than wrapping every endpoint, and notes that more tools do not always lead to better outcomes. A list_all_customers tool that dumps a thousand rows into context is a worse design than find_customer that takes a name or email and returns the three fields the model needs.
| Weak tool shape | Stronger tool shape |
|---|---|
| One tool per REST endpoint | One tool per task the user cares about |
| Returns raw database rows | Returns a short, readable summary with names instead of opaque IDs |
Open-ended query parameter | Typed parameters with enums and clear descriptions |
| Always returns everything | Supports a concise or detailed response, with pagination and sensible defaults |
We have found that the first draft of any tool surface is too wide. Cutting it down is the most valuable edit in the project.
How should the server report errors?
The MCP specification separates two kinds. Protocol errors are for problems with the request itself, such as an unknown tool or a malformed message, and they come back as standard JSON-RPC errors. Tool execution errors cover API failures, input validation problems, and business logic failures. These return a normal result with isError: true, so the model can read the message and try again.
The practical rule is to write execution errors for the model as the reader. Compare two messages.
- "Error 422." The model has nothing to act on.
- "Invalid departure date: must be in the future. Today is 2026-09-22." The model fixes the argument and retries.
The tools section of the specification uses almost exactly that second example, and it says clients should pass execution errors to the model so it can self-correct.
What breaks on stdio?
Local servers launched by a client use stdio, and the rule is strict. Per the stdio transport spec, the server must not write anything to stdout that is not a valid MCP message. Logging goes to stderr.
The failure mode is mundane. A developer adds a debug print, the output lands on stdout, the client fails to parse a line, and the server appears to hang or disconnect. If a server works in a test harness and dies inside a real client, check for this first.
Two more habits help. Exit promptly when stdin closes, because that is the portable shutdown signal. And keep each message on a single line, since messages are newline-delimited.
How do you handle authentication and scopes?
For stdio servers the spec says to retrieve credentials from the environment rather than follow the HTTP authorization flow. For remote servers over HTTP, authorization is optional in the spec, but you should treat it as required for anything touching real data.
The authorization specification builds on OAuth. The parts that bite in practice are these.
- Audience validation. Servers must validate that access tokens were issued specifically for them, and must not accept or pass through other tokens.
- Least-privilege scopes. Start with a minimal read scope and ask for more through a targeted challenge when a privileged tool is first used.
- Authorization on every call. Check what the caller is allowed to do inside each tool, not only at connection time.
- Protected resource metadata. Remote servers must publish it so clients can discover the authorization server.
We do not write OAuth plumbing from scratch if a vetted library or gateway exists. The security risks are covered in MCP security risks.
Keep a clear line between read tools and write tools from day one. Ship read-only first, watch real calls for a while, then add write tools one at a time behind explicit approval.
What should you log and test?
The spec's security considerations say servers must validate all tool inputs, apply access controls, rate limit, and sanitize outputs. It says clients should log tool usage for audit. We log on the server too, because the server is the part you control.
For each call, record who called, which tool, the arguments, the outcome, and the latency. Redact secrets and personal data before anything reaches the log store.
Testing has two layers.
- Deterministic tests for the server code: schema validation, error paths, auth failures, pagination.
- Evals for the model's behavior with your tools: a set of realistic tasks, scored on whether the model chose the right tool, passed correct arguments, and reached the right final answer.
Anthropic recommends building evaluations from real-world tasks that need multiple tool calls, running them programmatically, and reading the transcripts to find where the model got confused. The transcript review is where you find that two tool descriptions overlap, or that a parameter name is ambiguous. Fix the description, rerun, and compare. Our post on evals before production goes deeper on the method.
How do you ship and version it?
Treat the tool list as a public interface. Renaming a tool or changing its arguments breaks every client that has learned the old shape.
- Pin the protocol revision your server and clients target, and test against the fallback behavior for older clients.
- Return tools in a deterministic order. The spec recommends it for client caching and better prompt cache hit rates.
- Namespace tool names with a consistent prefix when your server may sit beside others, since names are only guaranteed unique within one server.
- Add, do not mutate. Introduce a new tool for new behavior and deprecate the old one on a schedule.
- Run the eval set on every change to a tool description, because small wording changes can shift model behavior.
Frequently asked questions
How long does it take to build an MCP server?
A narrow read-only server with two or three tools can be small. Most of the time goes into auth, error design, logging, and evals, not the protocol layer. Scope it as a pilot with a defined set of tools and a defined eval set.
Should one MCP server wrap our whole API?
Usually not. A wide surface confuses the model and raises risk. Expose the tasks people actually do, and add tools as the evals show a need.
Do we need OAuth for an internal server?
If it is remote and touches real data, you need real authentication and per-call authorization. For a local stdio server, credentials come from the environment, and the risk shifts to what that process can access.
Can we test an MCP server without a model?
You can test the server code that way, and you should. You still need model-in-the-loop evals to learn whether the model picks the right tool with the right arguments.
Takeaway
Build the smallest server that completes one real task, make its errors readable by a model, lock down auth, and measure with evals before widening it. For the concept behind it, see what is MCP. Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.