Agents and MCP · 7 MIN
How do you design tools for AI agents that models call correctly?
Good tool design decides whether an agent works. Learn how to scope tools, write descriptions, shape outputs, handle errors, and evaluate tool use.
Design tools for an agent the way you would design an interface for a new colleague who reads every word and cannot ask questions. Give each tool one clear job, a precise description, typed parameters, a compact readable result, and error messages that say how to fix the call. Most agent failures that look like model mistakes are tool design problems.
- Build tools around tasks the agent must complete, not around every endpoint your API has. More tools do not reliably produce better outcomes.
- The tool description is part of the prompt. Small wording changes can have large effects on whether the model picks and uses the tool correctly.
- Return high-signal, readable results with sensible limits, since every token a tool returns consumes the agent's context.
- Make mistakes hard to make: typed parameters, enums, and constraints that prevent invalid or dangerous calls.
- Write errors the model can act on, so it can correct itself and retry.
- Nactore designs tool surfaces and measures them with evals on real tasks before agents reach production.
What is a tool in an AI agent?
A tool is a function the model can ask the host application to run. The model sees a name, a description, and a schema for the inputs. It decides when to call it and with what arguments, the host executes the call, and the result goes back into the model's context.
The MCP specification for tools describes the same shape: a unique name, a description, an input schema in JSON Schema, and optionally an output schema. Whether you expose tools through MCP or define them inside your own agent runtime, the design principles are the same. If you are new to the protocol, start with what is MCP.
How many tools should an agent have, and what should each do?
Fewer than you think, each aimed at a task. Anthropic's guide to writing tools for agents advises building tools for specific, high-impact workflows rather than wrapping all your API endpoints, and observes that more tools do not always lead to better outcomes. It also suggests consolidating functionality, so one tool handles several operations under the hood.
| Instead of | Prefer | Why |
|---|---|---|
list_users, list_events, create_event as three calls | schedule_meeting that finds availability and books | One decision for the model, fewer round trips |
get_all_orders returning every field | find_order with a filter, returning the key fields | Smaller context, fewer wrong picks |
A generic run_query | Narrow, named lookups | Smaller blast radius and easier to test |
When you have many tools, or several servers, use consistent namespacing. The same guide suggests prefixes such as a service name followed by a resource, which helps the model choose between similar tools. The MCP spec also notes that names are only unique within one server, so aggregators should disambiguate.
How do you write a tool description the model will use correctly?
Write it like onboarding documentation for someone new to your domain. Say what the tool does, when to use it, when not to use it, and what each parameter means. Spell out any specialized terms and the formats you expect.
- Name the job in plain words.
search_invoicesbeatsquery_fin_db. - State the boundaries. "Use this for invoices only. For payments, use
search_payments." - Describe each parameter unambiguously. Prefer
customer_emailtouser, and say what a valid value looks like. - Add one short example where the format is easy to get wrong.
- Note side effects. Say whether the tool reads, writes, or sends something outside your system.
Anthropic reports that even small refinements to descriptions can yield dramatic improvements in agent performance. That is a reason to treat descriptions as code: version them and rerun your evals after any change. The MCP spec adds a caution for the other side of the wire. Descriptions and annotations from a server you do not trust should be considered untrusted.
What should a tool return?
Return what the model needs to decide the next step, and nothing more. The tools guide recommends high-signal information over raw flexibility, and replacing opaque identifiers with human-readable names where possible, since models handle names better than long IDs or mime types.
- Concise by default, detail on request. A
response_formatoption with concise and detailed modes lets the agent ask for more only when needed. - Pagination, filtering, and truncation with defaults. A tool that can return ten thousand rows eventually will.
- Guidance when truncating. Say "showing 20 of 340, narrow with a date range" so the agent narrows the search instead of retrying blindly.
- Structured output when you need it. MCP tools can declare an output schema, and servers must return results that conform to it.
Every token in a result competes for the agent's attention. Compact results also cost less and run faster.
How do you prevent bad calls?
Anthropic's guide on building effective agents borrows the term poka-yoke, mistake-proofing, for tool design. Change the interface so the error is hard to make.
- Use enums and constrained types instead of free text where the values are known.
- Require absolute identifiers instead of relative ones that can be misread.
- Validate on the server. The spec says servers must validate all tool inputs, enforce access controls, and rate limit.
- Separate read and write tools, and gate irreversible writes behind approval.
- Make destructive actions explicit. A tool named
delete_project_permanentlyis harder to call by accident thanupdate_projectwith a flag.
This also connects to safety. A narrow tool with limited credentials limits what a manipulated model can do, which we cover in MCP security risks.
How should tools report errors?
Report errors the model can use. The MCP spec distinguishes protocol errors, such as an unknown tool or malformed request, from tool execution errors, which are returned with isError: true. Execution errors such as bad input, API failures, and business rule violations should carry actionable text so the model can correct itself.
A useful pattern is: what went wrong, what a valid input looks like, and what to try next. "Date must be in the future. Today is 2026-07-20" teaches the model in one line. An opaque code teaches nothing. Also return partial results with a note when a call times out, instead of failing silently.
Read transcripts, not just scores. When an agent misuses a tool, the transcript usually shows why: two overlapping descriptions, an ambiguous parameter, or a result so long that the model lost the thread.
How do you test tool design?
Build an evaluation set of realistic tasks that need several tool calls, run it programmatically, and score outcomes. For each run, record whether the agent chose the right tools, the argument accuracy, the number of calls, errors, and tokens. The tools guide recommends reading transcripts to find where agents get confused, then revising and rerunning.
This is a loop. Change a description, rerun the set, compare. Keep a held-out set so you do not tune only to your own examples. Our post on AI evals before production walks through building the set.
Frequently asked questions
Is function calling the same as tool calling?
In practice, yes. Different vendors use different names for letting a model request a structured function call that the host application executes.
How long should a tool description be?
As long as needed to remove ambiguity, and no longer. A few sentences plus clear parameter descriptions is typical. Test changes against evals.
Should tool outputs be JSON or plain text?
Use whatever the model reads best for the task. Concise readable text often works well for the model, and structured output helps when downstream code also consumes the result.
Do we need MCP to do this well?
No. These principles apply to tools defined in your own code. MCP helps when you want to reuse tools across applications.
Takeaway
Fewer, sharper tools with clear descriptions and helpful errors beat a large surface of thin wrappers, and an eval set tells you which design is working. Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.