Writing tool definitions for a customer's API: six steps to get the model to pick the right tool and fill in the right parameters
The model never reads the customer's code or Swagger file. Everything it knows about the API is in the few lines of description you write, so a wrong tool call usually starts with those lines.
In brief
- The model chooses a tool based only on its name, description and inputs, so the description is the most important thing you control.
- Merge related endpoints into fewer tools, use names like acme_orders, and write order_id rather than order.
- Errors and responses are prompts too: rewrite errors so the model can correct itself, and return only the data it actually needs.
The model knows only what you write, so the revised version spells out everything the draft left open.
Graphic: FDE Times
Anthropic’s tool use documentation says that a detailed description is by far the most important factor in how well a tool performs. Before you think about switching models or frameworks, then, look again at the sentences you wrote yourself.
The reason is simple. According to Hugging Face’s Agents course, tool descriptions are inserted into the system prompt, so the model knows only what is written there. Paragon describes the same mechanism: the model picks a tool based on the name, the description and the inputs attached to each one.
The customer has an internal API, and you have to wrap it so an agent can use it. This guide works step by step through a hypothetical example, with a check after each step.
What will you build, and what do you need?
Imagine the customer is a delivery company called Acme. It has three endpoints: look up an order by ID, find orders by phone number, and cancel an order. Your goal is for a customer-service agent to call the right endpoint with the right parameters when a user asks in natural language.
Remember how function calling works. The Prompting Guide describes it as a way to turn natural language into valid API calls, but the model only generates arguments as JSON; actually calling the API is your code’s job. You need to be comfortable with JSON Schema, have the customer’s API documentation, and have a small wrapper layer sitting between the model and the API.
Step 1: See why the first draft fails
This is the kind of definition you often see when someone copies endpoint names straight in:
{
"name": "get",
"description": "Get order",
"input_schema": {
"type": "object",
"properties": { "order": { "type": "string" } },
"required": ["order"]
}
}
The name “get” says nothing about what it fetches or from which system. The parameter “order” could be an order ID, a whole order object or a sequence number. A two-word description says nothing about when to use the tool.
A post on Anthropic’s engineering blog recommends unambiguous parameter names, such as user_id instead of user. The same applies here: order_id is far clearer than order.
Check: read the definition aloud to a colleague who has never seen Acme’s API. If they ask “what is order?”, the model will misunderstand in exactly the same way.
Step 2: Merge endpoints and add a service prefix to names
Many people’s first instinct is one tool per endpoint. Anthropic’s documentation advises the opposite: consolidate related operations into fewer tools with an action parameter, because fewer, more capable tools leave the model less to deliberate over when choosing.
The same documentation recommends prefixing names with the service, as in github_list_prs or slack_send_message, so that selection stays clear as the tool library grows.
Applied to Acme, the three endpoints become one tool called acme_orders, with action taking one of three values: get, search, cancel. If the agent later gets tools from a CRM system, the acme_ prefix tells them apart immediately.
Check: list every tool the agent currently has. If there are two whose difference you have to read carefully to spot, merge or rename them.
Step 3: Write the description as if briefing a new hire
Anthropic suggests writing a description the way you would explain the tool to someone new to the team, spelling out all the context that insiders take for granted. Its documentation lists what to include: what the tool does, when to use it and when not to, what each parameter means, and any caveats.
Aim for at least 3-4 sentences per tool, more if the tool is complex. Note that each JSON string must sit on one line; if you want a line break inside a description, use the \n escape.
{
"name": "acme_orders",
"description": "Look up, search for, or cancel delivery orders in the Acme system. Use action='get' when the user has given an order ID; use action='search' when you only have the recipient's phone number. Only use action='cancel' when the user explicitly says they want to cancel, and only for orders that have not yet left the warehouse.
Do not use this tool for questions about shipping rates or compensation claims. Order IDs look like ORD- followed by 8 digits.",
"input_schema": {
"type": "object",
"properties": {
"action": { "type": "string", "enum": ["get", "search", "cancel"] },
"order_id": { "type": "string", "description": "Order ID in the form ORD-12345678. Required for get and cancel." },
"recipient_phone": { "type": "string", "description": "Recipient's phone number, digits only. Used with search." }
},
"required": ["action"]
}
}
The description tells the model to use get when the user gives an order ID, search when there is only the recipient’s phone number, and cancel only when the user explicitly asks to cancel and the order has not yet left the warehouse. It rules out questions about shipping rates and compensation claims, and states that order IDs are ORD- followed by 8 digits.
The order ID format, the “not yet left the warehouse” rule and the exclusion of claims are all hypothetical details. With a real customer, you have to dig these rules out, and they usually live in the operations team’s heads rather than in the API documentation.
Check: does the description contain a sentence starting with “Do not use…”?
Step 4: Add examples for inputs that are hard to format
For tools with complex or format-sensitive inputs, Anthropic’s API offers an optional input_examples field. Every example must validate against input_schema; an invalid one returns a 400 error. Examples also cost extra prompt tokens, so use them only when a format is genuinely often filled in wrong. The snippet below is simplified:
"input_examples": [
{ "action": "get", "order_id": "ORD-20481234" },
{ "action": "search", "recipient_phone": "0901234567" }
]
Check: send the request. If you get a 400 error, an example most likely deviates from the schema, for instance a misspelled field name or an action value outside the enum.
Step 5: Responses and errors are prompts too
Whatever the API returns goes back into the model’s context. Anthropic’s engineering blog advises that tools return only high-value information, and suggests pagination, filtering and truncation when payloads are large. The same post notes that error messages can be rewritten to say clearly what needs fixing, helping the model adjust on its next call.
The wrapper below is simplified. It separates two different failures: an ID in the wrong format (caught in the wrapper, before any API call) and an ID in the right format that does not exist (the API returns 404).
import re
# Simplified: wrapper between the model and the customer's API
ORDER_ID_PATTERN = re.compile(r"^ORD-\d{8}$")
def run_acme_orders(args):
if args["action"] in ("get", "cancel"):
order_id = args.get("order_id", "")
if not ORDER_ID_PATTERN.match(order_id):
return ("order_id has the wrong format: it must be ORD- + 8 digits, "
"e.g. ORD-20481234. Ask the user for the order ID again.")
resp = call_customer_api(args) # your HTTP function
if resp.status == 404:
return ("Order not found even though the ID has the right format.\n\n"
"If the user has the recipient's phone number, "
"call again with action='search'.")
order = resp.json()
return {"order_id": order["id"], "status": order["status"],
"eta": order["eta"]} # drop internal fields the model doesn't need
Compared with a bare “404 Not Found” string, each message here points the model down a different path: a malformed ID means asking the user again; a valid but non-existent ID means switching to a phone-number search. Trimming the response to three fields spares the model from wading through dozens of internal fields.
Check: make two test calls, one with a malformed ID and one with a well-formed ID that does not exist. The model’s next call should differ between the two cases, rather than retrying the same bad ID.
Step 6: Don’t trust your gut, run an eval
Anthropic reports that even small refinements to a description can improve results dramatically. That cuts both ways: changing one sentence can make things better, but it can also break a case that used to work.
The practical approach is to assemble about 20 questions taken from the customer’s support logs. For each, record the correct tool, action and arguments. Every time you edit a description, rerun the whole set and count the correct answers.
Traps to avoid
A common mistake is auto-generating tools from OpenAPI and leaving them as they are. That leads straight to what Anthropic’s documentation warns against: dozens of tools, one per endpoint, with descriptions copied from developer comments and too short for the model to tell them apart. Another mistake is leaving parameters named id, user or data with no format stated.
Nor should you pass the API’s raw JSON straight back, errors included, and then wonder why the model keeps repeating the same bad call.
What does this skill look like on a customer site?
On a customer site, the hard part is rarely JSON syntax. It is sitting with the operations team to extract answers to questions like “can an order be cancelled once it has left the warehouse?”, because that is exactly what goes into the “when not to use” sentence of the description.
If you are aiming for an FDE role, look for job descriptions that mention tool calling, agent integration or working directly with customer APIs.
On your CV, do not just write “integrated LLMs with internal APIs”. Say how many endpoints you merged into how many tools, and how accuracy on your eval set changed across description revisions. A before-and-after number like that shows a hiring manager you can do the work a customer will pay for.
Next time an agent calls the wrong API, don’t rush to swap the model. Open the description and read it as a new hire would on their first day.
Was this article useful?
Thanks for the feedback!