# Writing tool definitions for a customer's API: six steps to get the model to pick the right tool and fill in the right parameters

> The model never reads the customer's code or Swagger file. Everything it knows about the API is in the few lines of description you write, so a wrong tool call usually starts with those lines.

Original: https://fdetimes.net/en/guides/writing-tool-definitions-for-customer-apis/

Anthropic's tool use documentation says that a detailed description is by far the most important factor in how well a tool performs. Before you think about switching models or frameworks, then, look again at the sentences you wrote yourself.

The reason is simple. According to Hugging Face's Agents course, tool descriptions are inserted into the system prompt, so the model knows only what is written there. Paragon describes the same mechanism: the model picks a tool based on the name, the description and the inputs attached to each one.

The customer has an internal API, and you have to wrap it so an agent can use it. This guide works step by step through a hypothetical example, with a check after each step.

## What will you build, and what do you need?

Imagine the customer is a delivery company called Acme. It has three endpoints: look up an order by ID, find orders by phone number, and cancel an order. Your goal is for a customer-service agent to call the right endpoint with the right parameters when a user asks in natural language.

Remember how function calling works. The Prompting Guide describes it as a way to turn natural language into valid API calls, but the model only generates arguments as JSON; actually calling the API is your code's job. You need to be comfortable with JSON Schema, have the customer's API documentation, and have a small wrapper layer sitting between the model and the API.

## Step 1: See why the first draft fails

This is the kind of definition you often see when someone copies endpoint names straight in:

```json
{
  "name": "get",
  "description": "Get order",
  "input_schema": {
    "type": "object",
    "properties": { "order": { "type": "string" } },
    "required": ["order"]
  }
}
```

The name "get" says nothing about what it fetches or from which system. The parameter "order" could be an order ID, a whole order object or a sequence number. A two-word description says nothing about when to use the tool.

A post on Anthropic's engineering blog recommends unambiguous parameter names, such as `user_id` instead of `user`. The same applies here: `order_id` is far clearer than `order`.

**Check:** read the definition aloud to a colleague who has never seen Acme's API. If they ask "what is order?", the model will misunderstand in exactly the same way.

## Step 2: Merge endpoints and add a service prefix to names

Many people's first instinct is one tool per endpoint. Anthropic's documentation advises the opposite: consolidate related operations into fewer tools with an action parameter, because fewer, more capable tools leave the model less to deliberate over when choosing.

The same documentation recommends prefixing names with the service, as in `github_list_prs` or `slack_send_message`, so that selection stays clear as the tool library grows.

Applied to Acme, the three endpoints become one tool called `acme_orders`, with `action` taking one of three values: `get`, `search`, `cancel`. If the agent later gets tools from a CRM system, the `acme_` prefix tells them apart immediately.

**Check:** list every tool the agent currently has. If there are two whose difference you have to read carefully to spot, merge or rename them.

## Step 3: Write the description as if briefing a new hire

Anthropic suggests writing a description the way you would explain the tool to someone new to the team, spelling out all the context that insiders take for granted. Its documentation lists what to include: what the tool does, when to use it and when not to, what each parameter means, and any caveats.

Aim for at least 3-4 sentences per tool, more if the tool is complex. Note that each JSON string must sit on one line; if you want a line break inside a description, use the `\n` escape.

```json
{
  "name": "acme_orders",
  "description": "Look up, search for, or cancel delivery orders in the Acme system. Use action='get' when the user has given an order ID; use action='search' when you only have the recipient's phone number. Only use action='cancel' when the user explicitly says they want to cancel, and only for orders that have not yet left the warehouse.

Do not use this tool for questions about shipping rates or compensation claims. Order IDs look like ORD- followed by 8 digits.",
  "input_schema": {
    "type": "object",
    "properties": {
      "action": { "type": "string", "enum": ["get", "search", "cancel"] },
      "order_id": { "type": "string", "description": "Order ID in the form ORD-12345678. Required for get and cancel." },
      "recipient_phone": { "type": "string", "description": "Recipient's phone number, digits only. Used with search." }
    },
    "required": ["action"]
  }
}
```

The description tells the model to use `get` when the user gives an order ID, `search` when there is only the recipient's phone number, and `cancel` only when the user explicitly asks to cancel and the order has not yet left the warehouse. It rules out questions about shipping rates and compensation claims, and states that order IDs are ORD- followed by 8 digits.

The order ID format, the "not yet left the warehouse" rule and the exclusion of claims are all hypothetical details. With a real customer, you have to dig these rules out, and they usually live in the operations team's heads rather than in the API documentation.

**Check:** does the description contain a sentence starting with "Do not use…"?

## Step 4: Add examples for inputs that are hard to format

For tools with complex or format-sensitive inputs, Anthropic's API offers an optional `input_examples` field. Every example must validate against `input_schema`; an invalid one returns a 400 error. Examples also cost extra prompt tokens, so use them only when a format is genuinely often filled in wrong. The snippet below is simplified:

```json
"input_examples": [
  { "action": "get", "order_id": "ORD-20481234" },
  { "action": "search", "recipient_phone": "0901234567" }
]
```

**Check:** send the request. If you get a 400 error, an example most likely deviates from the schema, for instance a misspelled field name or an action value outside the enum.

## Step 5: Responses and errors are prompts too

Whatever the API returns goes back into the model's context. Anthropic's engineering blog advises that tools return only high-value information, and suggests pagination, filtering and truncation when payloads are large. The same post notes that error messages can be rewritten to say clearly what needs fixing, helping the model adjust on its next call.

The wrapper below is simplified. It separates two different failures: an ID in the wrong format (caught in the wrapper, before any API call) and an ID in the right format that does not exist (the API returns 404).

```python
import re

# Simplified: wrapper between the model and the customer's API
ORDER_ID_PATTERN = re.compile(r"^ORD-\d{8}$")

def run_acme_orders(args):
    if args["action"] in ("get", "cancel"):
        order_id = args.get("order_id", "")
        if not ORDER_ID_PATTERN.match(order_id):
            return ("order_id has the wrong format: it must be ORD- + 8 digits, "
                    "e.g. ORD-20481234. Ask the user for the order ID again.")
    resp = call_customer_api(args)  # your HTTP function
    if resp.status == 404:
        return ("Order not found even though the ID has the right format.\n\n"
                "If the user has the recipient's phone number, "
                "call again with action='search'.")
    order = resp.json()
    return {"order_id": order["id"], "status": order["status"],
            "eta": order["eta"]}  # drop internal fields the model doesn't need
```

Compared with a bare "404 Not Found" string, each message here points the model down a different path: a malformed ID means asking the user again; a valid but non-existent ID means switching to a phone-number search. Trimming the response to three fields spares the model from wading through dozens of internal fields.

**Check:** make two test calls, one with a malformed ID and one with a well-formed ID that does not exist. The model's next call should differ between the two cases, rather than retrying the same bad ID.

## Step 6: Don't trust your gut, run an eval

Anthropic reports that even small refinements to a description can improve results dramatically. That cuts both ways: changing one sentence can make things better, but it can also break a case that used to work.

The practical approach is to assemble about 20 questions taken from the customer's support logs. For each, record the correct tool, action and arguments. Every time you edit a description, rerun the whole set and count the correct answers.

**Key point:** A description is code: if you change it, rerun the tests.

## Traps to avoid

A common mistake is auto-generating tools from OpenAPI and leaving them as they are. That leads straight to what Anthropic's documentation warns against: dozens of tools, one per endpoint, with descriptions copied from developer comments and too short for the model to tell them apart. Another mistake is leaving parameters named `id`, `user` or `data` with no format stated.

Nor should you pass the API's raw JSON straight back, errors included, and then wonder why the model keeps repeating the same bad call.

## What does this skill look like on a customer site?

On a customer site, the hard part is rarely JSON syntax. It is sitting with the operations team to extract answers to questions like "can an order be cancelled once it has left the warehouse?", because that is exactly what goes into the "when not to use" sentence of the description.

If you are aiming for an FDE role, look for job descriptions that mention tool calling, agent integration or working directly with customer APIs.

On your CV, do not just write "integrated LLMs with internal APIs". Say how many endpoints you merged into how many tools, and how accuracy on your eval set changed across description revisions. A before-and-after number like that shows a hiring manager you can do the work a customer will pay for.

Next time an agent calls the wrong API, don't rush to swap the model. Open the description and read it as a new hire would on their first day.

**Try this week:**

- Pick an API you use, write a tool definition following the template in this article, with a description of at least 4 sentences that states clearly when not to use the tool.
- Assemble 20 real user questions, record the correct tool and arguments for each, then score how many the model gets right before and after each description edit.
- Rewrite that API's three most common error messages so that reading them tells the model what to fix on the next call.

## Sources

- [What are Tools? - Hugging Face Agents Course](https://huggingface.co/learn/agents-course/en/unit1/tools)

- [What is Tool Calling in Agents? (Paragon)](https://www.useparagon.com/blog/ai-building-blocks-what-is-tool-calling-a-guide-for-pms)

- [Function Calling with LLMs | Prompt Engineering Guide](https://www.promptingguide.ai/applications/function_calling)

- [Define tools - Claude API Docs](https://platform.claude.com/docs/en/docs/agents-and-tools/tool-use/implement-tool-use)

- [Writing effective tools for agents — with agents (Anthropic Engineering)](https://www.anthropic.com/engineering/writing-tools-for-agents)
