# Calling the Claude and OpenAI APIs directly: write your own tool-use loop, then compare with Gemini

> Before handing everything to a framework, work through a full cycle of requests, streaming and tool calls with Claude and OpenAI, and learn where Gemini differs. These are the things you will end up debugging at a client's office.

Bản gốc: https://fdetimes.net/en/guides/claude-openai-api-tool-use-loop/

A single parameter is enough to wreck a demo. On recent Claude models, setting `temperature`, `top_p` or `top_k` to anything other than the default makes the API return a 400 error. From the 4.6 line onwards, the prefill technique is also rejected, so code copied from an old blog post can break in front of the client.

Frameworks hide some of these details, but not forever. As an FDE, a client's system can fail at the lowest layer: a wrong header, a mismatched id, a lost conversation turn. The quickest way to learn to read those failures is to call the raw API yourself, at least once.

This guide walks you through a small script with a tool, `get_order_status(order_id)`, that looks up an order's status from a mock function. You will go through the full cycle of request, streaming and tool use with Claude and OpenAI, then compare how Gemini handles the same problem. You need Python, API keys set as environment variables, and the two SDKs `anthropic` and `openai`.

The code below is trimmed for readability and has no error handling or retries. Model names live in two variables, `MODEL` and `OA_MODEL`; fill in each provider's current model name from its docs.

## Step 1: why start with curl?

You need to see the raw request before letting an SDK wrap it. The Claude Messages API accepts `POST /v1/messages`, authenticates with the `x-api-key` header and also needs an `anthropic-version` header. In Anthropic's example, `max_tokens` and the `messages` array are the two required fields.

```bash
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model": "'"$CLAUDE_MODEL"'", "max_tokens": 512,
"messages": [{"role": "user", "content": "Chào bạn"}]}'
```

**Check:** you should get back JSON containing a reply. Now send a second turn containing only the new question. The model will remember nothing of the previous turn, because the Messages API is stateless: you must resend the full conversation history on every turn.

So when a client complains that the model "forgot" the previous turn, this is the first place to look.

## Step 2: streaming should show text gradually

A user staring at a blank screen for ten seconds will assume the system has hung. With Claude, you set `"stream": true` to receive the response in pieces via server-sent events. The Python SDK provides a `text_stream` helper, so you do not have to parse events yourself.

```python
import anthropic
client = anthropic.Anthropic()
history = [{"role": "user", "content": "Giải thích SSE trong 3 câu"}]

with client.messages.stream(model=MODEL, max_tokens=512,
messages=history) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
```

**Check:** text should appear gradually in the terminal. If the whole paragraph appears at once, check whether a proxy or gateway between your machine and the API is buffering the response before passing it on. On a client's internal network, that is a question worth asking early.

## Step 3: the model only proposes; your code runs the function

This is the most important idea in the whole guide. When Claude wants to use a client tool, it returns `stop_reason: "tool_use"` with one or more `tool_use` blocks. Your code runs the function and sends the result back in a `tool_result` block; server tools are different, since they run on Anthropic's infrastructure.

First, write a mock function that returns a string, declare the tool, and write a shared function that calls the API:

```python
def get_order_status(order_id): return f"Đơn {order_id}: đang giao"  # dữ liệu giả

tools = [{"name": "get_order_status",
"description": "Tra trạng thái đơn hàng theo mã",
"input_schema": {"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"]}}]

def ask():
return client.messages.create(model=MODEL, max_tokens=512,
tools=tools, messages=history)
```

Then put the question into `history` and write the loop, with a cap on the number of rounds to avoid looping forever:

```python
history = [{"role": "user", "content": "Đơn A123 tới đâu rồi?"}]
resp = ask()
for _ in range(5):
if resp.stop_reason != "tool_use":
break
history.append({"role": "assistant", "content": resp.content})
results = [{"type": "tool_result", "tool_use_id": b.id,
"content": get_order_status(**b.input)}
for b in resp.content if b.type == "tool_use"]
history.append({"role": "user", "content": results})
resp = ask()
```

**Check:** add a log line inside `get_order_status`. If the log line appears and the final answer contains the "đang giao" ("in transit") returned by the function, the loop is working.

Remember also that tools are not free: the API automatically inserts a hidden system prompt to enable tool use, and the `tools` parameter and the `tool_use` and `tool_result` blocks all count as tokens.

## Step 4: OpenAI Responses changes the names, not the logic

On the Responses API the loop is the same; only the names change. The function call sits in the `output` array as an item whose `type` is `function_call`, together with a `call_id`. You return the result as a `function_call_output` item carrying that same `call_id`.

Declare the tool in OpenAI's format, reusing the schema and mock function from step 3:

```python
import json
from openai import OpenAI
oa = OpenAI()
oa_tools = [{"type": "function", "name": "get_order_status",
"description": "Tra trạng thái đơn hàng theo mã",
"parameters": tools[0]["input_schema"]}]
inputs = [{"role": "user", "content": "Đơn A123 tới đâu rồi?"}]
```

Then run one round of function calls and send the results back:

```python
resp = oa.responses.create(model=OA_MODEL, input=inputs, tools=oa_tools)
for item in resp.output:
if item.type == "function_call":
args = json.loads(item.arguments)
inputs.append(item)
inputs.append({"type": "function_call_output",
"call_id": item.call_id,
"output": get_order_status(**args)})
resp = oa.responses.create(model=OA_MODEL, input=inputs, tools=oa_tools)
```

Note that this snippet handles exactly one round of function calls, unlike the capped loop in step 3. If the model proposes further calls after receiving the results, wrap it in a similar `for` loop that stops when `output` contains no more `function_call` items.

Streaming on the Responses API is also switched on with `stream=True`, but what you receive are semantic events, each with its own type. Before writing any handling logic, print every `event.type` to see what the event stream looks like:

```python
for event in oa.responses.create(model=OA_MODEL, input=inputs,
tools=oa_tools, stream=True):
print(event.type)
```

**Check:** you will see a sequence of events, each kind with its own name. Pick the event that carries new text, and check its name against the docs rather than guessing.

## Step 5: Gemini is changing APIs, so do not copy old code

With Gemini, the trap is choosing which API to use. The current docs recommend the Interactions API for new projects and classify `generateContent` as legacy, so tutorials online that use `generateContent` may be out of date. When calling over REST, the key is passed in the `x-goog-api-key` header.

```bash
# Khung rút gọn: endpoint và body lấy từ trang Interactions API
curl "$GEMINI_ENDPOINT" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "content-type: application/json" \
-d @request.json
```

(The comment notes that this is a trimmed skeleton: take the endpoint and body from the Interactions API page.)

This guide does not provide a ready-made tool-use loop for Gemini, because you should take the Interactions API's request structure straight from the docs. The principle is unchanged: Gemini does not run functions itself; it only proposes a call, and your application runs the function and sends the result back.

By default, the response is returned only once the model has finished generating, so you must enable streaming to receive it in pieces.

**Extension exercise:** rebuild the `get_order_status` loop on the Interactions API following the docs, with the same question about order A123, then compare how Gemini proposes function calls with the other two providers.

## Three APIs, one comparison table

| | Claude Messages | OpenAI Responses | Gemini |
|---|---|---|---|
| Streaming | `stream: true`, SSE | `stream=True`, semantic events | Enable streaming to receive pieces |
| Tool call signal | `stop_reason: "tool_use"` | `function_call` item in `output` | Model proposes a function call |
| Returning results | `tool_result` + `tool_use_id` | `function_call_output` + `call_id` | App sends the result back |
| History | Client resends everything | Can keep state on the server | Depends on the API you choose |

**Điểm mấu chốt:** With all three providers, the model only proposes a call and your code is what runs the function, so when a tool-use loop breaks, the first place to look is your own code.

## The failures you will meet on real projects

The most common is stale sampling configuration: if a client's config still has `temperature: 0.2` from an earlier model, a recent Claude model will return a 400. Next comes forgetting the id: when `tool_result` lacks `tool_use_id` or `function_call_output` lacks `call_id`, the model cannot tell which call a result belongs to. Then there is stuffing a few dozen tools into every request and being surprised when the token bill soars.

There is also an architectural decision you should spell out to the client. Anthropic positions the Messages API for self-written agent loops that need fine-grained control, while OpenAI says the Responses API preserves reasoning across turns, which Chat Completions cannot do, and runs hosted tools on the server side.

Simon Willison has also pointed out that Responses can manage conversation state on the server on your behalf.

As an FDE, the sensible move is to ask the client two questions in the first meeting: where is conversation history allowed to live, and who will run the tools?

A bank that wants to keep all logs inside its own systems will suit the stateless model better, while a startup that needs to ship quickly may prefer to hand some of the state management to the server.

## Turn the exercise into evidence on your CV

A line saying "experience with LLMs" on a CV does little to prove you understand the layer beneath the framework, yet that is the layer you will have to debug when the framework breaks at a client's site. Push this script to GitHub with Claude and OpenAI adapters sharing a common `call_tool` interface, plus a README recording each 400 error you hit and how you fixed it.

On your CV, state specifically that you wrote tool-use loops yourself on Claude Messages and OpenAI Responses, and if you did the extension exercise, add the Gemini Interactions API too.

When reading job descriptions, look for phrases such as "integrate with multiple model providers" or "build agent loops". That is exactly the exercise you have just done, and you can now describe it in terms of specific fields rather than just framework names.

**Thử ngay tuần này:**

- Call Claude with curl using the x-api-key and anthropic-version headers, then send a second turn that deliberately omits the history to see the model forget the previous turn
- Write a mock get_order_status tool and get it working with both Claude and OpenAI Responses, then print every event.type with stream=True enabled
- Push a small repo to GitHub with Claude and OpenAI adapters and a README listing each error you hit and how you fixed it, then link it from your CV

## Nguồn

- [Using the Messages API](https://platform.claude.com/docs/en/build-with-claude/working-with-messages)

- [Streaming messages](https://platform.claude.com/docs/en/build-with-claude/streaming)

- [Tool use with Claude](https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview)

- [Streaming API responses](https://developers.openai.com/api/docs/guides/streaming-responses)

- [Function calling](https://developers.openai.com/api/docs/guides/function-calling)

- [Why we built the Responses API](https://developers.openai.com/blog/responses-api/)

- [OpenAI API: Responses vs. Chat Completions](https://simonwillison.net/2025/Mar/11/responses-vs-chat-completions/)

- [Function calling with the Gemini API](https://ai.google.dev/gemini-api/docs/function-calling)

- [Gemini API | Google AI for Developers](https://ai.google.dev/gemini-api/docs)

- [Text generation | Gemini API | Google AI for Developers](https://ai.google.dev/gemini-api/docs/text-generation)
