Calling the Claude and OpenAI APIs directly: write your own tool-use loop, then compare with Gemini
Before handing everything to a framework, work through a full cycle of requests, streaming and tool calls with Claude and OpenAI, and learn where Gemini differs. These are the things you will end up debugging at a client's office.
In brief
- The Claude Messages API is stateless: you must resend the full history on every turn, and remember to return tool_result with its tool_use_id
- The OpenAI Responses API returns a function_call with a call_id in the output array; you send the result back as a function_call_output
- For new projects Gemini recommends the Interactions API and treats generateContent as legacy; with all three providers the model only proposes a call, and your code runs the function
A single parameter is enough to wreck a demo. On recent Claude models, setting temperature, top_p or top_k to anything other than the default makes the API return a 400 error. From the 4.6 line onwards, the prefill technique is also rejected, so code copied from an old blog post can break in front of the client.
Frameworks hide some of these details, but not forever. As an FDE, a client’s system can fail at the lowest layer: a wrong header, a mismatched id, a lost conversation turn. The quickest way to learn to read those failures is to call the raw API yourself, at least once.
This guide walks you through a small script with a tool, get_order_status(order_id), that looks up an order’s status from a mock function. You will go through the full cycle of request, streaming and tool use with Claude and OpenAI, then compare how Gemini handles the same problem. You need Python, API keys set as environment variables, and the two SDKs anthropic and openai.
The code below is trimmed for readability and has no error handling or retries. Model names live in two variables, MODEL and OA_MODEL; fill in each provider’s current model name from its docs.
Step 1: why start with curl?
You need to see the raw request before letting an SDK wrap it. The Claude Messages API accepts POST /v1/messages, authenticates with the x-api-key header and also needs an anthropic-version header. In Anthropic’s example, max_tokens and the messages array are the two required fields.
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model": "'"$CLAUDE_MODEL"'", "max_tokens": 512,
"messages": [{"role": "user", "content": "Chào bạn"}]}'
Check: you should get back JSON containing a reply. Now send a second turn containing only the new question. The model will remember nothing of the previous turn, because the Messages API is stateless: you must resend the full conversation history on every turn.
So when a client complains that the model “forgot” the previous turn, this is the first place to look.
Step 2: streaming should show text gradually
A user staring at a blank screen for ten seconds will assume the system has hung. With Claude, you set "stream": true to receive the response in pieces via server-sent events. The Python SDK provides a text_stream helper, so you do not have to parse events yourself.
import anthropic
client = anthropic.Anthropic()
history = [{"role": "user", "content": "Giải thích SSE trong 3 câu"}]
with client.messages.stream(model=MODEL, max_tokens=512,
messages=history) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
Check: text should appear gradually in the terminal. If the whole paragraph appears at once, check whether a proxy or gateway between your machine and the API is buffering the response before passing it on. On a client’s internal network, that is a question worth asking early.
Step 3: the model only proposes; your code runs the function
This is the most important idea in the whole guide. When Claude wants to use a client tool, it returns stop_reason: "tool_use" with one or more tool_use blocks. Your code runs the function and sends the result back in a tool_result block; server tools are different, since they run on Anthropic’s infrastructure.
First, write a mock function that returns a string, declare the tool, and write a shared function that calls the API:
def get_order_status(order_id): return f"Đơn {order_id}: đang giao" # dữ liệu giả
tools = [{"name": "get_order_status",
"description": "Tra trạng thái đơn hàng theo mã",
"input_schema": {"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"]}}]
def ask():
return client.messages.create(model=MODEL, max_tokens=512,
tools=tools, messages=history)
Then put the question into history and write the loop, with a cap on the number of rounds to avoid looping forever:
history = [{"role": "user", "content": "Đơn A123 tới đâu rồi?"}]
resp = ask()
for _ in range(5):
if resp.stop_reason != "tool_use":
break
history.append({"role": "assistant", "content": resp.content})
results = [{"type": "tool_result", "tool_use_id": b.id,
"content": get_order_status(**b.input)}
for b in resp.content if b.type == "tool_use"]
history.append({"role": "user", "content": results})
resp = ask()
Check: add a log line inside get_order_status. If the log line appears and the final answer contains the “đang giao” (“in transit”) returned by the function, the loop is working.
Remember also that tools are not free: the API automatically inserts a hidden system prompt to enable tool use, and the tools parameter and the tool_use and tool_result blocks all count as tokens.
Step 4: OpenAI Responses changes the names, not the logic
On the Responses API the loop is the same; only the names change. The function call sits in the output array as an item whose type is function_call, together with a call_id. You return the result as a function_call_output item carrying that same call_id.
Declare the tool in OpenAI’s format, reusing the schema and mock function from step 3:
import json
from openai import OpenAI
oa = OpenAI()
oa_tools = [{"type": "function", "name": "get_order_status",
"description": "Tra trạng thái đơn hàng theo mã",
"parameters": tools[0]["input_schema"]}]
inputs = [{"role": "user", "content": "Đơn A123 tới đâu rồi?"}]
Then run one round of function calls and send the results back:
resp = oa.responses.create(model=OA_MODEL, input=inputs, tools=oa_tools)
for item in resp.output:
if item.type == "function_call":
args = json.loads(item.arguments)
inputs.append(item)
inputs.append({"type": "function_call_output",
"call_id": item.call_id,
"output": get_order_status(**args)})
resp = oa.responses.create(model=OA_MODEL, input=inputs, tools=oa_tools)
Note that this snippet handles exactly one round of function calls, unlike the capped loop in step 3. If the model proposes further calls after receiving the results, wrap it in a similar for loop that stops when output contains no more function_call items.
Streaming on the Responses API is also switched on with stream=True, but what you receive are semantic events, each with its own type. Before writing any handling logic, print every event.type to see what the event stream looks like:
for event in oa.responses.create(model=OA_MODEL, input=inputs,
tools=oa_tools, stream=True):
print(event.type)
Check: you will see a sequence of events, each kind with its own name. Pick the event that carries new text, and check its name against the docs rather than guessing.
Step 5: Gemini is changing APIs, so do not copy old code
With Gemini, the trap is choosing which API to use. The current docs recommend the Interactions API for new projects and classify generateContent as legacy, so tutorials online that use generateContent may be out of date. When calling over REST, the key is passed in the x-goog-api-key header.
# Khung rút gọn: endpoint và body lấy từ trang Interactions API
curl "$GEMINI_ENDPOINT" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "content-type: application/json" \
-d @request.json
(The comment notes that this is a trimmed skeleton: take the endpoint and body from the Interactions API page.)
This guide does not provide a ready-made tool-use loop for Gemini, because you should take the Interactions API’s request structure straight from the docs. The principle is unchanged: Gemini does not run functions itself; it only proposes a call, and your application runs the function and sends the result back.
By default, the response is returned only once the model has finished generating, so you must enable streaming to receive it in pieces.
Extension exercise: rebuild the get_order_status loop on the Interactions API following the docs, with the same question about order A123, then compare how Gemini proposes function calls with the other two providers.
Three APIs, one comparison table
| Claude Messages | OpenAI Responses | Gemini | |
|---|---|---|---|
| Streaming | stream: true, SSE |
stream=True, semantic events |
Enable streaming to receive pieces |
| Tool call signal | stop_reason: "tool_use" |
function_call item in output |
Model proposes a function call |
| Returning results | tool_result + tool_use_id |
function_call_output + call_id |
App sends the result back |
| History | Client resends everything | Can keep state on the server | Depends on the API you choose |
The failures you will meet on real projects
The most common is stale sampling configuration: if a client’s config still has temperature: 0.2 from an earlier model, a recent Claude model will return a 400. Next comes forgetting the id: when tool_result lacks tool_use_id or function_call_output lacks call_id, the model cannot tell which call a result belongs to. Then there is stuffing a few dozen tools into every request and being surprised when the token bill soars.
There is also an architectural decision you should spell out to the client. Anthropic positions the Messages API for self-written agent loops that need fine-grained control, while OpenAI says the Responses API preserves reasoning across turns, which Chat Completions cannot do, and runs hosted tools on the server side.
Simon Willison has also pointed out that Responses can manage conversation state on the server on your behalf.
As an FDE, the sensible move is to ask the client two questions in the first meeting: where is conversation history allowed to live, and who will run the tools?
A bank that wants to keep all logs inside its own systems will suit the stateless model better, while a startup that needs to ship quickly may prefer to hand some of the state management to the server.
Turn the exercise into evidence on your CV
A line saying “experience with LLMs” on a CV does little to prove you understand the layer beneath the framework, yet that is the layer you will have to debug when the framework breaks at a client’s site. Push this script to GitHub with Claude and OpenAI adapters sharing a common call_tool interface, plus a README recording each 400 error you hit and how you fixed it.
On your CV, state specifically that you wrote tool-use loops yourself on Claude Messages and OpenAI Responses, and if you did the extension exercise, add the Gemini Interactions API too.
When reading job descriptions, look for phrases such as “integrate with multiple model providers” or “build agent loops”. That is exactly the exercise you have just done, and you can now describe it in terms of specific fields rather than just framework names.
10 sources
- Using the Messages API
- Streaming messages
- Tool use with Claude
- Streaming API responses
- Function calling
- Why we built the Responses API
- OpenAI API: Responses vs. Chat Completions · 2025-03-11
- Function calling with the Gemini API
- Gemini API | Google AI for Developers
- Text generation | Gemini API | Google AI for Developers