Guide · Developer tools
OpenAI vs Anthropic vs Gemini API message formats compared
OpenAI’s Chat Completions puts the system prompt in the message list and tool results in tool messages, while its Responses API uses instructions and typed items. Anthropic’s Messages API takes a top-level system and tool_use/tool_result blocks. Gemini’s generateContent uses systemInstruction, the roles user and model, and functionCall/functionResponse parts.
By Tahir NazirUpdated 13 min read
On this page
- How do the OpenAI, Anthropic and Gemini message formats differ?
- The same conversation in all four formats
- Where does the system prompt go?
- How do you send an image to each API?
- How do tool calls and tool results compare?
- How do the streaming events differ?
- How do the token usage fields compare?
- How to convert OpenAI messages to Anthropic or Gemini
- Questions people ask
How do the OpenAI, Anthropic and Gemini message formats differ?
All four APIs take a list of turns and return the model’s next turn, but they disagree on where the system prompt lives, which roles exist, how images and tool calls are encoded, and how tokens are reported. OpenAI has two: Chat Completions, a flat list of messages, and the newer Responses API, which models a conversation as typed items. Here is one conversation in all four shapes:
- 1System prompt
- 2User: text and a photo
- 3Model: calls check_stock
- 4Your app: the tool result
OpenAI Chat Completions
POST /v1/chat/completions
"role": "developer"A message like any other (“system” on older models)
"type": "image_url"Image as a URL or data: URL, next to a text part
"tool_calls": [{ "id" }]On the assistant message; arguments are a JSON string
"role": "tool"One message per call, matched by tool_call_id
OpenAI Responses
POST /v1/responses
"instructions"Not carried over by previous_response_id
"type": "input_image"Inside a user message, next to input_text
"type": "function_call"Its own item, not part of a message; has a call_id
"type": "function_call_output"Its own item, matched by call_id
Anthropic Messages
POST /v1/messages
"system"Messages hold user and assistant turns
"type": "image"A content block; base64 split into media_type and data
"type": "tool_use"In the assistant turn; input is a JSON object
"type": "tool_result"In a user turn, before any text, with tool_use_id
Gemini generateContent
POST /v1beta/models/{model}:generateContent
"systemInstruction"Text only; there is no system role
"inlineData"A part with mimeType and base64 data
"functionCall"Role “model”; args is an object; keep the thoughtSignature
"functionResponse"In a user turn; response must be a JSON object
| Chat Completions | Responses | Anthropic Messages | Gemini generateContent | |
|---|---|---|---|---|
| Turns go in | messages | input (a string or items) | messages | contents |
| Roles | developer, system, user, assistant, tool | developer, system, user, assistant | user, assistant | user, model |
| A turn’s pieces | Content parts | Content parts and typed items | Content blocks | Parts |
| Must also send | messages | Usually input (optional with previous_response_id) | messages, max_tokens | contents (the model goes in the URL) |
| Reply text is in | choices[0].message.content | output items (SDK: output_text) | content blocks | candidates[0].content.parts |
| Stop reason for a tool call | finish_reason: "tool_calls" | A function_call item in output | stop_reason: "tool_use" | None specific: look for functionCall parts |
From each provider’s API reference, 2026-10-11. Anthropic’s newer models also accept `system` turns mid-conversation; see below.
OpenAI recommends Responses for new projects and keeps Chat Completions supported. Google now recommends its newer Interactions API for new development and says generateContent “remains fully supported”; this guide covers generateContent, which Google still documents in full.
The same conversation in all four formats
The example is a bike-shop assistant. The system prompt tells it to check stock, the user sends a photo and asks what the part is, the model calls a check_stock tool, and your app returns the result. Each body below is the request that sends that tool result back, so it holds the whole history. The base64 image data is shortened.
{
"model": "gpt-6-astra",
"messages": [
{
"role": "developer",
"content": "You are a support assistant for a bike shop. Check stock before you answer."
},
{
"role": "user",
"content": [
{ "type": "text", "text": "What part is this, and do you have it in stock?" },
{ "type": "image_url", "image_url": { "url": "data:image/png;base64,iVBORw0KGgo..." } }
]
},
{
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_abc123",
"type": "function",
"function": { "name": "check_stock", "arguments": "{\"part\":\"rear derailleur\"}" }
}
]
},
{
"role": "tool",
"tool_call_id": "call_abc123",
"content": "{\"in_stock\":true,\"quantity\":4}"
}
],
"tools": [
{
"type": "function",
"function": {
"name": "check_stock",
"description": "Look up how many of a part the shop has in stock.",
"parameters": {
"type": "object",
"properties": { "part": { "type": "string" } },
"required": ["part"],
"additionalProperties": false
},
"strict": true
}
}
]
}OpenAI’s reference says developer messages replace system messages from o1 onwards. The call’s arguments is a string of JSON, and the result goes in a separate tool message whose tool_call_id matches the call’s id.
{
"model": "gpt-6-astra",
"instructions": "You are a support assistant for a bike shop. Check stock before you answer.",
"input": [
{
"role": "user",
"content": [
{ "type": "input_text", "text": "What part is this, and do you have it in stock?" },
{ "type": "input_image", "image_url": "data:image/png;base64,iVBORw0KGgo...", "detail": "auto" }
]
},
{
"type": "function_call",
"call_id": "call_abc123",
"name": "check_stock",
"arguments": "{\"part\":\"rear derailleur\"}"
},
{
"type": "function_call_output",
"call_id": "call_abc123",
"output": "{\"in_stock\":true,\"quantity\":4}"
}
],
"tools": [
{
"type": "function",
"name": "check_stock",
"description": "Look up how many of a part the shop has in stock.",
"parameters": {
"type": "object",
"properties": { "part": { "type": "string" } },
"required": ["part"],
"additionalProperties": false
},
"strict": true
}
]
}In Responses the call and its output are items of their own, linked by call_id (the call also has a separate item id). In real code you append everything from the previous response.output, including any reasoning items, rather than rebuilding the call by hand. Or let OpenAI keep the history: responses are stored by default, and previous_response_id chains them, but it doesn’t carry instructions over, so send them every time.
{
"model": "claude-opus-5-5",
"max_tokens": 1024,
"system": "You are a support assistant for a bike shop. Check stock before you answer.",
"messages": [
{
"role": "user",
"content": [
{ "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "iVBORw0KGgo..." } },
{ "type": "text", "text": "What part is this, and do you have it in stock?" }
]
},
{
"role": "assistant",
"content": [
{ "type": "tool_use", "id": "toolu_01A2b3C4", "name": "check_stock", "input": { "part": "rear derailleur" } }
]
},
{
"role": "user",
"content": [
{ "type": "tool_result", "tool_use_id": "toolu_01A2b3C4", "content": "{\"in_stock\":true,\"quantity\":4}" }
]
}
],
"tools": [
{
"name": "check_stock",
"description": "Look up how many of a part the shop has in stock.",
"input_schema": {
"type": "object",
"properties": { "part": { "type": "string" } },
"required": ["part"]
}
}
]
}Anthropic requires max_tokens. The image goes first because Anthropic says Claude works best when images come before text. The tool result is a block inside a user turn, not a separate role.
{
"systemInstruction": {
"parts": [{ "text": "You are a support assistant for a bike shop. Check stock before you answer." }]
},
"contents": [
{
"role": "user",
"parts": [
{ "inlineData": { "mimeType": "image/png", "data": "iVBORw0KGgo..." } },
{ "text": "What part is this, and do you have it in stock?" }
]
},
{
"role": "model",
"parts": [
{
"functionCall": { "id": "8f2b1a3c", "name": "check_stock", "args": { "part": "rear derailleur" } },
"thoughtSignature": "<signature from the model's reply>"
}
]
},
{
"role": "user",
"parts": [
{
"functionResponse": { "id": "8f2b1a3c", "name": "check_stock", "response": { "in_stock": true, "quantity": 4 } }
}
]
}
],
"tools": [
{
"functionDeclarations": [
{
"name": "check_stock",
"description": "Look up how many of a part the shop has in stock.",
"parameters": {
"type": "object",
"properties": { "part": { "type": "string" } },
"required": ["part"]
}
}
]
}
]
}Gemini names the model in the URL. The model’s turn is replayed exactly as it arrived, including the thoughtSignature (shown here as a placeholder), and the result’s response is a JSON object, not a string. Our Gemini playground uses the same format, streamed, with your own key.
Where does the system prompt go?
In Chat Completions it’s a message; everywhere else it’s a separate field:
- Chat Completions: a
developer(orsystem) message, which can appear anywhere inmessages. - Responses: the top-level
instructionsstring, ordeveloper/systemmessage items if you’re carrying over an old transcript. - Anthropic: the top-level
system, a string or an array of text blocks. Newer models (Claude Opus 5.5, Sonnet 5.5 and Haiku 5.5 among them) also acceptsystemturns later inmessages, but never as the first entry. - Gemini: the top-level
systemInstruction, a content object whose parts must be text. Gemini has no system role.
When converting from OpenAI, collect every system and developer message and join them into one field. Anthropic’s own OpenAI compatibility layer does exactly that: it concatenates them with a single newline and sends the result as system, wherever they appeared in the conversation.
How do you send an image to each API?
OpenAI takes one string, a URL or a data: URL; Anthropic and Gemini want the media type and the base64 data in separate fields:
| API | Part | Inline bytes | Other sources |
|---|---|---|---|
| Chat Completions | {"type": "image_url", "image_url": {"url": …}} | data:image/png;base64,… in url | An https URL; optional detail |
| Responses | {"type": "input_image", "image_url": …} | A data: URL in image_url | An https URL or a file_id; detail |
| Anthropic | {"type": "image", "source": {…}} | source: type: base64, media_type, data | type: url or type: file with a file_id |
| Gemini | {"inlineData": {…}} | mimeType and data | fileData with a fileUri: a Files API URI, a registered Cloud Storage URI, or a public or pre-signed HTTPS URL |
OpenAI and Anthropic accept PNG, JPEG, WebP and GIF (OpenAI non-animated only; Anthropic reads an animation’s first frame). Google’s image guide lists PNG, JPEG, WebP, HEIC and HEIF, and its API reference adds GIF and AVIF.
Every image costs input tokens, and each provider has its own formula, so check the usage fields of a real request rather than guessing.
How do tool calls and tool results compare?
The loop is the same everywhere: you declare tools, the model asks for one, you run it and send the result back with the call’s ID. The field names differ at every step:
| Step | Chat Completions | Responses | Anthropic | Gemini |
|---|---|---|---|---|
| Declare | {"type": "function", "function": {name, parameters}} | {"type": "function", name, parameters} | {name, description, input_schema} | tools[].functionDeclarations[] with parameters |
| Call arrives as | message.tool_calls[] | A function_call item | A tool_use block | A functionCall part |
| Arguments | function.arguments, a JSON string | arguments, a JSON string | input, an object | args, an object |
| ID to echo | id → tool_call_id | call_id → call_id | id → tool_use_id | id → id |
| Result goes in | A tool message | A function_call_output item | A tool_result block in a user turn | A functionResponse part in a user turn |
| Reporting a failure | Say so in content | Say so in output | is_error: true | An error key in response |
| Force a tool call | tool_choice: "required" | tool_choice: "required" | tool_choice: {"type": "any"} | toolConfig.functionCallingConfig.mode: "ANY" |
OpenAI says its function definitions are “externally tagged” in Chat Completions and “internally tagged” in Responses, hence the nesting difference.
Schema strictness differs too. Chat Completions is non-strict unless you set strict: true; Responses tries strict mode when you leave strict out and falls back if your schema can’t comply; Anthropic guarantees schema-valid inputs only for tools that set strict: true. OpenAI’s reference also warns that the model “does not always generate valid JSON” for arguments, so parse them defensively.
How do the streaming events differ?
All four stream over server-sent events, but Chat Completions sends unnamed chunks, Responses and Anthropic send named, typed events, and Gemini sends a series of partial responses:
| API | Turn it on | Text arrives in | Tool arguments arrive in | Usage | Finished when |
|---|---|---|---|---|---|
| Chat Completions | "stream": true | choices[0].delta.content | delta.tool_calls[].function.arguments fragments, by index | A final chunk with empty choices, only if stream_options.include_usage is set | data: [DONE] |
| Responses | "stream": true | response.output_text.delta | response.function_call_arguments.delta, then .done | In response.completed | response.completed |
| Anthropic | "stream": true | content_block_delta with text_delta | input_json_delta holding partial_json | Input in message_start; cumulative output in message_delta | message_stop |
| Gemini | :streamGenerateContent?alt=sse | Each event’s candidates[0].content.parts | functionCall parts | usageMetadata on the response objects | A chunk with finishReason |
Anthropic streams may also contain `ping` and `error` events, and Anthropic says new event types may be added, so ignore ones you don’t know.
Two details catch people out. Anthropic’s tool input streams as partial JSON strings: accumulate them and parse once the block’s content_block_stop arrives (or use a partial-JSON parser or the SDK’s streaming helpers); the final input is always an object. And Google warns that when a Gemini reply has no function call, its thought signature may arrive in a final part with empty text, so read the stream until finishReason appears.
How do the token usage fields compare?
The names differ, and so does what they include. The trap is caching: Anthropic’s input_tokens excludes cached tokens, while OpenAI and Gemini count them inside the prompt total.
| API | Input | Output | Cached input | Reasoning | Total |
|---|---|---|---|---|---|
| Chat Completions | usage.prompt_tokens | completion_tokens | prompt_tokens_details.cached_tokens (part of the input) | completion_tokens_details.reasoning_tokens | total_tokens |
| Responses | usage.input_tokens | output_tokens | input_tokens_details.cached_tokens and cache_write_tokens (part of the input) | output_tokens_details.reasoning_tokens | total_tokens |
| Anthropic | usage.input_tokens (uncached only) | output_tokens | cache_read_input_tokens and cache_creation_input_tokens (separate) | output_tokens_details.thinking_tokens | Not given: add them up |
| Gemini | usageMetadata.promptTokenCount | candidatesTokenCount | cachedContentTokenCount (part of the input) | thoughtsTokenCount | totalTokenCount |
Anthropic: total input = `cache_read_input_tokens` + `cache_creation_input_tokens` + `input_tokens`. Gemini: `totalTokenCount` = prompt + thoughts + candidates, so thoughts are not inside `candidatesTokenCount`.
If you log usage from several providers in one table, normalise it to total input, cached input and output before you compare or bill. The LLM cost calculator turns those counts into money, and prompt caching explained covers how the cached share is priced. The tokens themselves aren’t comparable either: each family has its own tokenizer, as counting tokens in Python and JavaScript shows, and the token counter compares them on your own text.
How to convert OpenAI messages to Anthropic or Gemini
Most conversion bugs come from the same handful of differences. Check each one:
- Hoist the system prompt into
systemorsystemInstruction, joining several messages into one. - Merge same-role turns. Anthropic combines consecutive
userorassistantturns into one, but in a user turn that carries tool results, thetool_resultblocks must come before any text. OpenAI’s separatetoolmessages therefore become one user turn. - Keep parallel results together for Gemini. After two calls, send both
functionResponseparts in one user turn. Google says interleaving calls and responses returns a 400 error. - Copy tool call IDs exactly. In Responses use
call_id, not the item’sid. Anthropic IDs must match^[a-zA-Z0-9_-]+$. Gemini 3 always returns anid; older histories may lack one, so generate your own for the other APIs. - Parse and wrap arguments and results. OpenAI’s
argumentsstring becomes an object; Gemini’sresponsemust be an object, so wrap plain text, as Google’s own example does with{"result": …}. - Split `data:` URLs into media type and base64 data for Anthropic and Gemini.
- Handle Gemini 3 thought signatures. Gemini 3 rejects a function-call turn without its signature. For calls Gemini didn’t generate (a history from another model), Google documents the placeholder
skip_thought_signature_validator. - Rename the schema field from
parameterstoinput_schemafor Anthropic, and addmax_tokens.
Here is a converter from Chat Completions messages to Anthropic’s system and messages. We ran it on the example above and on a conversation with two parallel calls and a follow-up question, and both outputs passed the same Anthropic schema check (the second became one user turn with two tool_result blocks, then the text):
import json
def openai_to_anthropic(messages):
"""Chat Completions messages -> Anthropic Messages (system, messages)."""
system, out = [], []
def add(role, blocks):
# Anthropic wants alternating turns: merge into the previous one if same role.
if out and out[-1]["role"] == role:
out[-1]["content"].extend(blocks)
else:
out.append({"role": role, "content": blocks})
for m in messages:
role, content = m["role"], m.get("content")
if role in ("system", "developer"):
system.append(content if isinstance(content, str) else "".join(p["text"] for p in content))
elif role == "tool":
add("user", [{"type": "tool_result", "tool_use_id": m["tool_call_id"], "content": content}])
else:
parts = [{"type": "text", "text": content}] if isinstance(content, str) else (content or [])
blocks = []
for p in parts:
if p["type"] == "text" and p["text"]: # Anthropic rejects empty text blocks
blocks.append({"type": "text", "text": p["text"]})
elif p["type"] == "image_url":
url = p["image_url"]["url"]
if url.startswith("data:"):
media_type, data = url[5:].split(";base64,", 1)
blocks.append({"type": "image", "source": {"type": "base64", "media_type": media_type, "data": data}})
else:
blocks.append({"type": "image", "source": {"type": "url", "url": url}})
for call in m.get("tool_calls") or []:
blocks.append({
"type": "tool_use",
"id": call["id"],
"name": call["function"]["name"],
"input": json.loads(call["function"]["arguments"]), # a string in OpenAI, an object here
})
add(role, blocks)
# tool_result blocks must come first in their user turn
for turn in out:
if turn["role"] == "user":
turn["content"].sort(key=lambda b: b["type"] != "tool_result")
return "\n".join(system), outIt covers text, images and function tools only; audio, files, refusals and reasoning need their own handling. If json.loads fails on a model’s arguments, repair the JSON first and read why LLMs return invalid JSON.
You can also skip conversion. Anthropic offers an OpenAI SDK compatibility layer, but describes it as “primarily intended to test and compare model capabilities” and ignores strict. Google serves Gemini through the OpenAI libraries at https://generativelanguage.googleapis.com/v1beta/openai/. Native formats give you each provider’s full feature set.
FAQ
Questions people ask
Can I use the OpenAI SDK with Claude?
Yes. Anthropic offers an OpenAI SDK compatibility layer: point the SDK at https://api.anthropic.com/v1/ with a Claude API key. Anthropic calls it primarily a way to test and compare models, not a long-term production solution. It hoists all system and developer messages into one system field and ignores strict on tools, so use the native Messages API for production features.
Can I use the OpenAI SDK with Gemini?
Yes. Google documents OpenAI library support for Gemini: set the base URL to https://generativelanguage.googleapis.com/v1beta/openai/ and use a Gemini API key, then call Chat Completions as usual. It’s the quickest way to try Gemini in OpenAI-shaped code, and Google says Gemini 3’s thought signatures work through it too. The native API documents every Gemini feature.
Should I use Chat Completions or the Responses API?
OpenAI recommends Responses for new projects and keeps Chat Completions supported. OpenAI says reasoning models perform better through Responses, and that from GPT-5.4 Chat Completions doesn’t support tool calling with a reasoning_effort other than none. Remember that Responses stores responses by default; set store: false if you don’t want that.
Why does Claude say “tool_use ids were found without tool_result blocks immediately after”?
Every tool_use block in an assistant turn needs a matching tool_result in the very next user turn, with nothing in between. The usual causes are a missing result for one of several parallel calls, results split across turns, or text placed before the tool_result blocks. Put all results first in one user message, each with the right tool_use_id.
Does Gemini have a system role?
No. Gemini’s contents accept only the roles user and model. The system prompt goes in the separate systemInstruction field, which holds text parts only. Tool results also go in a user turn, as functionResponse parts, rather than in a tool role.
Why does Gemini 3 return a 400 error after a function call?
Usually a missing thought signature. Gemini 3 attaches a thoughtSignature to the first function call part of each step, and you must send that part back unchanged with the rest of the turn. Replay the model’s content exactly as received rather than rebuilding it. For calls from another model, Google documents a placeholder value that skips the check.
Is Gemini’s generateContent API deprecated?
No. Google’s migration guide says generateContent “remains fully supported”, while recommending the newer Interactions API for new development. The Interactions API adds server-side history and typed steps. Google’s generateContent guides are now headed “Gemini Generate Content API (Legacy)”, but they’re still published and existing generateContent code keeps working.
Try it
Tools from this guide
- JSON RepairFix broken JSON from LLM output.
- Model ID ReferenceCopy current API model IDs for Claude, GPT, Gemini and more.
- AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- API Key CheckerCheck whether an API key works, without storing it.
- Gemini API PlaygroundTry the Gemini API free with your own key.
Keep reading