Every coding agent speaks one of three wire formats: OpenAI Chat Completions, the OpenAI Responses API, or Anthropic Messages. They all do the same job, and they disagree on nearly every detail. This post compares the OpenAI and Anthropic API formats field by field, explains how one gateway can expose the same model in more than one of them, and maps which agents use which.
POST /v1/chat/completions messages[] incl. system -> choices[0].message
POST /v1/responses instructions + input[] -> output[] items
POST /v1/messages system + messages[] -> content[] blocksThe same request in three formats#
A system prompt, one user message, a 200-token output limit.
{
"model": "deepseek/deepseek-v4.1-flash",
"max_tokens": 200,
"messages": [
{ "role": "system", "content": "You are a terse code reviewer." },
{ "role": "user", "content": "Is `except: pass` ever fine?" }
]
}{
"model": "deepseek/deepseek-v4.1-flash",
"max_output_tokens": 200,
"instructions": "You are a terse code reviewer.",
"input": "Is `except: pass` ever fine?"
}{
"model": "deepseek/deepseek-v4.1-flash",
"max_tokens": 200,
"system": "You are a terse code reviewer.",
"messages": [{ "role": "user", "content": "Is `except: pass` ever fine?" }]
}The replies differ just as much. Chat Completions returns choices[0].message.content. Responses returns an output array of typed items, where the text sits in an item of type message with output_text content (the SDKs add an output_text convenience property). Messages returns a content array of blocks, usually starting with {"type": "text", "text": "..."}.
Where the system prompt goes#
Chat Completions puts it in the message list as a message with role system. OpenAI's newer models use a developer role for the same purpose. That causes friction with non-OpenAI servers, which is why some clients avoid it: OpenClaw, for example, turns off the developer role for any host that isn't api.openai.com.
Responses has a top-level instructions field. You can also send system-style messages inside input, but instructions is the idiomatic place.
Messages has a top-level system field, either a string or an array of text blocks. There is no system role in the message list; Anthropic's docs say so explicitly. The roles in messages are user and assistant, nothing else.
This is the first thing a translation layer has to get right. A Chat Completions request with two system messages in different positions has no exact Messages equivalent; the translator has to merge them into the top-level field and lose their positions.
Output limits: max_tokens and its relatives#
The field names alone tell you the history:
| Format | Output limit field |
|---|---|
| Chat Completions | max_completion_tokens; the older max_tokens is deprecated by OpenAI but still widely sent and accepted by compatible servers |
| Responses | max_output_tokens |
| Messages | max_tokens |
In the OpenAI formats the limit is optional. In Anthropic's format, send max_tokens on every request: Anthropic's API has historically required it, every example in its docs includes it, and Anthropic-compatible servers commonly expect it.
On Tokens the limit has a second job. The gateway uses the requested output limit to reserve cost before forwarding the request, and on a low balance it may lower the limit to what your balance covers (never below 16 tokens). The gateway knows which field to adjust for each format.
Tool calling shapes#
This is where the formats differ most, and where most compatibility bugs live.
Defining a tool. Chat Completions wraps the definition: {"type": "function", "function": {"name", "description", "parameters"}}. Responses flattens it: {"type": "function", "name", "description", "parameters"}. Messages uses {"name", "description", "input_schema"}. All three use JSON Schema for the arguments; only the key names change.
The model calling a tool.
{
"role": "assistant",
"tool_calls": [
{
"id": "call_1",
"type": "function",
"function": { "name": "read_file", "arguments": "{\"path\": \"src/app.ts\"}" }
}
]
}{
"type": "function_call",
"call_id": "call_1",
"name": "read_file",
"arguments": "{\"path\": \"src/app.ts\"}"
}{
"type": "tool_use",
"id": "toolu_1",
"name": "read_file",
"input": { "path": "src/app.ts" }
}Note arguments versus input. In both OpenAI formats the arguments are a JSON string that you parse yourself, and OpenAI's reference warns that the model doesn't always produce valid JSON. In Messages, input is already an object. Code that passes arguments around without checking which one it has will break the moment it changes format.
Sending the result back. Chat Completions uses a message with role tool and a tool_call_id. Responses uses an input item {"type": "function_call_output", "call_id", "output"}. Messages puts a tool_result block, with tool_use_id and content, inside the next user message. Tool results are user turns in Anthropic's model of the conversation.
The stop signal differs too: finish_reason: "tool_calls" in Chat Completions, a function_call item in the Responses output, stop_reason: "tool_use" in Messages. The tool calling docs have runnable examples against Tokens.
Streaming events#
All three stream over Server-Sent Events, with different event vocabularies.
Chat Completions sends unnamed data: lines, each a chunk with choices[0].delta. Text arrives in delta.content, tool calls in delta.tool_calls fragments (the arguments string arrives in pieces), and the stream ends with data: [DONE]. Usage only appears, as a final chunk with an empty choices array, if you ask for it with stream_options: {"include_usage": true}.
Responses sends typed events: response.created, response.output_text.delta for text, response.function_call_arguments.delta for tool arguments, and response.completed carrying the final response, usage included.
Messages sends named events in a fixed structure: message_start (with input usage), then for each content block a content_block_start, one or more content_block_delta (text_delta, input_json_delta for tool arguments, thinking_delta for reasoning), and a content_block_stop. Then message_delta with the stop_reason and output usage, and message_stop. ping events can appear anywhere. Anthropic's docs note that the usage counts in message_delta are cumulative.
A client that buffers and only parses at the end will work with all three. A client that renders as it goes needs a separate parser for each. The streaming docs show both client sides against Tokens.
Usage fields#
| Format | Input | Output | Cache |
|---|---|---|---|
| Chat Completions | prompt_tokens | completion_tokens | prompt_tokens_details.cached_tokens, included in prompt_tokens |
| Responses | input_tokens | output_tokens | input_tokens_details.cached_tokens |
| Messages | input_tokens | output_tokens | cache_read_input_tokens and cache_creation_input_tokens, reported separately |
The trap is in the last column. In OpenAI's formats cached tokens are a subset of the input count. In Anthropic's, input_tokens counts only the uncached part, and cache reads and writes are added on top. Sum the Messages fields naively as if they were OpenAI's and you'll undercount input on cache-heavy workloads. Prompt caching explained goes into why the cache numbers matter for cost.
Why a gateway can serve both formats for the same model#
A model doesn't speak a wire format; the server in front of it does. Plenty of model makers now ship more than one. DeepSeek's API supports tool calls, the Responses API and the Anthropic format for its current models. Moonshot's Kimi API supports both OpenAI and Anthropic formats. When the source behind a model supports the format you called, a gateway can pass the request straight through.
When it doesn't, the gateway translates. On Tokens, for upstream sources that only speak Chat Completions, a /v1/messages request is converted into a chat completion, and the reply, including the SSE stream, is converted back into Messages blocks and events. That's why Claude Code can run against a model whose provider has never heard of Anthropic's format.
Translation covers the common core: system prompt, text turns, function tools, streaming text and tool arguments, usage. It's lossy at the edges, because some fields have no counterpart on the other side. Anthropic's cache_control markers, thinking configuration and beta tool types have nowhere to go in a Chat Completions request. That's the root of most "works in curl, fails in the agent" reports with non-Claude models in Claude Code; running Claude Code with other models covers the specific fields.
The Responses API on Tokens is passed through, not translated, so it works where the upstream implements it. Billing is the same for all three: the gateway reads token counts from whichever format the response uses.
Which agents use which format#
From each agent's own documentation, checked October 2026:
| Agent | Format it uses with a custom provider |
|---|---|
| Claude Code | Anthropic Messages only |
| Codex CLI | Responses only (wire_api = "responses") |
| OpenCode | Chat Completions via @ai-sdk/openai-compatible; Responses via @ai-sdk/openai |
| Crush | Chat Completions (openai-compat) or Anthropic (anthropic) |
| OpenClaw | openai-completions by default; also openai-responses and anthropic-messages |
| Hermes Agent | chat_completions; also anthropic_messages and codex_responses transports |
| Cline, Roo Code | Chat Completions ("OpenAI Compatible") |
| Kilo Code | Any of the three, by provider package |
| Continue | Chat Completions (provider: openai) |
| Zed | Chat Completions by default, Responses with chat_completions: false, or an Anthropic-compatible provider |
| Factory Droid | Any of the three (generic-chat-completion-api, openai, anthropic) |
| VS Code Copilot (Custom Endpoint) | Any of the three |
| Gemini CLI | Neither; it only accepts Gemini-format endpoints |
Which one to use in your own code#
My default for new code is Chat Completions. Every compatible server and every framework supports it, and with a gateway in front of many models it has the widest coverage. Use Messages if you're already on the Anthropic SDK or you want content blocks and Anthropic's cache markers. Use Responses if you have code built on it or a tool that requires it, and don't rely on server-side conversation state (previous_response_id, store) through a gateway: a retry or failover can land on a source that never saw the earlier response. Send the full history instead.
Whichever you pick, the base URLs on Tokens are https://tokens.bd/v1 for the OpenAI formats and https://tokens.bd for Anthropic clients. Reference pages: chat completions, responses, messages.
Third-party API and agent details checked on 2026-10-03.
Sources: https://developers.openai.com/api/reference/resources/chat · https://developers.openai.com/api/reference/resources/responses · https://platform.claude.com/docs/en/api/messages · https://platform.claude.com/docs/en/build-with-claude/streaming · https://api-docs.deepseek.com/quick_start/pricing · https://platform.kimi.ai/docs/pricing/chat · https://code.claude.com/docs/en/llm-gateway-protocol · https://learn.chatgpt.com/docs/config-file/config-advanced · https://opencode.ai/docs/providers/ · https://docs.openclaw.ai/concepts/model-providers/custom-providers · https://hermes-agent.nousresearch.com/docs/integrations/providers · https://zed.dev/docs/ai/use-api-access · https://docs.factory.com/cli/byok/overview · https://code.visualstudio.com/docs/agent-customization/language-models · https://github.com/google-gemini/gemini-cli/blob/main/docs/reference/configuration.md