Reasoning models (OpenAI calls them reasoning models, Anthropic calls the feature thinking) work through a problem before they write the answer. That extra work improves results on hard coding, math and planning tasks. It also costs tokens and time: the model produces thinking tokens you may never see, and you are billed for them.
This page covers what you can send, what comes back, how Tokens counts and bills the thinking tokens, and where the gateway changes your request. The parameter names and the meaning of each value belong to the provider that serves the model. Tokens forwards them and does not interpret them, so read the sections marked as provider behaviour as a summary of the maker's documentation (checked October 2026), not as a promise from Tokens.
The short version#
| You call | Control reasoning with | Reasoning text comes back as |
|---|---|---|
/v1/chat/completions | reasoning_effort | reasoning_content on the message or the stream delta, when the model exposes it |
/v1/messages | thinking and output_config.effort (older Claude models: thinking.budget_tokens) | thinking content blocks and thinking_delta stream events |
/v1/responses | reasoning.effort and reasoning.summary | A reasoning output item with a summary list |
Which of these a given model honours depends on the model. Some models always reason and have no switch, some take an effort level, some take a token budget, and some ignore every one of these fields.
Find out whether a model reasons#
The model catalog at /models lists context window, prices and a description for each model. It has no flag for "reasoning" or "thinking". To find out:
- Read the model's page and the maker's documentation. Choosing a model notes a few models that always think, such as Kimi K3 and GLM-5.3, and says that always-on thinking adds output tokens to every call.
- Send a test request and look at the answer. A reasoning model usually returns a
reasoning_contentfield, athinkingblock, or reasoning token details insideusage. A much largercompletion_tokensthan the visible answer also shows it.
import os
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
resp = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=2000,
messages=[{"role": "user", "content": "A train leaves at 09:10 and arrives at 11:45. How long is the trip?"}],
)
message = resp.choices[0].message
print("visible answer:", message.content)
print("reasoning text:", getattr(message, "reasoning_content", None))
print("usage:", resp.usage)reasoning_content is not part of OpenAI's own API. It is a convention some OpenAI-compatible providers use for the model's reasoning text, and the gateway passes it through when the provider sends it.
Chat Completions: reasoning_effort#
reasoning_effort is OpenAI's parameter for how much a reasoning model thinks. Per OpenAI's reasoning guide the accepted values depend on the model and include none, minimal, low, medium, high, xhigh and max; most current OpenAI models default to medium when you omit it. Some models reject some values: OpenAI documents that GPT-6 Astra rejects none with a 400, and that GPT-6.1 Sol rejects both none and minimal. Lower effort gives faster and cheaper answers. Higher effort suits hard problems.
curl https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"max_completion_tokens": 4000,
"reasoning_effort": "low",
"messages": [
{"role": "user", "content": "Find the bug: for i in range(len(xs)+1): total += xs[i]"}
]
}'import os
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
resp = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
max_completion_tokens=4000,
reasoning_effort="low",
messages=[
{"role": "user", "content": "Find the bug: for i in range(len(xs)+1): total += xs[i]"}
],
)
print(resp.choices[0].message.content)
print(resp.usage)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY });
const resp = await client.chat.completions.create({
model: "deepseek/deepseek-v4.1-flash",
max_completion_tokens: 4000,
reasoning_effort: "low",
messages: [{ role: "user", content: "Find the bug: for i in range(len(xs)+1): total += xs[i]" }],
});
console.log(resp.choices[0].message.content, resp.usage);What happens to the field depends on the model and the provider behind it. If the model does not take reasoning_effort, the provider either ignores it or answers with a 400, which reaches you as invalid_request (errors). Some makers use their own field names for the same idea. The gateway does not validate Chat Completions fields, so a provider-specific field in the body is forwarded as it is; with the OpenAI SDKs you can add one with extra_body. Whether the provider honours it is the maker's decision, so check their documentation.
Two OpenAI details from its documentation that matter in practice: reasoning models count their thinking against the output limit, and Chat Completions does not support tool calling with a reasoning_effort other than none starting with GPT-5.4. For tool use with those models, use reasoning_effort: "none" or the Responses API.
What the gateway rewrites for OpenAI reasoning models#
OpenAI's o-series and GPT-5 and later models refuse max_tokens and any temperature or top_p other than 1. Many clients send them anyway, so for a Chat Completions request to one of these models (matched by the model's name) the gateway edits the body before forwarding:
max_tokensis renamed tomax_completion_tokens(if you sent both, yours is kept).temperatureandtop_pare removed unless they are exactly 1.
This applies on /v1/chat/completions only, not on Messages or Responses, and not to other makers' models. Models from other makers get your parameters as sent, and some of them reject or ignore temperature while they reason.
If your balance cannot cover the max_tokens or max_completion_tokens you asked for, the gateway can lower it (never below 16) as described in Chat Completions. On a reasoning model that cut can use up the whole allowance on thinking and leave an empty or short answer with finish_reason: "length". Set a limit your balance covers, or top up in billing.
Messages: thinking and effort#
On /v1/messages Anthropic's models take a thinking object and, on current models, an output_config.effort level. The gateway forwards both unchanged to a provider that speaks the Messages API. Per Anthropic's documentation (checked October 2026, thinking, extended thinking, effort):
- Current Claude models (4.7 and later, including the 5.x line): use
thinking: {"type": "adaptive"}and control depth withoutput_config: {"effort": "low" | "medium" | "high" | "xhigh" | "max"}. Which levels exist depends on the model. On several 5.x models thinking is already on without anythinkingfield. The old formthinking: {"type": "enabled", "budget_tokens": N}returns a 400 on these models. - Claude 4.6: the
enabledform still works but is deprecated. - Claude 4.5 and earlier: only the
enabledform exists.budget_tokensis a target that must be at least 1,024 and less thanmax_tokens. - Seeing the thinking text: on many current models the
thinkingfield of each block comes back empty by default (displayisomitted). Send"display": "summarized"inside thethinkingobject to get a summary of the reasoning. Anthropic does not return the raw chain of thought under any setting.
import os
import anthropic
client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"])
message = client.messages.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=8000,
thinking={"type": "adaptive", "display": "summarized"},
output_config={"effort": "medium"},
messages=[{"role": "user", "content": "Plan a safe rollout for a database column rename."}],
)
for block in message.content:
if block.type == "thinking":
print("thinking summary:", block.thinking)
elif block.type == "text":
print("answer:", block.text)
print(message.usage)import os
import anthropic
client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"])
message = client.messages.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=8000,
thinking={"type": "enabled", "budget_tokens": 4000}, # at least 1024, below max_tokens
messages=[{"role": "user", "content": "Plan a safe rollout for a database column rename."}],
)
for block in message.content:
if block.type == "thinking":
print("thinking:", block.thinking)
elif block.type == "text":
print("answer:", block.text)curl https://tokens.bd/v1/messages \
-H "x-api-key: $TOKENS_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"max_tokens": 8000,
"thinking": {"type": "adaptive", "display": "summarized"},
"output_config": {"effort": "medium"},
"messages": [
{"role": "user", "content": "Plan a safe rollout for a database column rename."}
]
}'output_config needs a recent Anthropic SDK. These examples show the request shape for Claude models; use the form your model's maker documents. Which form a model takes is not recorded in the Tokens catalog.
A response with thinking has thinking blocks before the text block:
{
"content": [
{ "type": "thinking", "thinking": "The rename needs a two-step deploy...", "signature": "EosnCkYICxIM..." },
{ "type": "text", "text": "Roll it out in three steps: add the new column..." }
],
"stop_reason": "end_turn",
"usage": { "input_tokens": 24, "output_tokens": 912 }
}The signature is encrypted data that lets the model continue its own reasoning. Anthropic requires thinking blocks (and any redacted_thinking blocks, which hold encrypted content with no readable text) to be sent back unchanged when you return tool results in a tool-use loop. Keep the whole assistant content array, not just the text; the loop in Tool calling shows {"role": "assistant", "content": resp.content}. Anthropic also documents that manual extended thinking only allows tool_choice of auto or none.
What survives translation#
Each model is served by one or more providers, and the gateway prefers one that speaks the same format as your request. When a model is only available in the other format, the gateway translates (see Messages), and reasoning is handled like this:
| Direction | Reasoning controls in your request | Reasoning in the answer |
|---|---|---|
| Messages request, OpenAI-format provider | thinking and output_config are not carried | A reasoning_content the provider sends is not turned into a thinking block; it is dropped |
| Chat Completions request, Anthropic-format provider | reasoning_effort is not carried | thinking blocks become message.reasoning_content (and delta.reasoning_content when streaming). redacted_thinking blocks and signature values are dropped |
Consequences:
- On a translated path you cannot switch reasoning on or change its level with those fields. Whether the model reasons is then the model's default.
- The usage numbers still include the tokens the model spent thinking, so you are billed for reasoning you cannot see.
- A Chat Completions request cannot carry Anthropic's thinking signatures, so a tool-use loop that needs thinking blocks passed back should use
/v1/messages.
If you need a particular reasoning setting to apply, call the endpoint in the model's native format, and confirm with a test request that the answer carries reasoning (a thinking block or reasoning_content) or that usage changes with the setting.
Responses API#
On POST https://tokens.bd/v1/responses OpenAI's parameter is reasoning, with effort and, for a readable summary, summary. The gateway forwards the body as it is, and support depends on the provider behind the model (see Responses API).
import os
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
resp = client.responses.create(
model="deepseek/deepseek-v4.1-flash",
input="Find the bug: for i in range(len(xs)+1): total += xs[i]",
reasoning={"effort": "low", "summary": "auto"},
max_output_tokens=4000,
)
print(resp.output_text)
print(resp.usage)OpenAI documents that raw reasoning tokens are never returned. With summary set, the response has a reasoning output item whose summary list holds a readable summary. OpenAI also says to pass reasoning items back to the model together with function-call outputs when you continue a tool loop.
Streaming reasoning#
With "stream": true the reasoning arrives before the answer, and the stream stays quiet for a while if the model thinks first.
Chat Completions. Providers that expose reasoning text send it as delta.reasoning_content, followed by delta.content for the answer. Translated Anthropic thinking is delivered the same way. A client that ignores unknown fields shows only the answer.
import os
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
stream = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=4000,
stream=True,
stream_options={"include_usage": True},
messages=[{"role": "user", "content": "Why does 0.1 + 0.2 != 0.3 in floating point?"}],
)
for chunk in stream:
if chunk.usage:
print("\nusage:", chunk.usage)
if not chunk.choices:
continue
delta = chunk.choices[0].delta
reasoning = getattr(delta, "reasoning_content", None)
if reasoning:
print(reasoning, end="", flush=True) # reasoning text, if the model sends it
if delta.content:
print(delta.content, end="", flush=True)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY });
const stream = await client.chat.completions.create({
model: "deepseek/deepseek-v4.1-flash",
max_tokens: 4000,
stream: true,
stream_options: { include_usage: true },
messages: [{ role: "user", content: "Why does 0.1 + 0.2 != 0.3 in floating point?" }],
});
for await (const chunk of stream) {
if (chunk.usage) console.log("\nusage:", chunk.usage);
const delta = chunk.choices[0]?.delta as
| { content?: string | null; reasoning_content?: string | null }
| undefined;
if (delta?.reasoning_content) process.stdout.write(delta.reasoning_content);
if (delta?.content) process.stdout.write(delta.content);
}Messages. Thinking arrives as content_block_delta events with a thinking_delta, then one signature_delta just before the block closes, and the text blocks follow. If display is omitted, the thinking_delta events carry an empty string and only the signature arrives.
import os
import anthropic
client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"])
stream = client.messages.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=8000,
stream=True,
thinking={"type": "adaptive", "display": "summarized"},
messages=[{"role": "user", "content": "Why does 0.1 + 0.2 != 0.3 in floating point?"}],
)
for event in stream:
if event.type == "content_block_delta":
if event.delta.type == "thinking_delta":
print(event.delta.thinking, end="", flush=True)
elif event.delta.type == "text_delta":
print(event.delta.text, end="", flush=True)More on stream formats, disconnects and timeouts is in streaming.
Token counts and billing#
Reasoning tokens are output tokens. Both OpenAI and Anthropic document that the tokens a model spends thinking are billed as output tokens, and they count toward the output limit (max_completion_tokens or max_tokens) even when the reasoning text is not returned to you.
- Tokens bills what the provider reports in
usage: input, output, and cache reads and writes, each at the model's catalog price for that kind of token (prices). Reasoning tokens are part of the output count, so they are priced at the model's output rate. There is no separate reasoning price. - The visible answer can be a small part of the bill. A response with a 300-token answer and 1,200 thinking tokens is billed as 1,500 output tokens. Usage analytics (usage) show output tokens as billed.
- Summaries and omitted thinking do not reduce the bill. Anthropic documents that you are charged for the full thinking tokens, not the summary, and that
display: "omitted"only reduces latency. - Some providers report the split inside
usage, for examplecompletion_tokens_details.reasoning_tokenson Chat Completions,output_tokens_details.reasoning_tokenson Responses andoutput_tokens_details.thinking_tokenson Messages. The fields come from the provider and are passed through; they are informational, and the output total is what is billed. - If a provider reports no usage at all, the gateway estimates from the response size, counting
reasoning_contenttext along with the answer. - Anthropic documents that its newer models keep thinking blocks from earlier turns in context and bill them as input on the later turns. A long tool loop with thinking resends them each round, so cost grows with the number of rounds.
Before forwarding, the gateway reserves the worst case against your balance using your output limit (8,192 tokens if you set none), as explained in Chat Completions. A large limit for a reasoning model, such as 32,000 or more, reserves a large amount even if the model finishes early. You are charged for actual usage.
How much to allow#
- Set the output limit high enough for thinking plus the answer. If the model uses it all on thinking, you get a cut-off or empty answer with
finish_reason: "length"(Chat Completions) orstop_reason: "max_tokens"(Messages). OpenAI suggests reserving at least 25,000 tokens for reasoning and output when you start out on its reasoning models, then tuning down. - Start at low or medium effort and raise it only if answers are not good enough. Higher effort means more thinking tokens and more latency.
- For routine work (renames, formatting, simple edits) a non-reasoning model or
reasoning_effort: "none"where the model allows it is cheaper and faster.
Latency and timeouts#
A model that thinks before it answers can be silent for a long time before the first token. Stream the response so your client can show progress and so the connection stays active.
The gateway also watches for silence. When several providers serve a model, it waits for a streaming provider's first event for 30 seconds by default (the operator can change this), then tries the next provider if one is configured; a non-streaming request gets at least 120 seconds. The last provider in the chain has no early cut-off. A provider that sends nothing for 600 seconds ends in 504 upstream_timeout (errors). If a slow reasoning model times out on non-streaming calls, switch to streaming and set your own client timeout above the longest answer you expect.
Troubleshooting#
| Symptom | Likely cause | Fix |
|---|---|---|
400 invalid_request after adding reasoning_effort | The model or provider does not accept that field or value | Remove it, or use a value the maker documents for that model. |
400 on thinking with budget_tokens | The Claude model only supports adaptive thinking | Use thinking: {"type": "adaptive"} and output_config.effort. |
400 on max_tokens or temperature | An OpenAI reasoning model on a path where the gateway does not rewrite them | On Chat Completions the gateway rewrites them; on Messages or Responses use the model's own parameter names. |
Empty answer, finish_reason: "length" | Thinking used the whole output limit | Raise max_tokens or max_completion_tokens, or lower the effort. |
| No reasoning text in the response | The model does not expose it, display is omitted, or the path dropped it | Set display: "summarized" on Claude models, or see the translation table above. |
thinking block missing on Messages for a non-Claude model | The model is served in OpenAI format and reasoning_content is dropped | Read the reasoning on Chat Completions instead. |
| Bill larger than the visible answer | Thinking tokens are output tokens | Lower the effort or budget, or use a smaller model. |
504 upstream_timeout on a long non-streaming call | The model thought for longer than the upstream wait | Stream the response. |
| Tool loop fails with a thinking-block error on Claude | Thinking blocks were not passed back unchanged | Resend the full assistant content, including thinking and redacted_thinking blocks, on /v1/messages. |