# Reasoning and thinking models

> Use reasoning models through Tokens: reasoning_effort on Chat Completions, thinking and effort on Messages, how reasoning text is returned and streamed, how reasoning tokens are counted and billed, and what the gateway rewrites or drops.

Reasoning models (OpenAI calls them reasoning models, Anthropic calls the feature thinking) work through a problem before they write the answer. That extra work improves results on hard coding, math and planning tasks. It also costs tokens and time: the model produces thinking tokens you may never see, and you are billed for them.

This page covers what you can send, what comes back, how Tokens counts and bills the thinking tokens, and where the gateway changes your request. The parameter names and the meaning of each value belong to the provider that serves the model. Tokens forwards them and does not interpret them, so read the sections marked as provider behaviour as a summary of the maker's documentation (checked October 2026), not as a promise from Tokens.

## The short version

| You call                    | Control reasoning with                                             | Reasoning text comes back as                                   |
| --------------------------- | ------------------------------------------------------------------ | -------------------------------------------------------------- |
| `/v1/chat/completions`      | `reasoning_effort`                                                 | `reasoning_content` on the message or the stream delta, when the model exposes it |
| `/v1/messages`              | `thinking` and `output_config.effort` (older Claude models: `thinking.budget_tokens`) | `thinking` content blocks and `thinking_delta` stream events |
| `/v1/responses`             | `reasoning.effort` and `reasoning.summary`                         | A `reasoning` output item with a `summary` list                |

Which of these a given model honours depends on the model. Some models always reason and have no switch, some take an effort level, some take a token budget, and some ignore every one of these fields.

## Find out whether a model reasons

The model catalog at [/models](/models) lists context window, prices and a description for each model. It has no flag for "reasoning" or "thinking". To find out:

- Read the model's page and the maker's documentation. [Choosing a model](/docs/choosing-a-model) notes a few models that always think, such as Kimi K3 and GLM-5.3, and says that always-on thinking adds output tokens to every call.
- Send a test request and look at the answer. A reasoning model usually returns a `reasoning_content` field, a `thinking` block, or reasoning token details inside `usage`. A much larger `completion_tokens` than the visible answer also shows it.

```python
import os
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

resp = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=2000,
    messages=[{"role": "user", "content": "A train leaves at 09:10 and arrives at 11:45. How long is the trip?"}],
)

message = resp.choices[0].message
print("visible answer:", message.content)
print("reasoning text:", getattr(message, "reasoning_content", None))
print("usage:", resp.usage)
```

`reasoning_content` is not part of OpenAI's own API. It is a convention some OpenAI-compatible providers use for the model's reasoning text, and the gateway passes it through when the provider sends it.

## Chat Completions: reasoning_effort

`reasoning_effort` is OpenAI's parameter for how much a reasoning model thinks. Per OpenAI's reasoning guide the accepted values depend on the model and include `none`, `minimal`, `low`, `medium`, `high`, `xhigh` and `max`; most current OpenAI models default to `medium` when you omit it. Some models reject some values: OpenAI documents that GPT-6 Astra rejects `none` with a 400, and that GPT-6.1 Sol rejects both `none` and `minimal`. Lower effort gives faster and cheaper answers. Higher effort suits hard problems.

:::code-tabs

```bash title="cURL"
curl https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "max_completion_tokens": 4000,
    "reasoning_effort": "low",
    "messages": [
      {"role": "user", "content": "Find the bug: for i in range(len(xs)+1): total += xs[i]"}
    ]
  }'
```

```python title="Python"
import os
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

resp = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    max_completion_tokens=4000,
    reasoning_effort="low",
    messages=[
        {"role": "user", "content": "Find the bug: for i in range(len(xs)+1): total += xs[i]"}
    ],
)
print(resp.choices[0].message.content)
print(resp.usage)
```

```typescript title="Node.js"
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY });

const resp = await client.chat.completions.create({
  model: "deepseek/deepseek-v4.1-flash",
  max_completion_tokens: 4000,
  reasoning_effort: "low",
  messages: [{ role: "user", content: "Find the bug: for i in range(len(xs)+1): total += xs[i]" }],
});
console.log(resp.choices[0].message.content, resp.usage);
```

:::

What happens to the field depends on the model and the provider behind it. If the model does not take `reasoning_effort`, the provider either ignores it or answers with a 400, which reaches you as `invalid_request` ([errors](/docs/errors)). Some makers use their own field names for the same idea. The gateway does not validate Chat Completions fields, so a provider-specific field in the body is forwarded as it is; with the OpenAI SDKs you can add one with `extra_body`. Whether the provider honours it is the maker's decision, so check their documentation.

Two OpenAI details from its documentation that matter in practice: reasoning models count their thinking against the output limit, and Chat Completions does not support tool calling with a `reasoning_effort` other than `none` starting with GPT-5.4. For tool use with those models, use `reasoning_effort: "none"` or the Responses API.

### What the gateway rewrites for OpenAI reasoning models

OpenAI's o-series and GPT-5 and later models refuse `max_tokens` and any `temperature` or `top_p` other than 1. Many clients send them anyway, so for a Chat Completions request to one of these models (matched by the model's name) the gateway edits the body before forwarding:

- `max_tokens` is renamed to `max_completion_tokens` (if you sent both, yours is kept).
- `temperature` and `top_p` are removed unless they are exactly 1.

This applies on `/v1/chat/completions` only, not on Messages or Responses, and not to other makers' models. Models from other makers get your parameters as sent, and some of them reject or ignore `temperature` while they reason.

If your balance cannot cover the `max_tokens` or `max_completion_tokens` you asked for, the gateway can lower it (never below 16) as described in [Chat Completions](/docs/chat-completions). On a reasoning model that cut can use up the whole allowance on thinking and leave an empty or short answer with `finish_reason: "length"`. Set a limit your balance covers, or top up in [billing](/dashboard/billing).

## Messages: thinking and effort

On `/v1/messages` Anthropic's models take a `thinking` object and, on current models, an `output_config.effort` level. The gateway forwards both unchanged to a provider that speaks the Messages API. Per Anthropic's documentation (checked October 2026, [thinking](https://platform.claude.com/docs/en/build-with-claude/thinking), [extended thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking), [effort](https://platform.claude.com/docs/en/build-with-claude/effort)):

- **Current Claude models (4.7 and later, including the 5.x line):** use `thinking: {"type": "adaptive"}` and control depth with `output_config: {"effort": "low" | "medium" | "high" | "xhigh" | "max"}`. Which levels exist depends on the model. On several 5.x models thinking is already on without any `thinking` field. The old form `thinking: {"type": "enabled", "budget_tokens": N}` returns a 400 on these models.
- **Claude 4.6:** the `enabled` form still works but is deprecated.
- **Claude 4.5 and earlier:** only the `enabled` form exists. `budget_tokens` is a target that must be at least 1,024 and less than `max_tokens`.
- **Seeing the thinking text:** on many current models the `thinking` field of each block comes back empty by default (`display` is `omitted`). Send `"display": "summarized"` inside the `thinking` object to get a summary of the reasoning. Anthropic does not return the raw chain of thought under any setting.

:::code-tabs

```python title="Python: current Claude models"
import os
import anthropic

client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"])

message = client.messages.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=8000,
    thinking={"type": "adaptive", "display": "summarized"},
    output_config={"effort": "medium"},
    messages=[{"role": "user", "content": "Plan a safe rollout for a database column rename."}],
)

for block in message.content:
    if block.type == "thinking":
        print("thinking summary:", block.thinking)
    elif block.type == "text":
        print("answer:", block.text)
print(message.usage)
```

```python title="Python: Claude 4.5 and earlier"
import os
import anthropic

client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"])

message = client.messages.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=8000,
    thinking={"type": "enabled", "budget_tokens": 4000},  # at least 1024, below max_tokens
    messages=[{"role": "user", "content": "Plan a safe rollout for a database column rename."}],
)

for block in message.content:
    if block.type == "thinking":
        print("thinking:", block.thinking)
    elif block.type == "text":
        print("answer:", block.text)
```

```bash title="cURL"
curl https://tokens.bd/v1/messages \
  -H "x-api-key: $TOKENS_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "max_tokens": 8000,
    "thinking": {"type": "adaptive", "display": "summarized"},
    "output_config": {"effort": "medium"},
    "messages": [
      {"role": "user", "content": "Plan a safe rollout for a database column rename."}
    ]
  }'
```

:::

`output_config` needs a recent Anthropic SDK. These examples show the request shape for Claude models; use the form your model's maker documents. Which form a model takes is not recorded in the Tokens catalog.

A response with thinking has `thinking` blocks before the `text` block:

```json
{
  "content": [
    { "type": "thinking", "thinking": "The rename needs a two-step deploy...", "signature": "EosnCkYICxIM..." },
    { "type": "text", "text": "Roll it out in three steps: add the new column..." }
  ],
  "stop_reason": "end_turn",
  "usage": { "input_tokens": 24, "output_tokens": 912 }
}
```

The `signature` is encrypted data that lets the model continue its own reasoning. Anthropic requires thinking blocks (and any `redacted_thinking` blocks, which hold encrypted content with no readable text) to be sent back unchanged when you return tool results in a tool-use loop. Keep the whole assistant `content` array, not just the text; the loop in [Tool calling](/docs/tool-calling) shows `{"role": "assistant", "content": resp.content}`. Anthropic also documents that manual extended thinking only allows `tool_choice` of `auto` or `none`.

## What survives translation

Each model is served by one or more providers, and the gateway prefers one that speaks the same format as your request. When a model is only available in the other format, the gateway translates (see [Messages](/docs/messages)), and reasoning is handled like this:

| Direction                                    | Reasoning controls in your request                              | Reasoning in the answer                                                                                      |
| -------------------------------------------- | --------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Messages request, OpenAI-format provider     | `thinking` and `output_config` are **not carried**              | A `reasoning_content` the provider sends is **not** turned into a `thinking` block; it is dropped            |
| Chat Completions request, Anthropic-format provider | `reasoning_effort` is **not carried**                    | `thinking` blocks become `message.reasoning_content` (and `delta.reasoning_content` when streaming). `redacted_thinking` blocks and `signature` values are dropped |

Consequences:

- On a translated path you cannot switch reasoning on or change its level with those fields. Whether the model reasons is then the model's default.
- The usage numbers still include the tokens the model spent thinking, so you are billed for reasoning you cannot see.
- A Chat Completions request cannot carry Anthropic's thinking signatures, so a tool-use loop that needs thinking blocks passed back should use `/v1/messages`.

If you need a particular reasoning setting to apply, call the endpoint in the model's native format, and confirm with a test request that the answer carries reasoning (a `thinking` block or `reasoning_content`) or that `usage` changes with the setting.

## Responses API

On `POST https://tokens.bd/v1/responses` OpenAI's parameter is `reasoning`, with `effort` and, for a readable summary, `summary`. The gateway forwards the body as it is, and support depends on the provider behind the model (see [Responses API](/docs/responses)).

```python
import os
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

resp = client.responses.create(
    model="deepseek/deepseek-v4.1-flash",
    input="Find the bug: for i in range(len(xs)+1): total += xs[i]",
    reasoning={"effort": "low", "summary": "auto"},
    max_output_tokens=4000,
)
print(resp.output_text)
print(resp.usage)
```

OpenAI documents that raw reasoning tokens are never returned. With `summary` set, the response has a `reasoning` output item whose `summary` list holds a readable summary. OpenAI also says to pass reasoning items back to the model together with function-call outputs when you continue a tool loop.

## Streaming reasoning

With `"stream": true` the reasoning arrives before the answer, and the stream stays quiet for a while if the model thinks first.

**Chat Completions.** Providers that expose reasoning text send it as `delta.reasoning_content`, followed by `delta.content` for the answer. Translated Anthropic thinking is delivered the same way. A client that ignores unknown fields shows only the answer.

:::code-tabs

```python title="Python"
import os
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

stream = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=4000,
    stream=True,
    stream_options={"include_usage": True},
    messages=[{"role": "user", "content": "Why does 0.1 + 0.2 != 0.3 in floating point?"}],
)

for chunk in stream:
    if chunk.usage:
        print("\nusage:", chunk.usage)
    if not chunk.choices:
        continue
    delta = chunk.choices[0].delta
    reasoning = getattr(delta, "reasoning_content", None)
    if reasoning:
        print(reasoning, end="", flush=True)  # reasoning text, if the model sends it
    if delta.content:
        print(delta.content, end="", flush=True)
```

```typescript title="Node.js"
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY });

const stream = await client.chat.completions.create({
  model: "deepseek/deepseek-v4.1-flash",
  max_tokens: 4000,
  stream: true,
  stream_options: { include_usage: true },
  messages: [{ role: "user", content: "Why does 0.1 + 0.2 != 0.3 in floating point?" }],
});

for await (const chunk of stream) {
  if (chunk.usage) console.log("\nusage:", chunk.usage);
  const delta = chunk.choices[0]?.delta as
    | { content?: string | null; reasoning_content?: string | null }
    | undefined;
  if (delta?.reasoning_content) process.stdout.write(delta.reasoning_content);
  if (delta?.content) process.stdout.write(delta.content);
}
```

:::

**Messages.** Thinking arrives as `content_block_delta` events with a `thinking_delta`, then one `signature_delta` just before the block closes, and the text blocks follow. If `display` is `omitted`, the `thinking_delta` events carry an empty string and only the signature arrives.

```python
import os
import anthropic

client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"])

stream = client.messages.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=8000,
    stream=True,
    thinking={"type": "adaptive", "display": "summarized"},
    messages=[{"role": "user", "content": "Why does 0.1 + 0.2 != 0.3 in floating point?"}],
)

for event in stream:
    if event.type == "content_block_delta":
        if event.delta.type == "thinking_delta":
            print(event.delta.thinking, end="", flush=True)
        elif event.delta.type == "text_delta":
            print(event.delta.text, end="", flush=True)
```

More on stream formats, disconnects and timeouts is in [streaming](/docs/streaming).

## Token counts and billing

Reasoning tokens are output tokens. Both OpenAI and Anthropic document that the tokens a model spends thinking are billed as output tokens, and they count toward the output limit (`max_completion_tokens` or `max_tokens`) even when the reasoning text is not returned to you.

- Tokens bills what the provider reports in `usage`: input, output, and cache reads and writes, each at the model's catalog price for that kind of token ([prices](/models)). Reasoning tokens are part of the output count, so they are priced at the model's output rate. There is no separate reasoning price.
- The visible answer can be a small part of the bill. A response with a 300-token answer and 1,200 thinking tokens is billed as 1,500 output tokens. Usage analytics ([usage](/docs/usage-and-alerts)) show output tokens as billed.
- Summaries and omitted thinking do not reduce the bill. Anthropic documents that you are charged for the full thinking tokens, not the summary, and that `display: "omitted"` only reduces latency.
- Some providers report the split inside `usage`, for example `completion_tokens_details.reasoning_tokens` on Chat Completions, `output_tokens_details.reasoning_tokens` on Responses and `output_tokens_details.thinking_tokens` on Messages. The fields come from the provider and are passed through; they are informational, and the output total is what is billed.
- If a provider reports no usage at all, the gateway estimates from the response size, counting `reasoning_content` text along with the answer.
- Anthropic documents that its newer models keep thinking blocks from earlier turns in context and bill them as input on the later turns. A long tool loop with thinking resends them each round, so cost grows with the number of rounds.

Before forwarding, the gateway reserves the worst case against your balance using your output limit (8,192 tokens if you set none), as explained in [Chat Completions](/docs/chat-completions). A large limit for a reasoning model, such as 32,000 or more, reserves a large amount even if the model finishes early. You are charged for actual usage.

## How much to allow

- Set the output limit high enough for thinking plus the answer. If the model uses it all on thinking, you get a cut-off or empty answer with `finish_reason: "length"` (Chat Completions) or `stop_reason: "max_tokens"` (Messages). OpenAI suggests reserving at least 25,000 tokens for reasoning and output when you start out on its reasoning models, then tuning down.
- Start at low or medium effort and raise it only if answers are not good enough. Higher effort means more thinking tokens and more latency.
- For routine work (renames, formatting, simple edits) a non-reasoning model or `reasoning_effort: "none"` where the model allows it is cheaper and faster.

## Latency and timeouts

A model that thinks before it answers can be silent for a long time before the first token. Stream the response so your client can show progress and so the connection stays active.

The gateway also watches for silence. When several providers serve a model, it waits for a streaming provider's first event for 30 seconds by default (the operator can change this), then tries the next provider if one is configured; a non-streaming request gets at least 120 seconds. The last provider in the chain has no early cut-off. A provider that sends nothing for 600 seconds ends in `504 upstream_timeout` ([errors](/docs/errors)). If a slow reasoning model times out on non-streaming calls, switch to streaming and set your own client timeout above the longest answer you expect.

## Troubleshooting

| Symptom                                                  | Likely cause                                                                | Fix                                                                                          |
| -------------------------------------------------------- | --------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------- |
| 400 `invalid_request` after adding `reasoning_effort`    | The model or provider does not accept that field or value                   | Remove it, or use a value the maker documents for that model.                                |
| 400 on `thinking` with `budget_tokens`                   | The Claude model only supports adaptive thinking                            | Use `thinking: {"type": "adaptive"}` and `output_config.effort`.                             |
| 400 on `max_tokens` or `temperature`                     | An OpenAI reasoning model on a path where the gateway does not rewrite them | On Chat Completions the gateway rewrites them; on Messages or Responses use the model's own parameter names. |
| Empty answer, `finish_reason: "length"`                  | Thinking used the whole output limit                                        | Raise `max_tokens` or `max_completion_tokens`, or lower the effort.                          |
| No reasoning text in the response                        | The model does not expose it, `display` is `omitted`, or the path dropped it | Set `display: "summarized"` on Claude models, or see the translation table above.            |
| `thinking` block missing on Messages for a non-Claude model | The model is served in OpenAI format and `reasoning_content` is dropped  | Read the reasoning on Chat Completions instead.                                              |
| Bill larger than the visible answer                      | Thinking tokens are output tokens                                           | Lower the effort or budget, or use a smaller model.                                          |
| 504 `upstream_timeout` on a long non-streaming call      | The model thought for longer than the upstream wait                         | Stream the response.                                                                         |
| Tool loop fails with a thinking-block error on Claude    | Thinking blocks were not passed back unchanged                              | Resend the full assistant `content`, including `thinking` and `redacted_thinking` blocks, on `/v1/messages`. |

---
Page: https://tokens.bd/docs/reasoning
