# Token Counting

> POST /v1/messages/count_tokens returns the input size of a request without running the model. How exact it is, what it costs, and how to count or estimate tokens for chat, completions, responses and embeddings.

Token counts decide what a request costs, whether it fits a model's context window and how much of your plan it uses. Tokens gives you two ways to know them: a counting endpoint that tells you the input size before you send, and a `usage` object in every response that tells you what was billed. This page covers both, and how to estimate when you have neither.

## Count tokens with POST /v1/messages/count_tokens

`POST https://tokens.bd/v1/messages/count_tokens` takes the same body as [`/v1/messages`](/docs/messages) and returns the number of input tokens. It never runs the model and generates no text. It speaks the Anthropic format, so Anthropic SDKs use the bare host `https://tokens.bd` as their base URL.

:::code-tabs

```bash title="cURL"
curl -i https://tokens.bd/v1/messages/count_tokens \
  -H "x-api-key: $TOKENS_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "system": "You are a concise senior engineer.",
    "messages": [
      {"role": "user", "content": "Explain idempotency keys in two sentences."}
    ]
  }'
```

```python title="Python"
import os
import anthropic

client = anthropic.Anthropic(
    base_url="https://tokens.bd",
    api_key=os.environ["TOKENS_API_KEY"],
)

count = client.messages.count_tokens(
    model="deepseek/deepseek-v4.1-flash",
    system="You are a concise senior engineer.",
    messages=[{"role": "user", "content": "Explain idempotency keys in two sentences."}],
)
print(count.input_tokens)
```

```typescript title="Node.js"
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  baseURL: "https://tokens.bd",
  apiKey: process.env.TOKENS_API_KEY,
});

const count = await client.messages.countTokens({
  model: "deepseek/deepseek-v4.1-flash",
  system: "You are a concise senior engineer.",
  messages: [{ role: "user", content: "Explain idempotency keys in two sentences." }],
});
console.log(count.input_tokens);
```

:::

The response is one field:

```json
{ "input_tokens": 31 }
```

The number is illustrative. `model` is required, as on every Tokens endpoint, and `max_tokens` is not needed. You can send `system`, `messages` with text and images, and `tools`; the count includes all of them. Authentication is the same as on `/v1/messages`: `x-api-key` or `Authorization: Bearer`. As elsewhere, send `content-type: application/json`.

## Exact count or estimate

The model decides how tokens are counted, so Tokens answers in one of two ways:

| What happens                                                                                                  | How you can tell                               |
| ------------------------------------------------------------------------------------------------------------- | ---------------------------------------------- |
| A provider that serves the model speaks the Anthropic Messages protocol, so the request goes to it and you get its own count. | No special header.                             |
| No provider for the model speaks that protocol, or the one that does is down or answers 404, 405, 429 or a 5xx error. Tokens counts locally. | Response header `x-tokens-estimated: true`.    |

A model served only through an OpenAI-format provider has no native counting, so it always gets the estimate. Check the header (`curl -i` prints it) before you trust a number to the token.

The local estimate is rough by design. It counts the text in `system`, `messages` and `tools` at about four characters per token, adds a flat 1,600 tokens for each image or document block and a few tokens for each message, and returns at least 1. Those constants can change. It is fine for "will this fit" and "about how big is this", and it is not an exact count. Code, JSON and non-English text such as Bengali are the cases where four characters per token is least accurate, because they usually need more tokens per character than English prose. Measure your own data (see below) before you size a budget.

Counting follows Anthropic's own endpoint when a provider forwards it. Based on Anthropic's token counting documentation, checked October 2026: the count is itself an estimate that can differ from the billed number by a small amount, it counts system prompts, tools, images and PDFs, and server tools, the MCP connector and `url` or `file` image and document sources are rejected (send images and PDFs as base64). Counts also depend on the model's tokenizer, so count against the model id you will send.

## What counting costs and what it counts toward

- **Not billed.** A count never settles a charge, uses no plan credits and writes no usage record.
- **Still checked like a request.** Counting goes through the same admission step as inference, so it needs a valid key, a model your key and plan can call, and an account with a plan or wallet balance. It counts toward your per-minute request limit and holds a concurrency slot while it runs. A used-up usage window (429 `window_exhausted`) or a key restricted to other models blocks it too. Unlike `GET /v1/models` and `GET /v1/tokens/usage`, it is not free of rate limits. See [rate limits](/docs/rate-limits).
- **Same size limit.** The body can be up to 10 MB, images included.
- **Needs the Messages endpoint.** If the gateway has turned off the Anthropic Messages endpoint, counting returns 404 `anthropic_protocol_disabled` and you should use the methods below.

Errors on this endpoint use Anthropic's error shape, with the Tokens `code` inside, as described in [Messages](/docs/messages). Codes are in [errors](/docs/errors).

## Count tokens for other endpoints

There is no counting endpoint for chat completions, completions, responses or embeddings. Use these instead.

### Read the usage that comes back

Every successful inference response tells you what the provider counted, and that is what Tokens bills:

| Endpoint                | Where the numbers are                                                                                        |
| ----------------------- | ------------------------------------------------------------------------------------------------------------ |
| `/v1/chat/completions`  | `usage.prompt_tokens`, `usage.completion_tokens`, `usage.total_tokens`; cached input in `usage.prompt_tokens_details.cached_tokens` when reported |
| `/v1/completions`       | `usage.prompt_tokens`, `usage.completion_tokens`, `usage.total_tokens`                                       |
| `/v1/embeddings`        | `usage.prompt_tokens`, `usage.total_tokens`                                                                  |
| `/v1/responses`         | `usage.input_tokens`, `usage.output_tokens`                                                                  |
| `/v1/messages`          | `usage.input_tokens`, `usage.output_tokens`, plus cache fields when the provider reports them                |

For streams, chat completions only include `usage` when you set `stream_options: {"include_usage": true}`; Anthropic-format streams carry it in the `message_start` and `message_delta` events. Details are in [streaming](/docs/streaming). Per-request costs appear in the [usage analytics](/docs/usage-and-alerts), and remaining plan windows and wallet balance are in `GET /v1/tokens/usage` ([models and usage](/docs/models-and-usage)).

### Measure with a one-token request

To learn the exact prompt size for a chat model before running a long generation, send the real prompt with `max_tokens` set to 1 and read `usage.prompt_tokens`:

```bash
curl https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "messages": [{"role": "user", "content": "Explain idempotency keys in two sentences."}],
    "max_tokens": 1
  }'
```

This is billed: you pay for the input tokens and at most one output token. It also counts as a request. Some models reject or ignore very small `max_tokens` values, so raise it slightly if you get a 400. For a cheaper check on the Messages format, use the counting endpoint above.

### Estimate with a rule of thumb

For a quick budget, divide the number of characters by four for English prose. The gateway does the same for its own pre-flight check: before forwarding, it reserves the worst-case cost using the request body size divided by four for input, and `max_tokens` (8,192 if you leave it out) for output. That reservation is a safety margin and not a bill: you are charged for the `usage` the provider reports. It does mean a large `max_tokens` can make a request fail on funds or a spend cap even when the real answer would be short, so set `max_tokens` close to what you need. See [chat completions](/docs/chat-completions).

A local tokenizer library gives exact counts only for the family of models it was built for. For other models it is another estimate, so use it for sizing, and use the `usage` object for truth.

## Check that a prompt fits before sending

A request fits when the prompt tokens plus `max_tokens` stay inside the model's context window, which is on the model's page in [the catalog](/models). This Python helper counts first and then sets `max_tokens` from the space left:

```python
import os
import anthropic

client = anthropic.Anthropic(
    base_url="https://tokens.bd",
    api_key=os.environ["TOKENS_API_KEY"],
)

CONTEXT_WINDOW = 128_000  # read this from the model's catalog page
messages = [{"role": "user", "content": open("big-file.txt").read()}]

used = client.messages.count_tokens(model="deepseek/deepseek-v4.1-flash", messages=messages).input_tokens
room = CONTEXT_WINDOW - used
if room < 1_000:
    raise SystemExit(f"Prompt is {used} tokens; too close to the context window.")

reply = client.messages.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=min(4_000, room),
    messages=messages,
)
print(reply.usage)
```

Replace `CONTEXT_WINDOW` with the real value for your model. When the count is an estimate, leave extra room, for example 10 percent.

## Related

- [Messages](/docs/messages) for the endpoint whose body counting takes.
- [Chat completions](/docs/chat-completions), [Embeddings](/docs/embeddings) and [Legacy completions](/docs/legacy-completions) for the `usage` object on each.
- [Plans and wallet](/docs/plans-and-wallet) for how tokens turn into credits and cost.

---
Page: https://tokens.bd/docs/token-counting
