Skip to content

Token Counting

POST /v1/messages/count_tokens returns the input size of a request without running the model. How exact it is, what it costs, and how to count or estimate tokens for chat, completions, responses and embeddings.

On this page

Token counts decide what a request costs, whether it fits a model's context window and how much of your plan it uses. Tokens gives you two ways to know them: a counting endpoint that tells you the input size before you send, and a usage object in every response that tells you what was billed. This page covers both, and how to estimate when you have neither.

Count tokens with POST /v1/messages/count_tokens#

POST https://tokens.bd/v1/messages/count_tokens takes the same body as /v1/messages and returns the number of input tokens. It never runs the model and generates no text. It speaks the Anthropic format, so Anthropic SDKs use the bare host https://tokens.bd as their base URL.

curl -i https://tokens.bd/v1/messages/count_tokens \
  -H "x-api-key: $TOKENS_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "system": "You are a concise senior engineer.",
    "messages": [
      {"role": "user", "content": "Explain idempotency keys in two sentences."}
    ]
  }'

The response is one field:

json
{ "input_tokens": 31 }

The number is illustrative. model is required, as on every Tokens endpoint, and max_tokens is not needed. You can send system, messages with text and images, and tools; the count includes all of them. Authentication is the same as on /v1/messages: x-api-key or Authorization: Bearer. As elsewhere, send content-type: application/json.

Exact count or estimate#

The model decides how tokens are counted, so Tokens answers in one of two ways:

What happensHow you can tell
A provider that serves the model speaks the Anthropic Messages protocol, so the request goes to it and you get its own count.No special header.
No provider for the model speaks that protocol, or the one that does is down or answers 404, 405, 429 or a 5xx error. Tokens counts locally.Response header x-tokens-estimated: true.

A model served only through an OpenAI-format provider has no native counting, so it always gets the estimate. Check the header (curl -i prints it) before you trust a number to the token.

The local estimate is rough by design. It counts the text in system, messages and tools at about four characters per token, adds a flat 1,600 tokens for each image or document block and a few tokens for each message, and returns at least 1. Those constants can change. It is fine for "will this fit" and "about how big is this", and it is not an exact count. Code, JSON and non-English text such as Bengali are the cases where four characters per token is least accurate, because they usually need more tokens per character than English prose. Measure your own data (see below) before you size a budget.

Counting follows Anthropic's own endpoint when a provider forwards it. Based on Anthropic's token counting documentation, checked October 2026: the count is itself an estimate that can differ from the billed number by a small amount, it counts system prompts, tools, images and PDFs, and server tools, the MCP connector and url or file image and document sources are rejected (send images and PDFs as base64). Counts also depend on the model's tokenizer, so count against the model id you will send.

What counting costs and what it counts toward#

  • Not billed. A count never settles a charge, uses no plan credits and writes no usage record.
  • Still checked like a request. Counting goes through the same admission step as inference, so it needs a valid key, a model your key and plan can call, and an account with a plan or wallet balance. It counts toward your per-minute request limit and holds a concurrency slot while it runs. A used-up usage window (429 window_exhausted) or a key restricted to other models blocks it too. Unlike GET /v1/models and GET /v1/tokens/usage, it is not free of rate limits. See rate limits.
  • Same size limit. The body can be up to 10 MB, images included.
  • Needs the Messages endpoint. If the gateway has turned off the Anthropic Messages endpoint, counting returns 404 anthropic_protocol_disabled and you should use the methods below.

Errors on this endpoint use Anthropic's error shape, with the Tokens code inside, as described in Messages. Codes are in errors.

Count tokens for other endpoints#

There is no counting endpoint for chat completions, completions, responses or embeddings. Use these instead.

Read the usage that comes back#

Every successful inference response tells you what the provider counted, and that is what Tokens bills:

EndpointWhere the numbers are
/v1/chat/completionsusage.prompt_tokens, usage.completion_tokens, usage.total_tokens; cached input in usage.prompt_tokens_details.cached_tokens when reported
/v1/completionsusage.prompt_tokens, usage.completion_tokens, usage.total_tokens
/v1/embeddingsusage.prompt_tokens, usage.total_tokens
/v1/responsesusage.input_tokens, usage.output_tokens
/v1/messagesusage.input_tokens, usage.output_tokens, plus cache fields when the provider reports them

For streams, chat completions only include usage when you set stream_options: {"include_usage": true}; Anthropic-format streams carry it in the message_start and message_delta events. Details are in streaming. Per-request costs appear in the usage analytics, and remaining plan windows and wallet balance are in GET /v1/tokens/usage (models and usage).

Measure with a one-token request#

To learn the exact prompt size for a chat model before running a long generation, send the real prompt with max_tokens set to 1 and read usage.prompt_tokens:

bash
curl https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "messages": [{"role": "user", "content": "Explain idempotency keys in two sentences."}],
    "max_tokens": 1
  }'

This is billed: you pay for the input tokens and at most one output token. It also counts as a request. Some models reject or ignore very small max_tokens values, so raise it slightly if you get a 400. For a cheaper check on the Messages format, use the counting endpoint above.

Estimate with a rule of thumb#

For a quick budget, divide the number of characters by four for English prose. The gateway does the same for its own pre-flight check: before forwarding, it reserves the worst-case cost using the request body size divided by four for input, and max_tokens (8,192 if you leave it out) for output. That reservation is a safety margin and not a bill: you are charged for the usage the provider reports. It does mean a large max_tokens can make a request fail on funds or a spend cap even when the real answer would be short, so set max_tokens close to what you need. See chat completions.

A local tokenizer library gives exact counts only for the family of models it was built for. For other models it is another estimate, so use it for sizing, and use the usage object for truth.

Check that a prompt fits before sending#

A request fits when the prompt tokens plus max_tokens stay inside the model's context window, which is on the model's page in the catalog. This Python helper counts first and then sets max_tokens from the space left:

python
import os
import anthropic

client = anthropic.Anthropic(
    base_url="https://tokens.bd",
    api_key=os.environ["TOKENS_API_KEY"],
)

CONTEXT_WINDOW = 128_000  # read this from the model's catalog page
messages = [{"role": "user", "content": open("big-file.txt").read()}]

used = client.messages.count_tokens(model="deepseek/deepseek-v4.1-flash", messages=messages).input_tokens
room = CONTEXT_WINDOW - used
if room < 1_000:
    raise SystemExit(f"Prompt is {used} tokens; too close to the context window.")

reply = client.messages.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=min(4_000, room),
    messages=messages,
)
print(reply.usage)

Replace CONTEXT_WINDOW with the real value for your model. When the count is an estimate, leave extra room, for example 10 percent.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.