Tokens serves the Anthropic Messages API at POST https://tokens.bd/v1/messages. This is the endpoint Claude Code and the Anthropic SDKs call, so you can point them at Tokens by changing the base URL and key. You can request any model in your catalog through it, not only Claude models.
Base URL and headers#
Anthropic SDKs and Claude Code add /v1/messages themselves, so their base URL is the bare host: https://tokens.bd.
| Header | Value | Required |
|---|---|---|
x-api-key or Authorization | tok_live_your_key or Bearer tok_live_your_key | Yes |
anthropic-version | 2023-06-01 | Recommended; forwarded upstream |
content-type | application/json | Yes |
The SDKs send x-api-key and anthropic-version for you. Claude Code with ANTHROPIC_AUTH_TOKEN sends a Bearer header instead; both work.
Send a Messages API request with curl#
curl https://tokens.bd/v1/messages \
-H "x-api-key: $TOKENS_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"max_tokens": 512,
"system": "You are a concise senior engineer.",
"messages": [
{"role": "user", "content": "Explain idempotency keys in two sentences."}
]
}'max_tokens is required by the Messages API. A typical response:
{
"id": "msg_01AbCdEf",
"type": "message",
"role": "assistant",
"model": "deepseek/deepseek-v4.1-flash",
"content": [
{ "type": "text", "text": "An idempotency key is a client-chosen id sent with a request..." }
],
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": { "input_tokens": 31, "output_tokens": 58 }
}Use the Anthropic Python SDK#
import os
import anthropic
client = anthropic.Anthropic(
base_url="https://tokens.bd",
api_key=os.environ["TOKENS_API_KEY"],
)
message = client.messages.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=512,
messages=[{"role": "user", "content": "Explain idempotency keys in two sentences."}],
)
print(message.content[0].text)
print(message.usage)The TypeScript SDK works the same way: new Anthropic({ baseURL: "https://tokens.bd", apiKey: process.env.TOKENS_API_KEY }).
For Claude Code, the settings live in ~/.claude/settings.json; see the setup steps in the quickstart or the snippet at Connect your agent.
Non-Claude models through /v1/messages#
When the model you name is served by an OpenAI-compatible upstream, the gateway translates your Messages request into a chat completions request and translates the answer back. system, messages, max_tokens, stop_sequences, temperature, top_p, tools and tool_choice are carried across. Anthropic-only content (prompt caching markers, extended thinking blocks, documents) may be dropped or rejected in translation, so test those features against the specific model before depending on them.
Streaming events#
Set "stream": true and the response is Server-Sent Events in Anthropic's format:
event: message_start
data: {"type":"message_start","message":{"id":"msg_01AbCdEf","type":"message","role":"assistant","content":[],"model":"deepseek/deepseek-v4.1-flash","stop_reason":null,"usage":{"input_tokens":31,"output_tokens":0}}}
event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"An idempotency"}}
event: content_block_stop
data: {"type":"content_block_stop","index":0}
event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"output_tokens":58}}
event: message_stop
data: {"type":"message_stop"}Tool calls stream as a content_block_start with "type": "tool_use" followed by input_json_delta deltas. With the Python SDK, client.messages.stream(...) handles the event parsing. More on disconnects and timeouts in streaming.
Error shapes on /v1/messages#
There are two shapes, depending on where the error came from.
Errors raised by the gateway itself (bad key, no balance, rate limits) use the same OpenAI-style body as every other endpoint:
{
"error": {
"message": "Invalid API key.",
"type": "authentication_error",
"code": "invalid_api_key",
"param": null,
"request_id": "8f0c7a4e-2b1d-4c55-9a51-3f7e2d9b6c10"
}
}Errors returned by the upstream provider are reshaped into Anthropic's format:
{
"type": "error",
"error": {
"type": "invalid_request_error",
"message": "The request was rejected by the upstream provider."
}
}Upstream error messages are replaced with a generic one so internal provider details don't leak. The Anthropic SDKs raise exceptions by HTTP status (AuthenticationError, RateLimitError and so on), so both shapes surface as the right exception class. For the specific code, read x-tokens-request-id from the response headers and check errors.
Headers that are not forwarded#
The gateway forwards a short allow-list of request headers upstream. anthropic-version is on it. anthropic-beta is not, so beta features that depend on that header (for example, newer tool types or extended context flags) won't be enabled. If a feature only works with a beta header, assume it is unavailable through Tokens.
POST /v1/messages/count_tokens is also routed so that clients which call it don't break, but its behavior depends on the upstream serving the model. Don't build budgeting logic on it; use the usage object in responses instead.