If your app calls the Anthropic Messages API, moving it to Tokens takes three changes: the base URL, the API key and the model id. Tokens serves POST /v1/messages, streaming and POST /v1/messages/count_tokens in Anthropic's format, so the Anthropic SDKs and your message-handling code keep working. What needs care is the set of Anthropic-only features (prompt caching, thinking, citations, the Files and Batches APIs), because they depend on which provider serves the model you pick.
What changes and what stays the same#
| Setting | Anthropic | Tokens | Where you set it |
|---|---|---|---|
| Base URL | https://api.anthropic.com (SDK default) | https://tokens.bd, without /v1 | base_url (Python) or baseURL (Node.js), or ANTHROPIC_BASE_URL |
| API key | An Anthropic key | tok_live_... from API keys | api_key or apiKey, or ANTHROPIC_API_KEY |
| Model id | An Anthropic id | An alias in provider/model form from /models | The model field of every request |
anthropic-workspace-id | Selects a workspace | Not used. Remove it. | Client options |
The SDKs add /v1/messages themselves, so the base URL is the bare host. If you set https://tokens.bd/v1 instead, requests go to /v1/v1/messages and fail with 404. Calling the endpoint yourself with curl is different: the full URL is https://tokens.bd/v1/messages.
The Anthropic Python and TypeScript SDKs read ANTHROPIC_API_KEY, ANTHROPIC_AUTH_TOKEN and ANTHROPIC_BASE_URL when you pass nothing in code (checked in the SDK sources, October 2026). Be careful in a shell where you also run Claude Code or other tools that read the same variables. See the Anthropic SDK page and Claude Code.
What stays the same:
- The request and response shapes of
POST /v1/messages:system,messages, content blocks,max_tokens(still required),tools,tool_choice,stop_sequencesand theusageobject. - Server-Sent Events streaming with the usual event names (
message_start,content_block_delta,message_stop).client.messages.stream(...)works. - Authentication. Tokens accepts the key in
x-api-key(what the SDKs send) or inAuthorization: Bearer, so the SDK needs no auth change. If both arrive, the Bearer header wins. - The
anthropic-versionheader. It is forwarded to the provider, and the default2023-06-01is used when you send none. - Error bodies on
/v1/messagesin Anthropic's shape, so the SDKs raise the same exception classes.
Before and after#
import os
import anthropic
-client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
+client = anthropic.Anthropic(
+ base_url="https://tokens.bd",
+ api_key=os.environ["TOKENS_API_KEY"],
+)
message = client.messages.create(
- model="your-claude-model",
+ model="deepseek/deepseek-v4.1-flash",
max_tokens=512,
messages=[{"role": "user", "content": "Explain idempotency keys in two sentences."}],
)
print(message.content[0].text) import Anthropic from "@anthropic-ai/sdk";
-const client = new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY });
+const client = new Anthropic({
+ baseURL: "https://tokens.bd",
+ apiKey: process.env.TOKENS_API_KEY,
+});
const message = await client.messages.create({
- model: "your-claude-model",
+ model: "deepseek/deepseek-v4.1-flash",
max_tokens: 512,
messages: [{ role: "user", content: "Explain idempotency keys in two sentences." }],
});
console.log(message.content[0].text);-curl https://api.anthropic.com/v1/messages \
- -H "x-api-key: $ANTHROPIC_API_KEY" \
+curl https://tokens.bd/v1/messages \
+ -H "x-api-key: $TOKENS_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
- "model": "your-claude-model",
+ "model": "deepseek/deepseek-v4.1-flash",
"max_tokens": 512,
"messages": [{"role": "user", "content": "Explain idempotency keys in two sentences."}]
}'Choose the model id#
Anthropic's ids and Tokens ids are different lists. Tokens ids are aliases in provider/model form, and they do not always match the provider's own id. The Anthropic SDK lists them for you, because GET /v1/models answers in Anthropic's format when the request carries anthropic-version:
for m in client.models.list():
print(m.id, m.display_name)With curl, send anthropic-version to get that shape. With only Authorization: Bearer you get the OpenAI list shape. In both, the list holds only models this key can call, so a key with an allowed-models list, or an account with no plan or wallet balance, sees fewer.
The Anthropic format also works for models from other makers, not only Claude models. That is a feature, but it is also a model change: answers, tool-call behaviour and token counts differ. Prices, context windows and capabilities are on /models; Choosing a model helps you pick.
Which provider serves the model decides the features#
Each model on Tokens is served by one or more providers. When a provider speaks the Messages API itself, your request goes to it unchanged except for the model id, so thinking blocks, prompt caching and anthropic-beta features come back untouched. Tokens prefers such a provider when one serves the model. When a model is only served by a provider that speaks the OpenAI format, Tokens translates your Messages request into a chat completions request and the answer back. The catalog does not say which case applies to a model, so test every feature you depend on against the exact model you choose.
What the translation carries: system (text, joined into one system message), messages with text, image, tool_use and tool_result blocks, max_tokens, stop_sequences, temperature, top_p, stream, tools and tool_choice. Everything else is not copied.
| Anthropic feature | Provider speaks Messages | Translated model |
|---|---|---|
| Streaming, client tools, images | Passed through | Supported (translated) |
Prompt caching (cache_control) | Passed through. usage carries the cache counts the provider reports. | Markers dropped. Caching, if any, is the provider's own and not under your control. |
Extended or adaptive thinking | Passed through | Parameter dropped. No thinking blocks. |
document blocks and citations | Passed through | Document blocks dropped. No citations. |
| Anthropic-defined server tools (web search, code execution, computer use) | Passed to the provider as sent. Tokens does not run them, so it is up to the provider. | Every tools entry becomes a function tool, so these do not work. |
anthropic-beta header | Forwarded | No effect |
top_k, metadata | Passed through | Dropped |
Tokens billing reads the provider's usage numbers, and cached input is priced at the model's cache-read rate where the catalog has one. Prompt caching also saves money only on models whose provider supports it; see Choosing a model and Messages.
Endpoints and features Tokens does not have#
Tokens serves these Anthropic endpoints: POST /v1/messages, POST /v1/messages/count_tokens and GET /v1/models. Every other path returns 404 unsupported_endpoint, in Anthropic's error shape (not_found_error).
| Anthropic API | On Tokens |
|---|---|
Message Batches (/v1/messages/batches) | Not supported. Send requests one by one, or run your own queue. |
Files API (/v1/files) | Not supported. Send images and documents inline as base64, or as an image URL. A file_id in a content block cannot work. |
| Skills, Managed Agents, Agents, Sessions | Not supported. |
GET /v1/models/{id} (models.retrieve) | Not supported. List the models and filter. |
Models API fields capabilities, lifecycle | Not returned. The list has type, id, display_name and created_at, and it has no pagination: has_more is false. |
POST /v1/messages/count_tokens | Supported. See below. |
Count tokens#
POST /v1/messages/count_tokens takes the body of a Messages request without max_tokens and returns {"input_tokens": N}. It never runs the model and is never billed. When a provider that serves the model counts tokens natively you get its exact count. Otherwise Tokens estimates (about four characters per token, plus a fixed amount per image and per turn) and sets the response header x-tokens-estimated: true. Use the estimate for budgeting only.
Differences that can bite#
Rate limits#
Anthropic limits requests, input tokens and output tokens per minute for each model class and reports them in anthropic-ratelimit-* headers. Tokens limits requests per minute per account (60 by default, or your plan's value) and concurrent requests per account (10 with a plan, 3 without). The rate limits page lists no token-per-minute limit. More keys do not raise either limit.
- A 429 carries
retry-afterin seconds. There are noanthropic-ratelimit-*headers. PollGET /v1/tokens/usageto see a plan window. - Both Anthropic SDKs retry 429 and 5xx twice by default and honour
retry-after.window_exhaustedandmodel_limit_reachedcan mean hours or days, so retrying them is useless. Setmax_retries=0(maxRetries: 0in TypeScript) and handle retries yourself if you want control; the backoff example in rate limits shows how. - A job that fans out many streams in parallel can hit
concurrency_limitfirst. Limit its parallelism.
Errors#
Errors on /v1/messages use Anthropic's shape: {"type": "error", "error": {"type", "message"}}. Errors that Tokens raises itself (key, credit, plan and limit errors) add error.code, and carry a request_id. Read error.code to tell causes apart. The full list is in errors.
| Situation | Anthropic | Tokens |
|---|---|---|
| Out of credit | 402 billing_error | 402 billing_error, code insufficient_credits, no_funding or outstanding_debt |
| A spend limit you set | 400 invalid_request_error, or 429 for some workspaces | 403 permission_error, code monthly_spend_cap_exceeded for a key's cap |
| Key problem | 401 authentication_error, 403 permission_error | Same types, with codes such as invalid_api_key, key_inactive, key_expired |
| Rate limited | 429 rate_limit_error | 429 rate_limit_error, code rate_limited, concurrency_limit or window_exhausted |
| Provider failure | 500 api_error, 529 overloaded_error | 502 upstream_unreachable, 503 no_upstream_available, 504 upstream_timeout |
| Request too large | 413 request_too_large at 32 MB | 413 at 10 MB (code request_entity_too_large) |
A 403 is not retried by the SDKs. If your code treated the spend-limit 400 or 429 as "stop and alert", map the Tokens 403 monthly_spend_cap_exceeded to the same behaviour.
The provider's own error text is replaced by a generic message, so a 400 for an unsupported parameter does not say which one. Check the request against the model's page on /models.
The request id header is different#
Anthropic returns a request-id header, and the SDKs expose it as _request_id. Tokens does not return that header, so _request_id is None. Read x-tokens-request-id instead:
raw = client.messages.with_raw_response.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=64,
messages=[{"role": "user", "content": "ping"}],
)
print(raw.headers.get("x-tokens-request-id"))
message = raw.parse()Tokens passes on only a short list of provider response headers (content-type, cache-control and retry-after), so anthropic-organization-id, anthropic-workspace-id and the rate-limit headers do not reach you. Support searches by x-tokens-request-id.
Output limits and the credit reservation#
Anthropic does not count max_tokens against your output-token rate limit, so many apps set it high. Tokens reserves the worst-case cost of a request before it forwards it, using your max_tokens. On a low balance or a key close to its cap, a large max_tokens can be refused. On a low balance Tokens can lower it to what you can afford (never below 16), and the answer then ends with stop_reason: "max_tokens". You pay for the tokens used, not the reservation. Set max_tokens to what you need.
More differences#
- No CORS. Calls from a browser fail. Call Tokens from a server.
- Body size. Up to 10 MB, against 32 MB at Anthropic. Large base64 images or PDFs count.
- Latency. Tokens adds one network hop.
- Privacy. Tokens stores usage metadata, not prompt content, and the provider that serves the model sees the prompt. See security and privacy.
- Billing. Tokens bills in USD or BDT from a plan or a wallet, at the catalog price of each model. See plans and wallet.
- Claude Code and other agents. Agents that use the Anthropic protocol have their own setup. See Claude Code.
Test the switch safely#
- Create a second key at /dashboard/keys with a low monthly spend cap and an allowed-models list with only the models you are testing. Neither can be edited later. Tokens counts the worst-case cost of each request against the cap.
- Read the base URL, key and model id from configuration. Both Anthropic SDKs accept the three values from environment variables, so a deploy can switch without a code change.
- Run both for a while. Replay logged requests, or mirror a share of live traffic and discard the Tokens answers. Compare:
| Check | How |
|---|---|
| Quality | Run your own prompts or evals. A different model gives different answers. |
| Prompt caching | Read usage.cache_read_input_tokens on the second call with the same prefix. Zero means no cache hit on this model. |
| Thinking, citations, server tools | Send one request that uses each, to the exact model. Check the response has the blocks you expect. |
| Tool calls | tool_use inputs match your schema. |
stop_reason | More max_tokens than before means the limit is too low for this model. |
| Cost per finished task | Compare Anthropic's invoice with the cost in usage. |
| Errors | Count them by error.code. |
- Ramp up with a feature flag or a percentage. Create the production key with the cap you want before you cut over, because a rotation or a revoke takes effect at once.
Roll back#
- Keep your Anthropic key and billing active until Tokens has carried production traffic for a full billing cycle.
- Put the old base URL (or unset
ANTHROPIC_BASE_URL), key and model ids back in the configuration, and deploy or flip the flag. - Revoke the Tokens key you no longer use at /dashboard/keys. Your wallet balance stays on your account; see the refund policy.
If you use the Files or Batches APIs, those parts never left Anthropic, so they need no rollback.
Where to go next#
- Messages for headers, streaming events and the translation rules.
- Anthropic SDK for Python and TypeScript setup.
- Errors and Rate limits.
- Migrate from OpenAI and Migrate from OpenRouter.
Sources, checked October 2026: Anthropic API overview, errors, rate limits, prompt caching, citations, List Models and the anthropic-sdk-python and anthropic-sdk-typescript sources. Tokens behaviour is from the gateway code and the pages linked above. It has not been tested against a live Anthropic account.