Skip to content

Migrate from Anthropic

Move an app that calls the Anthropic Messages API to Tokens: base URL without /v1, key header, model ids, which Anthropic-only features pass through and which do not, how errors and limits differ, testing and rollback.

On this page

If your app calls the Anthropic Messages API, moving it to Tokens takes three changes: the base URL, the API key and the model id. Tokens serves POST /v1/messages, streaming and POST /v1/messages/count_tokens in Anthropic's format, so the Anthropic SDKs and your message-handling code keep working. What needs care is the set of Anthropic-only features (prompt caching, thinking, citations, the Files and Batches APIs), because they depend on which provider serves the model you pick.

What changes and what stays the same#

SettingAnthropicTokensWhere you set it
Base URLhttps://api.anthropic.com (SDK default)https://tokens.bd, without /v1base_url (Python) or baseURL (Node.js), or ANTHROPIC_BASE_URL
API keyAn Anthropic keytok_live_... from API keysapi_key or apiKey, or ANTHROPIC_API_KEY
Model idAn Anthropic idAn alias in provider/model form from /modelsThe model field of every request
anthropic-workspace-idSelects a workspaceNot used. Remove it.Client options

The SDKs add /v1/messages themselves, so the base URL is the bare host. If you set https://tokens.bd/v1 instead, requests go to /v1/v1/messages and fail with 404. Calling the endpoint yourself with curl is different: the full URL is https://tokens.bd/v1/messages.

The Anthropic Python and TypeScript SDKs read ANTHROPIC_API_KEY, ANTHROPIC_AUTH_TOKEN and ANTHROPIC_BASE_URL when you pass nothing in code (checked in the SDK sources, October 2026). Be careful in a shell where you also run Claude Code or other tools that read the same variables. See the Anthropic SDK page and Claude Code.

What stays the same:

  • The request and response shapes of POST /v1/messages: system, messages, content blocks, max_tokens (still required), tools, tool_choice, stop_sequences and the usage object.
  • Server-Sent Events streaming with the usual event names (message_start, content_block_delta, message_stop). client.messages.stream(...) works.
  • Authentication. Tokens accepts the key in x-api-key (what the SDKs send) or in Authorization: Bearer, so the SDK needs no auth change. If both arrive, the Bearer header wins.
  • The anthropic-version header. It is forwarded to the provider, and the default 2023-06-01 is used when you send none.
  • Error bodies on /v1/messages in Anthropic's shape, so the SDKs raise the same exception classes.

Before and after#

 import os
 import anthropic

-client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
+client = anthropic.Anthropic(
+    base_url="https://tokens.bd",
+    api_key=os.environ["TOKENS_API_KEY"],
+)

 message = client.messages.create(
-    model="your-claude-model",
+    model="deepseek/deepseek-v4.1-flash",
     max_tokens=512,
     messages=[{"role": "user", "content": "Explain idempotency keys in two sentences."}],
 )
 print(message.content[0].text)

Choose the model id#

Anthropic's ids and Tokens ids are different lists. Tokens ids are aliases in provider/model form, and they do not always match the provider's own id. The Anthropic SDK lists them for you, because GET /v1/models answers in Anthropic's format when the request carries anthropic-version:

python
for m in client.models.list():
    print(m.id, m.display_name)

With curl, send anthropic-version to get that shape. With only Authorization: Bearer you get the OpenAI list shape. In both, the list holds only models this key can call, so a key with an allowed-models list, or an account with no plan or wallet balance, sees fewer.

The Anthropic format also works for models from other makers, not only Claude models. That is a feature, but it is also a model change: answers, tool-call behaviour and token counts differ. Prices, context windows and capabilities are on /models; Choosing a model helps you pick.

Which provider serves the model decides the features#

Each model on Tokens is served by one or more providers. When a provider speaks the Messages API itself, your request goes to it unchanged except for the model id, so thinking blocks, prompt caching and anthropic-beta features come back untouched. Tokens prefers such a provider when one serves the model. When a model is only served by a provider that speaks the OpenAI format, Tokens translates your Messages request into a chat completions request and the answer back. The catalog does not say which case applies to a model, so test every feature you depend on against the exact model you choose.

What the translation carries: system (text, joined into one system message), messages with text, image, tool_use and tool_result blocks, max_tokens, stop_sequences, temperature, top_p, stream, tools and tool_choice. Everything else is not copied.

Anthropic featureProvider speaks MessagesTranslated model
Streaming, client tools, imagesPassed throughSupported (translated)
Prompt caching (cache_control)Passed through. usage carries the cache counts the provider reports.Markers dropped. Caching, if any, is the provider's own and not under your control.
Extended or adaptive thinkingPassed throughParameter dropped. No thinking blocks.
document blocks and citationsPassed throughDocument blocks dropped. No citations.
Anthropic-defined server tools (web search, code execution, computer use)Passed to the provider as sent. Tokens does not run them, so it is up to the provider.Every tools entry becomes a function tool, so these do not work.
anthropic-beta headerForwardedNo effect
top_k, metadataPassed throughDropped

Tokens billing reads the provider's usage numbers, and cached input is priced at the model's cache-read rate where the catalog has one. Prompt caching also saves money only on models whose provider supports it; see Choosing a model and Messages.

Endpoints and features Tokens does not have#

Tokens serves these Anthropic endpoints: POST /v1/messages, POST /v1/messages/count_tokens and GET /v1/models. Every other path returns 404 unsupported_endpoint, in Anthropic's error shape (not_found_error).

Anthropic APIOn Tokens
Message Batches (/v1/messages/batches)Not supported. Send requests one by one, or run your own queue.
Files API (/v1/files)Not supported. Send images and documents inline as base64, or as an image URL. A file_id in a content block cannot work.
Skills, Managed Agents, Agents, SessionsNot supported.
GET /v1/models/{id} (models.retrieve)Not supported. List the models and filter.
Models API fields capabilities, lifecycleNot returned. The list has type, id, display_name and created_at, and it has no pagination: has_more is false.
POST /v1/messages/count_tokensSupported. See below.

Count tokens#

POST /v1/messages/count_tokens takes the body of a Messages request without max_tokens and returns {"input_tokens": N}. It never runs the model and is never billed. When a provider that serves the model counts tokens natively you get its exact count. Otherwise Tokens estimates (about four characters per token, plus a fixed amount per image and per turn) and sets the response header x-tokens-estimated: true. Use the estimate for budgeting only.

Differences that can bite#

Rate limits#

Anthropic limits requests, input tokens and output tokens per minute for each model class and reports them in anthropic-ratelimit-* headers. Tokens limits requests per minute per account (60 by default, or your plan's value) and concurrent requests per account (10 with a plan, 3 without). The rate limits page lists no token-per-minute limit. More keys do not raise either limit.

  • A 429 carries retry-after in seconds. There are no anthropic-ratelimit-* headers. Poll GET /v1/tokens/usage to see a plan window.
  • Both Anthropic SDKs retry 429 and 5xx twice by default and honour retry-after. window_exhausted and model_limit_reached can mean hours or days, so retrying them is useless. Set max_retries=0 (maxRetries: 0 in TypeScript) and handle retries yourself if you want control; the backoff example in rate limits shows how.
  • A job that fans out many streams in parallel can hit concurrency_limit first. Limit its parallelism.

Errors#

Errors on /v1/messages use Anthropic's shape: {"type": "error", "error": {"type", "message"}}. Errors that Tokens raises itself (key, credit, plan and limit errors) add error.code, and carry a request_id. Read error.code to tell causes apart. The full list is in errors.

SituationAnthropicTokens
Out of credit402 billing_error402 billing_error, code insufficient_credits, no_funding or outstanding_debt
A spend limit you set400 invalid_request_error, or 429 for some workspaces403 permission_error, code monthly_spend_cap_exceeded for a key's cap
Key problem401 authentication_error, 403 permission_errorSame types, with codes such as invalid_api_key, key_inactive, key_expired
Rate limited429 rate_limit_error429 rate_limit_error, code rate_limited, concurrency_limit or window_exhausted
Provider failure500 api_error, 529 overloaded_error502 upstream_unreachable, 503 no_upstream_available, 504 upstream_timeout
Request too large413 request_too_large at 32 MB413 at 10 MB (code request_entity_too_large)

A 403 is not retried by the SDKs. If your code treated the spend-limit 400 or 429 as "stop and alert", map the Tokens 403 monthly_spend_cap_exceeded to the same behaviour.

The provider's own error text is replaced by a generic message, so a 400 for an unsupported parameter does not say which one. Check the request against the model's page on /models.

The request id header is different#

Anthropic returns a request-id header, and the SDKs expose it as _request_id. Tokens does not return that header, so _request_id is None. Read x-tokens-request-id instead:

python
raw = client.messages.with_raw_response.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=64,
    messages=[{"role": "user", "content": "ping"}],
)
print(raw.headers.get("x-tokens-request-id"))
message = raw.parse()

Tokens passes on only a short list of provider response headers (content-type, cache-control and retry-after), so anthropic-organization-id, anthropic-workspace-id and the rate-limit headers do not reach you. Support searches by x-tokens-request-id.

Output limits and the credit reservation#

Anthropic does not count max_tokens against your output-token rate limit, so many apps set it high. Tokens reserves the worst-case cost of a request before it forwards it, using your max_tokens. On a low balance or a key close to its cap, a large max_tokens can be refused. On a low balance Tokens can lower it to what you can afford (never below 16), and the answer then ends with stop_reason: "max_tokens". You pay for the tokens used, not the reservation. Set max_tokens to what you need.

More differences#

  • No CORS. Calls from a browser fail. Call Tokens from a server.
  • Body size. Up to 10 MB, against 32 MB at Anthropic. Large base64 images or PDFs count.
  • Latency. Tokens adds one network hop.
  • Privacy. Tokens stores usage metadata, not prompt content, and the provider that serves the model sees the prompt. See security and privacy.
  • Billing. Tokens bills in USD or BDT from a plan or a wallet, at the catalog price of each model. See plans and wallet.
  • Claude Code and other agents. Agents that use the Anthropic protocol have their own setup. See Claude Code.

Test the switch safely#

  1. Create a second key at /dashboard/keys with a low monthly spend cap and an allowed-models list with only the models you are testing. Neither can be edited later. Tokens counts the worst-case cost of each request against the cap.
  2. Read the base URL, key and model id from configuration. Both Anthropic SDKs accept the three values from environment variables, so a deploy can switch without a code change.
  3. Run both for a while. Replay logged requests, or mirror a share of live traffic and discard the Tokens answers. Compare:
CheckHow
QualityRun your own prompts or evals. A different model gives different answers.
Prompt cachingRead usage.cache_read_input_tokens on the second call with the same prefix. Zero means no cache hit on this model.
Thinking, citations, server toolsSend one request that uses each, to the exact model. Check the response has the blocks you expect.
Tool callstool_use inputs match your schema.
stop_reasonMore max_tokens than before means the limit is too low for this model.
Cost per finished taskCompare Anthropic's invoice with the cost in usage.
ErrorsCount them by error.code.
  1. Ramp up with a feature flag or a percentage. Create the production key with the cap you want before you cut over, because a rotation or a revoke takes effect at once.

Roll back#

  1. Keep your Anthropic key and billing active until Tokens has carried production traffic for a full billing cycle.
  2. Put the old base URL (or unset ANTHROPIC_BASE_URL), key and model ids back in the configuration, and deploy or flip the flag.
  3. Revoke the Tokens key you no longer use at /dashboard/keys. Your wallet balance stays on your account; see the refund policy.

If you use the Files or Batches APIs, those parts never left Anthropic, so they need no rollback.

Where to go next#

Sources, checked October 2026: Anthropic API overview, errors, rate limits, prompt caching, citations, List Models and the anthropic-sdk-python and anthropic-sdk-typescript sources. Tokens behaviour is from the gateway code and the pages linked above. It has not been tested against a live Anthropic account.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.