# Migrate from Anthropic

> Move an app that calls the Anthropic Messages API to Tokens: base URL without /v1, key header, model ids, which Anthropic-only features pass through and which do not, how errors and limits differ, testing and rollback.

If your app calls the Anthropic Messages API, moving it to Tokens takes three changes: the base URL, the API key and the model id. Tokens serves `POST /v1/messages`, streaming and `POST /v1/messages/count_tokens` in Anthropic's format, so the Anthropic SDKs and your message-handling code keep working. What needs care is the set of Anthropic-only features (prompt caching, thinking, citations, the Files and Batches APIs), because they depend on which provider serves the model you pick.

## What changes and what stays the same

| Setting           | Anthropic                                  | Tokens                                              | Where you set it                                                    |
| ----------------- | ------------------------------------------ | --------------------------------------------------- | ------------------------------------------------------------------- |
| Base URL          | `https://api.anthropic.com` (SDK default)  | `https://tokens.bd`, **without `/v1`**         | `base_url` (Python) or `baseURL` (Node.js), or `ANTHROPIC_BASE_URL` |
| API key           | An Anthropic key                           | `tok_live_...` from [API keys](/docs/api-keys)      | `api_key` or `apiKey`, or `ANTHROPIC_API_KEY`                       |
| Model id          | An Anthropic id                            | An alias in `provider/model` form from `/models`    | The `model` field of every request                                  |
| `anthropic-workspace-id` | Selects a workspace                 | Not used. Remove it.                                | Client options                                                      |

The SDKs add `/v1/messages` themselves, so the base URL is the bare host. If you set `https://tokens.bd/v1` instead, requests go to `/v1/v1/messages` and fail with 404. Calling the endpoint yourself with curl is different: the full URL is `https://tokens.bd/v1/messages`.

The Anthropic Python and TypeScript SDKs read `ANTHROPIC_API_KEY`, `ANTHROPIC_AUTH_TOKEN` and `ANTHROPIC_BASE_URL` when you pass nothing in code (checked in the SDK sources, October 2026). Be careful in a shell where you also run Claude Code or other tools that read the same variables. See the [Anthropic SDK page](/docs/anthropic-sdk) and [Claude Code](/docs/claude-code).

What stays the same:

- The request and response shapes of `POST /v1/messages`: `system`, `messages`, content blocks, `max_tokens` (still required), `tools`, `tool_choice`, `stop_sequences` and the `usage` object.
- Server-Sent Events streaming with the usual event names (`message_start`, `content_block_delta`, `message_stop`). `client.messages.stream(...)` works.
- Authentication. Tokens accepts the key in `x-api-key` (what the SDKs send) or in `Authorization: Bearer`, so the SDK needs no auth change. If both arrive, the Bearer header wins.
- The `anthropic-version` header. It is forwarded to the provider, and the default `2023-06-01` is used when you send none.
- Error bodies on `/v1/messages` in Anthropic's shape, so the SDKs raise the same exception classes.

## Before and after

:::code-tabs

```diff title="Python"
 import os
 import anthropic

-client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
+client = anthropic.Anthropic(
+    base_url="https://tokens.bd",
+    api_key=os.environ["TOKENS_API_KEY"],
+)

 message = client.messages.create(
-    model="your-claude-model",
+    model="deepseek/deepseek-v4.1-flash",
     max_tokens=512,
     messages=[{"role": "user", "content": "Explain idempotency keys in two sentences."}],
 )
 print(message.content[0].text)
```

```diff title="Node.js"
 import Anthropic from "@anthropic-ai/sdk";

-const client = new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY });
+const client = new Anthropic({
+  baseURL: "https://tokens.bd",
+  apiKey: process.env.TOKENS_API_KEY,
+});

 const message = await client.messages.create({
-  model: "your-claude-model",
+  model: "deepseek/deepseek-v4.1-flash",
   max_tokens: 512,
   messages: [{ role: "user", content: "Explain idempotency keys in two sentences." }],
 });
 console.log(message.content[0].text);
```

```diff title="curl"
-curl https://api.anthropic.com/v1/messages \
-  -H "x-api-key: $ANTHROPIC_API_KEY" \
+curl https://tokens.bd/v1/messages \
+  -H "x-api-key: $TOKENS_API_KEY" \
   -H "anthropic-version: 2023-06-01" \
   -H "content-type: application/json" \
   -d '{
-    "model": "your-claude-model",
+    "model": "deepseek/deepseek-v4.1-flash",
     "max_tokens": 512,
     "messages": [{"role": "user", "content": "Explain idempotency keys in two sentences."}]
   }'
```

:::

### Choose the model id

Anthropic's ids and Tokens ids are different lists. Tokens ids are aliases in `provider/model` form, and they do not always match the provider's own id. The Anthropic SDK lists them for you, because `GET /v1/models` answers in Anthropic's format when the request carries `anthropic-version`:

```python
for m in client.models.list():
    print(m.id, m.display_name)
```

With curl, send `anthropic-version` to get that shape. With only `Authorization: Bearer` you get the OpenAI list shape. In both, the list holds only models this key can call, so a key with an allowed-models list, or an account with no plan or wallet balance, sees fewer.

The Anthropic format also works for models from other makers, not only Claude models. That is a feature, but it is also a model change: answers, tool-call behaviour and token counts differ. Prices, context windows and capabilities are on [/models](/models); [Choosing a model](/docs/choosing-a-model) helps you pick.

## Which provider serves the model decides the features

Each model on Tokens is served by one or more providers. When a provider speaks the Messages API itself, your request goes to it unchanged except for the model id, so thinking blocks, prompt caching and `anthropic-beta` features come back untouched. Tokens prefers such a provider when one serves the model. When a model is only served by a provider that speaks the OpenAI format, Tokens translates your Messages request into a chat completions request and the answer back. The catalog does not say which case applies to a model, so test every feature you depend on against the exact model you choose.

What the translation carries: `system` (text, joined into one system message), `messages` with text, image, `tool_use` and `tool_result` blocks, `max_tokens`, `stop_sequences`, `temperature`, `top_p`, `stream`, `tools` and `tool_choice`. Everything else is not copied.

| Anthropic feature                      | Provider speaks Messages                                | Translated model                                                              |
| -------------------------------------- | ------------------------------------------------------- | ----------------------------------------------------------------------------- |
| Streaming, client tools, images        | Passed through                                          | Supported (translated)                                                        |
| Prompt caching (`cache_control`)       | Passed through. `usage` carries the cache counts the provider reports.         | Markers dropped. Caching, if any, is the provider's own and not under your control. |
| Extended or adaptive `thinking`        | Passed through                                          | Parameter dropped. No thinking blocks.                                        |
| `document` blocks and `citations`      | Passed through                                          | Document blocks dropped. No citations.                                        |
| Anthropic-defined server tools (web search, code execution, computer use) | Passed to the provider as sent. Tokens does not run them, so it is up to the provider. | Every `tools` entry becomes a function tool, so these do not work.         |
| `anthropic-beta` header                | Forwarded                                               | No effect                                                                     |
| `top_k`, `metadata`                    | Passed through                                          | Dropped                                                                       |

Tokens billing reads the provider's `usage` numbers, and cached input is priced at the model's cache-read rate where the catalog has one. Prompt caching also saves money only on models whose provider supports it; see [Choosing a model](/docs/choosing-a-model) and [Messages](/docs/messages).

## Endpoints and features Tokens does not have

Tokens serves these Anthropic endpoints: `POST /v1/messages`, `POST /v1/messages/count_tokens` and `GET /v1/models`. Every other path returns 404 `unsupported_endpoint`, in Anthropic's error shape (`not_found_error`).

| Anthropic API                                  | On Tokens                                                                                                                  |
| ---------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- |
| Message Batches (`/v1/messages/batches`)       | Not supported. Send requests one by one, or run your own queue.                                                            |
| Files API (`/v1/files`)                        | Not supported. Send images and documents inline as base64, or as an image URL. A `file_id` in a content block cannot work. |
| Skills, Managed Agents, Agents, Sessions       | Not supported.                                                                                                             |
| `GET /v1/models/{id}` (`models.retrieve`)      | Not supported. List the models and filter.                                                                                 |
| Models API fields `capabilities`, `lifecycle`  | Not returned. The list has `type`, `id`, `display_name` and `created_at`, and it has no pagination: `has_more` is `false`. |
| `POST /v1/messages/count_tokens`               | Supported. See below.                                                                                                      |

### Count tokens

`POST /v1/messages/count_tokens` takes the body of a Messages request without `max_tokens` and returns `{"input_tokens": N}`. It never runs the model and is never billed. When a provider that serves the model counts tokens natively you get its exact count. Otherwise Tokens estimates (about four characters per token, plus a fixed amount per image and per turn) and sets the response header `x-tokens-estimated: true`. Use the estimate for budgeting only.

## Differences that can bite

### Rate limits

Anthropic limits requests, input tokens and output tokens per minute for each model class and reports them in `anthropic-ratelimit-*` headers. Tokens limits **requests per minute per account** (60 by default, or your plan's value) and **concurrent requests per account** (10 with a plan, 3 without). The [rate limits](/docs/rate-limits) page lists no token-per-minute limit. More keys do not raise either limit.

- A 429 carries `retry-after` in seconds. There are no `anthropic-ratelimit-*` headers. Poll `GET /v1/tokens/usage` to see a plan window.
- Both Anthropic SDKs retry 429 and 5xx twice by default and honour `retry-after`. `window_exhausted` and `model_limit_reached` can mean hours or days, so retrying them is useless. Set `max_retries=0` (`maxRetries: 0` in TypeScript) and handle retries yourself if you want control; the backoff example in [rate limits](/docs/rate-limits) shows how.
- A job that fans out many streams in parallel can hit `concurrency_limit` first. Limit its parallelism.

### Errors

Errors on `/v1/messages` use Anthropic's shape: `{"type": "error", "error": {"type", "message"}}`. Errors that Tokens raises itself (key, credit, plan and limit errors) add `error.code`, and carry a `request_id`. Read `error.code` to tell causes apart. The full list is in [errors](/docs/errors).

| Situation                | Anthropic                                         | Tokens                                                                              |
| ------------------------ | ------------------------------------------------- | ----------------------------------------------------------------------------------- |
| Out of credit            | 402 `billing_error`                               | 402 `billing_error`, code `insufficient_credits`, `no_funding` or `outstanding_debt` |
| A spend limit you set    | 400 `invalid_request_error`, or 429 for some workspaces | 403 `permission_error`, code `monthly_spend_cap_exceeded` for a key's cap     |
| Key problem              | 401 `authentication_error`, 403 `permission_error`| Same types, with codes such as `invalid_api_key`, `key_inactive`, `key_expired`     |
| Rate limited             | 429 `rate_limit_error`                            | 429 `rate_limit_error`, code `rate_limited`, `concurrency_limit` or `window_exhausted` |
| Provider failure         | 500 `api_error`, 529 `overloaded_error`           | 502 `upstream_unreachable`, 503 `no_upstream_available`, 504 `upstream_timeout`     |
| Request too large        | 413 `request_too_large` at 32 MB                  | 413 at **10 MB** (code `request_entity_too_large`)                                  |

A 403 is not retried by the SDKs. If your code treated the spend-limit 400 or 429 as "stop and alert", map the Tokens 403 `monthly_spend_cap_exceeded` to the same behaviour.

The provider's own error text is replaced by a generic message, so a 400 for an unsupported parameter does not say which one. Check the request against the model's page on [/models](/models).

### The request id header is different

Anthropic returns a `request-id` header, and the SDKs expose it as `_request_id`. Tokens does not return that header, so `_request_id` is `None`. Read `x-tokens-request-id` instead:

```python
raw = client.messages.with_raw_response.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=64,
    messages=[{"role": "user", "content": "ping"}],
)
print(raw.headers.get("x-tokens-request-id"))
message = raw.parse()
```

Tokens passes on only a short list of provider response headers (`content-type`, `cache-control` and `retry-after`), so `anthropic-organization-id`, `anthropic-workspace-id` and the rate-limit headers do not reach you. [Support](/docs/support) searches by `x-tokens-request-id`.

### Output limits and the credit reservation

Anthropic does not count `max_tokens` against your output-token rate limit, so many apps set it high. Tokens reserves the worst-case cost of a request before it forwards it, using your `max_tokens`. On a low balance or a key close to its cap, a large `max_tokens` can be refused. On a low balance Tokens can lower it to what you can afford (never below 16), and the answer then ends with `stop_reason: "max_tokens"`. You pay for the tokens used, not the reservation. Set `max_tokens` to what you need.

### More differences

- **No CORS.** Calls from a browser fail. Call Tokens from a server.
- **Body size.** Up to 10 MB, against 32 MB at Anthropic. Large base64 images or PDFs count.
- **Latency.** Tokens adds one network hop.
- **Privacy.** Tokens stores usage metadata, not prompt content, and the provider that serves the model sees the prompt. See [security and privacy](/docs/security-and-privacy).
- **Billing.** Tokens bills in USD or BDT from a plan or a wallet, at the catalog price of each model. See [plans and wallet](/docs/plans-and-wallet).
- **Claude Code and other agents.** Agents that use the Anthropic protocol have their own setup. See [Claude Code](/docs/claude-code).

## Test the switch safely

1. **Create a second key** at [/dashboard/keys](/dashboard/keys) with a **low monthly spend cap** and an **allowed-models list** with only the models you are testing. Neither can be edited later. Tokens counts the worst-case cost of each request against the cap.
2. **Read the base URL, key and model id from configuration.** Both Anthropic SDKs accept the three values from environment variables, so a deploy can switch without a code change.
3. **Run both for a while.** Replay logged requests, or mirror a share of live traffic and discard the Tokens answers. Compare:

| Check                                | How                                                                                                  |
| ------------------------------------ | ---------------------------------------------------------------------------------------------------- |
| Quality                              | Run your own prompts or evals. A different model gives different answers.                            |
| Prompt caching                       | Read `usage.cache_read_input_tokens` on the second call with the same prefix. Zero means no cache hit on this model. |
| Thinking, citations, server tools    | Send one request that uses each, to the exact model. Check the response has the blocks you expect.   |
| Tool calls                           | `tool_use` inputs match your schema.                                                                 |
| `stop_reason`                        | More `max_tokens` than before means the limit is too low for this model.                             |
| Cost per finished task               | Compare Anthropic's invoice with the cost in [usage](/docs/usage-and-alerts).                         |
| Errors                               | Count them by `error.code`.                                                                          |

4. **Ramp up** with a feature flag or a percentage. Create the production key with the cap you want before you cut over, because a rotation or a revoke takes effect at once.

## Roll back

1. Keep your Anthropic key and billing active until Tokens has carried production traffic for a full billing cycle.
2. Put the old base URL (or unset `ANTHROPIC_BASE_URL`), key and model ids back in the configuration, and deploy or flip the flag.
3. Revoke the Tokens key you no longer use at [/dashboard/keys](/dashboard/keys). Your wallet balance stays on your account; see the [refund policy](/refund-policy).

If you use the Files or Batches APIs, those parts never left Anthropic, so they need no rollback.

## Where to go next

- [Messages](/docs/messages) for headers, streaming events and the translation rules.
- [Anthropic SDK](/docs/anthropic-sdk) for Python and TypeScript setup.
- [Errors](/docs/errors) and [Rate limits](/docs/rate-limits).
- [Migrate from OpenAI](/docs/migrate-from-openai) and [Migrate from OpenRouter](/docs/migrate-from-openrouter).

Sources, checked October 2026: Anthropic [API overview](https://platform.claude.com/docs/en/api/overview), [errors](https://platform.claude.com/docs/en/api/errors), [rate limits](https://platform.claude.com/docs/en/api/rate-limits), [prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching), [citations](https://platform.claude.com/docs/en/build-with-claude/citations), [List Models](https://platform.claude.com/docs/en/api/models/list) and the [anthropic-sdk-python](https://github.com/anthropics/anthropic-sdk-python) and [anthropic-sdk-typescript](https://github.com/anthropics/anthropic-sdk-typescript) sources. Tokens behaviour is from the gateway code and the pages linked above. It has not been tested against a live Anthropic account.

---
Page: https://tokens.bd/docs/migrate-from-anthropic
