# Migrate from OpenRouter

> Move an app from OpenRouter to Tokens: the base URL, key and model ids to change, what happens to OpenRouter routing fields, headers and model suffixes, how errors and limits differ, and how to test and roll back.

OpenRouter and Tokens both give you one OpenAI-compatible endpoint for models from many providers, so an app written for OpenRouter usually needs the same three changes as any OpenAI app: the base URL, the key and the model id. The work is in what OpenRouter does on top of the OpenAI format: provider routing, model fallbacks, model suffixes, attribution headers and cost fields. Tokens does not do these. This page says exactly what happens to each one.

## What changes and what stays the same

| Setting  | OpenRouter                                         | Tokens                                              | Where you set it                       |
| -------- | -------------------------------------------------- | --------------------------------------------------- | -------------------------------------- |
| Base URL | `https://openrouter.ai/api/v1`                     | `https://tokens.bd/v1`                               | `base_url` or `baseURL` in the SDK     |
| API key  | An OpenRouter key                                  | `tok_live_...` from [API keys](/docs/api-keys)      | `api_key` or `apiKey`                  |
| Model id | `provider/model`, an OpenRouter slug               | `provider/model`, a Tokens alias from `/models`     | The `model` field of every request     |
| Headers  | `HTTP-Referer`, `X-Title` (attribution, optional)  | Not used. Remove them.                              | `default_headers` or `defaultHeaders`  |

What stays the same:

- The OpenAI request and response format of `POST /v1/chat/completions`, with `messages`, `tools`, `stream` and the `usage` object. See [Chat Completions](/docs/chat-completions).
- `Authorization: Bearer <key>` authentication, and Server-Sent Events streaming.
- Any OpenAI SDK, or the Vercel AI SDK, LangChain and similar libraries you pointed at OpenRouter. Change their base URL and key in the same place.

## Before and after

:::code-tabs

```diff title="Python"
 import os
 from openai import OpenAI

 client = OpenAI(
-    base_url="https://openrouter.ai/api/v1",
-    api_key=os.environ["OPENROUTER_API_KEY"],
-    default_headers={
-        "HTTP-Referer": "https://example.com",
-        "X-Title": "My app",
-    },
+    base_url="https://tokens.bd/v1",
+    api_key=os.environ["TOKENS_API_KEY"],
 )

 resp = client.chat.completions.create(
-    model="provider/openrouter-model-slug",
+    model="deepseek/deepseek-v4.1-flash",
     messages=[{"role": "user", "content": "What does HTTP 429 mean?"}],
     max_tokens=300,
 )
 print(resp.choices[0].message.content)
```

```diff title="Node.js"
 import OpenAI from "openai";

 const client = new OpenAI({
-  baseURL: "https://openrouter.ai/api/v1",
-  apiKey: process.env.OPENROUTER_API_KEY,
-  defaultHeaders: { "HTTP-Referer": "https://example.com", "X-Title": "My app" },
+  baseURL: "https://tokens.bd/v1",
+  apiKey: process.env.TOKENS_API_KEY,
 });

 const resp = await client.chat.completions.create({
-  model: "provider/openrouter-model-slug",
+  model: "deepseek/deepseek-v4.1-flash",
   messages: [{ role: "user", content: "What does HTTP 429 mean?" }],
   max_tokens: 300,
 });
 console.log(resp.choices[0].message.content);
```

```diff title="curl"
-curl https://openrouter.ai/api/v1/chat/completions \
-  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
-  -H "HTTP-Referer: https://example.com" \
-  -H "X-Title: My app" \
+curl https://tokens.bd/v1/chat/completions \
+  -H "Authorization: Bearer $TOKENS_API_KEY" \
   -H "Content-Type: application/json" \
   -d '{
-    "model": "provider/openrouter-model-slug",
+    "model": "deepseek/deepseek-v4.1-flash",
     "messages": [{"role": "user", "content": "What does HTTP 429 mean?"}],
     "max_tokens": 300
   }'
```

:::

Keep `Content-Type: application/json` when you use curl. Without it Tokens does not read the body and answers 400 `invalid_request`.

If you pass the settings through the `OPENAI_BASE_URL` and `OPENAI_API_KEY` variables, which the official OpenAI Python and Node.js SDKs read (checked in the SDK sources, October 2026), change the variables and the model id and nothing else.

## Model ids look the same and are not

OpenRouter slugs and Tokens aliases both use `provider/model`. They are separate lists. A slug that works on OpenRouter can be unknown on Tokens, or can name a different version, and Tokens aliases do not always match the provider's own id. Do not rewrite ids by pattern.

1. List what your key can call: `curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY"`. Each `data[].id` is valid for that key.
2. Prices, context windows and capabilities are on [/models](/models). The API list has ids only. See [Choosing a model](/docs/choosing-a-model).
3. Keep a table from your old ids to the new ones in configuration, so the mapping is not scattered in code.

The `model` value must match an alias exactly. An unknown id returns 404 `model_not_found`.

## OpenRouter features, one by one

This is what Tokens does with each OpenRouter feature. It comes from the gateway code, not from guesses.

**Model suffixes.** OpenRouter documents `:free`, `:nitro`, `:floor`, `:exacto` (and the deprecated `:online`, `:thinking`, `:extended`) as part of the model id. Tokens looks up the whole string as the alias, so `some/model:nitro` returns 404 `model_not_found`. Remove the suffix. There is no equivalent: Tokens has no suffixes and no per-request provider sorting. A cheaper or faster model is a different model id.

**Router models.** `openrouter/auto` and other OpenRouter router ids are not in the Tokens catalog (404 `model_not_found`). Pick a fixed model.

**Provider routing (`provider`).** Tokens does not read this field. It does not reject it either: for a model whose provider speaks the OpenAI protocol (the usual case), the rest of the request body is passed to the provider as sent. Whether the provider ignores the field or answers 400 is up to the provider, and Tokens did not check any provider's behaviour. Remove it. Which provider serves a model is Tokens' choice, not yours, and the `model` field of the response is always the id you asked for.

**Fallbacks (`models`, `route`).** Not implemented. The same rule as above: forwarded, not acted on, so the field does nothing. Tokens does fail over, but only between sources of the same model: on 429, 502, 503, 504 or a dropped connection, and before any byte of the answer has reached you, it retries on another source of that model. It never switches to a different model. If you want a model fallback, do it in your code:

```python
import os

import openai
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], max_retries=0)

MODELS = ["deepseek/deepseek-v4.1-flash", "your-second-model-id"]  # ids from GET /v1/models


def ask(messages):
    last = None
    for model in MODELS:
        try:
            return client.chat.completions.create(model=model, messages=messages, max_tokens=512)
        except openai.APIStatusError as e:
            if e.status_code not in (429, 500, 502, 503, 504) or e.code == "window_exhausted":
                raise
            last = e
    raise last
```

`window_exhausted` is excluded because it is an account-wide plan window: another model will hit it too.

**Prompt transforms and plugins (`transforms`, `plugins`).** OpenRouter lets you compress long prompts, parse files, search the web and repair responses through these. Tokens does none of it. The fields are forwarded to the provider as unknown body fields, like `provider`. Trim the prompt yourself: Tokens does not shorten a prompt that is over the model's context window, so the provider decides what happens to it.

**`reasoning` and `usage` fields.** Forwarded like the others. Where a model supports reasoning, the provider decides what it does with them. The `usage: {"include": true}` request field is not needed: OpenRouter's documentation calls it deprecated and returns usage anyway, and Tokens' answers always carry the `usage` object the provider sent.

**Attribution headers.** `HTTP-Referer`, `X-Title` and the `X-OpenRouter-*` headers only matter on OpenRouter, for app pages and rankings. Tokens forwards a short allow-list of request headers to providers (`content-type`, `accept`, `openai-beta` and `anthropic-version`, among a few) and drops the rest without an error. Sending them is harmless and has no effect. Remove them so the code does not suggest otherwise.

**If you are not sure a field is safe.** Test it against the model you plan to use, on the test key described below. A request that works is the only proof that a provider accepts a field.

When a model is served only by a provider that speaks Anthropic's Messages protocol, Tokens translates your chat completions request instead of forwarding it. Only these fields are carried: `messages` (text, images, tool calls and tool results), `max_tokens` or `max_completion_tokens`, `temperature`, `top_p`, `stop`, `stream`, `tools`, `tool_choice`, `parallel_tool_calls` and `user`. Everything else is dropped, including `n`, `response_format`, `seed`, `logprobs` and the penalties. `temperature` is capped at 1, and `max_tokens` defaults to 4096 when you send none.

## Response and usage differences

- **`model`** in the response is always the id you asked for, whichever provider answered. OpenRouter's `model` shows the model it routed to, so code that logged it to see the routing outcome gets nothing new.
- **Cost fields.** OpenRouter adds `cost`, `cost_details` and `native_finish_reason` to the response. Tokens does not add anything to the provider's body. The cost of each request is in your [usage dashboard](/docs/usage-and-alerts) and the CSV export, priced at the model's catalog price. If you read `usage.cost` to bill your own customers, switch to the export, or compute the cost from the token counts and the price on [/models](/models).
- **Usage lookups.** There is no `GET /generation?id=` or `GET /key`. `GET /v1/tokens/usage` returns plan windows, wallet balance and the calling key's cap ([Models and usage](/docs/models-and-usage)). The response header `x-tokens-request-id` is the id to log and to quote to [support](/docs/support).
- **Streaming.** The usage chunk at the end of a stream appears when you send `stream_options: {"include_usage": true}`. See [Streaming](/docs/streaming).

## Differences that can bite

### Errors

OpenRouter errors are `{"error": {"code": 429, "message": "...", "metadata": {...}}}` with a numeric `code`. Tokens errors are `{"error": {"message", "type", "code", "param", "request_id"}}` with a **string** `code` such as `rate_limited`. Code that reads `error.code` as a number, or reads `error.metadata`, needs a change: use the HTTP status for the class of error and `error.code` for the cause. The full table is in [errors](/docs/errors).

| Situation               | OpenRouter                                    | Tokens                                                                   |
| ----------------------- | --------------------------------------------- | ------------------------------------------------------------------------ |
| Out of credit           | 402                                           | 402 `insufficient_credits`, `no_funding` or `outstanding_debt`           |
| Rate limited            | 429                                           | 429 `rate_limited`, `concurrency_limit`, `window_exhausted` or `model_limit_reached` |
| Key limit you set       | Key credit limit: 402                         | 403 `monthly_spend_cap_exceeded` (see [API keys](/docs/api-keys))        |
| Timeout                 | 408                                           | 504 `upstream_timeout`                                                   |
| Model down              | 502                                           | 502 `upstream_unreachable` or the provider's 5xx as `upstream_error`     |
| No provider available   | 503                                           | 503 `no_upstream_available`                                              |

OpenRouter documents that a failure after streaming starts arrives as an error event in a 200 response. Keep any in-body error check you added for that.

The provider's own error message is replaced by a generic one, and `metadata.provider_*` details do not exist.

### Rate limits

OpenRouter limits free models to 20 requests per minute and puts no platform cap on paid models. Tokens limits **requests per minute per account** (60 by default, or your plan's value) and **concurrent requests per account** (10 with a plan, 3 without). An agent or batch job that ran with high parallelism on OpenRouter can hit `concurrency_limit`; cap its parallelism. More keys do not raise the limits.

- A 429 carries `Retry-After` in seconds. There are no `X-RateLimit-*` headers. Poll `GET /v1/tokens/usage` to see a plan window.
- `window_exhausted` and `model_limit_reached` can mean hours or days. Do not retry them in a loop.
- See [rate limits](/docs/rate-limits) for the numbers and a backoff example.

### Other differences

- **No `/api/v1` path.** Tokens' path is `/v1`: `https://tokens.bd/v1`.
- **Unsupported endpoints.** Images, audio, files, batches, assistants, fine-tuning and moderations return 404 `unsupported_endpoint`. Embeddings work only for embedding models. Supported: `/v1/chat/completions`, `/v1/responses`, `/v1/completions` (legacy), `/v1/embeddings`, `/v1/models`, `/v1/messages` and `/v1/messages/count_tokens`.
- **Parameters depend on the model.** Besides `model` and `n` (1 to 4), Tokens does not validate the body. Tool calling, `response_format`, vision and reasoning depend on the model. The gateway renames `max_tokens` to `max_completion_tokens` and removes `temperature` and `top_p` unless they equal 1 for OpenAI o-series and GPT-5 and later models on chat completions.
- **Output reservation.** Tokens reserves the worst-case cost of a request, using `max_tokens` or 8,192 output tokens if you set none. A large `max_tokens` on a low balance or a nearly full key cap can be refused or lowered. Set it to what you need.
- **No CORS.** Calls from a browser fail. Call Tokens from a server.
- **Body size.** Up to 10 MB.
- **Privacy.** Tokens stores usage metadata, not prompt content. The provider that serves a model sees the prompt and its policy applies. See [security and privacy](/docs/security-and-privacy).
- **Payment.** Tokens bills in USD or BDT from a plan or a wallet; there are no OpenRouter credits. See [plans and wallet](/docs/plans-and-wallet).

If you call OpenRouter's Anthropic-style Messages endpoint with an Anthropic SDK, read [Migrate from Anthropic](/docs/migrate-from-anthropic) too. Tokens serves `POST /v1/messages`, with a base URL without `/v1`.

## Test the switch safely

1. **Create a second key** at [/dashboard/keys](/dashboard/keys) with a **low monthly spend cap** and an **allowed-models list** that holds only the models you are testing. Neither can be edited afterwards. Tokens counts the worst-case cost of each request against the cap.
2. **Read the base URL, key and model id from configuration**, not from constants, so the switch is a config change.
3. **Run both for a while.** Replay logged requests, or mirror a share of live traffic to Tokens and discard its answers. Compare the same prompts on both:

| Check                  | How                                                                                         |
| ---------------------- | ------------------------------------------------------------------------------------------- |
| Quality                | Run your own prompts or evals. The two models in a pair are rarely the same model.           |
| Fields you stopped sending | Run once without `provider`, `models`, `transforms` and the headers. Did anything depend on them? |
| Tool calls             | Valid JSON arguments for your schema, on the model you chose.                               |
| `finish_reason`        | More `length` than before means `max_tokens` is too low.                                    |
| Cost per finished task | OpenRouter's `usage.cost` against the cost in [usage](/docs/usage-and-alerts).              |
| Errors                 | Count them by `error.code`.                                                                 |

4. **Ramp up** with a feature flag or a percentage. Create the production key with the cap you want before you cut over.

## Roll back

1. Keep your OpenRouter key and credits until Tokens has carried production traffic for a full billing cycle.
2. Put the old base URL, key, model ids and headers back in the configuration and deploy or flip the flag.
3. Revoke the Tokens key you no longer use at [/dashboard/keys](/dashboard/keys). Your wallet balance stays on your account; see the [refund policy](/refund-policy).

Because the id mapping lives in configuration, rolling back is the same change in reverse.

## Where to go next

- [Chat Completions](/docs/chat-completions), [Streaming](/docs/streaming) and [Tool calling](/docs/tool-calling).
- [Errors](/docs/errors) and [Rate limits](/docs/rate-limits).
- [Migrate from OpenAI](/docs/migrate-from-openai) for the shared parts of an OpenAI-style switch.

Sources, checked October 2026: OpenRouter [API overview](https://openrouter.ai/docs/api-reference/overview), [errors](https://openrouter.ai/docs/api-reference/errors), [limits](https://openrouter.ai/docs/api-reference/limits), [provider routing](https://openrouter.ai/docs/guides/routing/provider-selection), [model fallbacks](https://openrouter.ai/docs/guides/routing/model-fallbacks), [model variants](https://openrouter.ai/docs/guides/routing/model-variants), [usage accounting](https://openrouter.ai/docs/guides/guides/usage-accounting) and [app attribution](https://openrouter.ai/docs/app-attribution). Tokens behaviour is from the gateway code and the pages linked above. It has not been tested against a live OpenRouter account.

---
Page: https://tokens.bd/docs/migrate-from-openrouter
