OpenRouter and Tokens both give you one OpenAI-compatible endpoint for models from many providers, so an app written for OpenRouter usually needs the same three changes as any OpenAI app: the base URL, the key and the model id. The work is in what OpenRouter does on top of the OpenAI format: provider routing, model fallbacks, model suffixes, attribution headers and cost fields. Tokens does not do these. This page says exactly what happens to each one.
What changes and what stays the same#
| Setting | OpenRouter | Tokens | Where you set it |
|---|---|---|---|
| Base URL | https://openrouter.ai/api/v1 | https://tokens.bd/v1 | base_url or baseURL in the SDK |
| API key | An OpenRouter key | tok_live_... from API keys | api_key or apiKey |
| Model id | provider/model, an OpenRouter slug | provider/model, a Tokens alias from /models | The model field of every request |
| Headers | HTTP-Referer, X-Title (attribution, optional) | Not used. Remove them. | default_headers or defaultHeaders |
What stays the same:
- The OpenAI request and response format of
POST /v1/chat/completions, withmessages,tools,streamand theusageobject. See Chat Completions. Authorization: Bearer <key>authentication, and Server-Sent Events streaming.- Any OpenAI SDK, or the Vercel AI SDK, LangChain and similar libraries you pointed at OpenRouter. Change their base URL and key in the same place.
Before and after#
import os
from openai import OpenAI
client = OpenAI(
- base_url="https://openrouter.ai/api/v1",
- api_key=os.environ["OPENROUTER_API_KEY"],
- default_headers={
- "HTTP-Referer": "https://example.com",
- "X-Title": "My app",
- },
+ base_url="https://tokens.bd/v1",
+ api_key=os.environ["TOKENS_API_KEY"],
)
resp = client.chat.completions.create(
- model="provider/openrouter-model-slug",
+ model="deepseek/deepseek-v4.1-flash",
messages=[{"role": "user", "content": "What does HTTP 429 mean?"}],
max_tokens=300,
)
print(resp.choices[0].message.content) import OpenAI from "openai";
const client = new OpenAI({
- baseURL: "https://openrouter.ai/api/v1",
- apiKey: process.env.OPENROUTER_API_KEY,
- defaultHeaders: { "HTTP-Referer": "https://example.com", "X-Title": "My app" },
+ baseURL: "https://tokens.bd/v1",
+ apiKey: process.env.TOKENS_API_KEY,
});
const resp = await client.chat.completions.create({
- model: "provider/openrouter-model-slug",
+ model: "deepseek/deepseek-v4.1-flash",
messages: [{ role: "user", content: "What does HTTP 429 mean?" }],
max_tokens: 300,
});
console.log(resp.choices[0].message.content);-curl https://openrouter.ai/api/v1/chat/completions \
- -H "Authorization: Bearer $OPENROUTER_API_KEY" \
- -H "HTTP-Referer: https://example.com" \
- -H "X-Title: My app" \
+curl https://tokens.bd/v1/chat/completions \
+ -H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
- "model": "provider/openrouter-model-slug",
+ "model": "deepseek/deepseek-v4.1-flash",
"messages": [{"role": "user", "content": "What does HTTP 429 mean?"}],
"max_tokens": 300
}'Keep Content-Type: application/json when you use curl. Without it Tokens does not read the body and answers 400 invalid_request.
If you pass the settings through the OPENAI_BASE_URL and OPENAI_API_KEY variables, which the official OpenAI Python and Node.js SDKs read (checked in the SDK sources, October 2026), change the variables and the model id and nothing else.
Model ids look the same and are not#
OpenRouter slugs and Tokens aliases both use provider/model. They are separate lists. A slug that works on OpenRouter can be unknown on Tokens, or can name a different version, and Tokens aliases do not always match the provider's own id. Do not rewrite ids by pattern.
- List what your key can call:
curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY". Eachdata[].idis valid for that key. - Prices, context windows and capabilities are on /models. The API list has ids only. See Choosing a model.
- Keep a table from your old ids to the new ones in configuration, so the mapping is not scattered in code.
The model value must match an alias exactly. An unknown id returns 404 model_not_found.
OpenRouter features, one by one#
This is what Tokens does with each OpenRouter feature. It comes from the gateway code, not from guesses.
Model suffixes. OpenRouter documents :free, :nitro, :floor, :exacto (and the deprecated :online, :thinking, :extended) as part of the model id. Tokens looks up the whole string as the alias, so some/model:nitro returns 404 model_not_found. Remove the suffix. There is no equivalent: Tokens has no suffixes and no per-request provider sorting. A cheaper or faster model is a different model id.
Router models. openrouter/auto and other OpenRouter router ids are not in the Tokens catalog (404 model_not_found). Pick a fixed model.
Provider routing (provider). Tokens does not read this field. It does not reject it either: for a model whose provider speaks the OpenAI protocol (the usual case), the rest of the request body is passed to the provider as sent. Whether the provider ignores the field or answers 400 is up to the provider, and Tokens did not check any provider's behaviour. Remove it. Which provider serves a model is Tokens' choice, not yours, and the model field of the response is always the id you asked for.
Fallbacks (models, route). Not implemented. The same rule as above: forwarded, not acted on, so the field does nothing. Tokens does fail over, but only between sources of the same model: on 429, 502, 503, 504 or a dropped connection, and before any byte of the answer has reached you, it retries on another source of that model. It never switches to a different model. If you want a model fallback, do it in your code:
import os
import openai
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], max_retries=0)
MODELS = ["deepseek/deepseek-v4.1-flash", "your-second-model-id"] # ids from GET /v1/models
def ask(messages):
last = None
for model in MODELS:
try:
return client.chat.completions.create(model=model, messages=messages, max_tokens=512)
except openai.APIStatusError as e:
if e.status_code not in (429, 500, 502, 503, 504) or e.code == "window_exhausted":
raise
last = e
raise lastwindow_exhausted is excluded because it is an account-wide plan window: another model will hit it too.
Prompt transforms and plugins (transforms, plugins). OpenRouter lets you compress long prompts, parse files, search the web and repair responses through these. Tokens does none of it. The fields are forwarded to the provider as unknown body fields, like provider. Trim the prompt yourself: Tokens does not shorten a prompt that is over the model's context window, so the provider decides what happens to it.
reasoning and usage fields. Forwarded like the others. Where a model supports reasoning, the provider decides what it does with them. The usage: {"include": true} request field is not needed: OpenRouter's documentation calls it deprecated and returns usage anyway, and Tokens' answers always carry the usage object the provider sent.
Attribution headers. HTTP-Referer, X-Title and the X-OpenRouter-* headers only matter on OpenRouter, for app pages and rankings. Tokens forwards a short allow-list of request headers to providers (content-type, accept, openai-beta and anthropic-version, among a few) and drops the rest without an error. Sending them is harmless and has no effect. Remove them so the code does not suggest otherwise.
If you are not sure a field is safe. Test it against the model you plan to use, on the test key described below. A request that works is the only proof that a provider accepts a field.
When a model is served only by a provider that speaks Anthropic's Messages protocol, Tokens translates your chat completions request instead of forwarding it. Only these fields are carried: messages (text, images, tool calls and tool results), max_tokens or max_completion_tokens, temperature, top_p, stop, stream, tools, tool_choice, parallel_tool_calls and user. Everything else is dropped, including n, response_format, seed, logprobs and the penalties. temperature is capped at 1, and max_tokens defaults to 4096 when you send none.
Response and usage differences#
modelin the response is always the id you asked for, whichever provider answered. OpenRouter'smodelshows the model it routed to, so code that logged it to see the routing outcome gets nothing new.- Cost fields. OpenRouter adds
cost,cost_detailsandnative_finish_reasonto the response. Tokens does not add anything to the provider's body. The cost of each request is in your usage dashboard and the CSV export, priced at the model's catalog price. If you readusage.costto bill your own customers, switch to the export, or compute the cost from the token counts and the price on /models. - Usage lookups. There is no
GET /generation?id=orGET /key.GET /v1/tokens/usagereturns plan windows, wallet balance and the calling key's cap (Models and usage). The response headerx-tokens-request-idis the id to log and to quote to support. - Streaming. The usage chunk at the end of a stream appears when you send
stream_options: {"include_usage": true}. See Streaming.
Differences that can bite#
Errors#
OpenRouter errors are {"error": {"code": 429, "message": "...", "metadata": {...}}} with a numeric code. Tokens errors are {"error": {"message", "type", "code", "param", "request_id"}} with a string code such as rate_limited. Code that reads error.code as a number, or reads error.metadata, needs a change: use the HTTP status for the class of error and error.code for the cause. The full table is in errors.
| Situation | OpenRouter | Tokens |
|---|---|---|
| Out of credit | 402 | 402 insufficient_credits, no_funding or outstanding_debt |
| Rate limited | 429 | 429 rate_limited, concurrency_limit, window_exhausted or model_limit_reached |
| Key limit you set | Key credit limit: 402 | 403 monthly_spend_cap_exceeded (see API keys) |
| Timeout | 408 | 504 upstream_timeout |
| Model down | 502 | 502 upstream_unreachable or the provider's 5xx as upstream_error |
| No provider available | 503 | 503 no_upstream_available |
OpenRouter documents that a failure after streaming starts arrives as an error event in a 200 response. Keep any in-body error check you added for that.
The provider's own error message is replaced by a generic one, and metadata.provider_* details do not exist.
Rate limits#
OpenRouter limits free models to 20 requests per minute and puts no platform cap on paid models. Tokens limits requests per minute per account (60 by default, or your plan's value) and concurrent requests per account (10 with a plan, 3 without). An agent or batch job that ran with high parallelism on OpenRouter can hit concurrency_limit; cap its parallelism. More keys do not raise the limits.
- A 429 carries
Retry-Afterin seconds. There are noX-RateLimit-*headers. PollGET /v1/tokens/usageto see a plan window. window_exhaustedandmodel_limit_reachedcan mean hours or days. Do not retry them in a loop.- See rate limits for the numbers and a backoff example.
Other differences#
- No
/api/v1path. Tokens' path is/v1:https://tokens.bd/v1. - Unsupported endpoints. Images, audio, files, batches, assistants, fine-tuning and moderations return 404
unsupported_endpoint. Embeddings work only for embedding models. Supported:/v1/chat/completions,/v1/responses,/v1/completions(legacy),/v1/embeddings,/v1/models,/v1/messagesand/v1/messages/count_tokens. - Parameters depend on the model. Besides
modelandn(1 to 4), Tokens does not validate the body. Tool calling,response_format, vision and reasoning depend on the model. The gateway renamesmax_tokenstomax_completion_tokensand removestemperatureandtop_punless they equal 1 for OpenAI o-series and GPT-5 and later models on chat completions. - Output reservation. Tokens reserves the worst-case cost of a request, using
max_tokensor 8,192 output tokens if you set none. A largemax_tokenson a low balance or a nearly full key cap can be refused or lowered. Set it to what you need. - No CORS. Calls from a browser fail. Call Tokens from a server.
- Body size. Up to 10 MB.
- Privacy. Tokens stores usage metadata, not prompt content. The provider that serves a model sees the prompt and its policy applies. See security and privacy.
- Payment. Tokens bills in USD or BDT from a plan or a wallet; there are no OpenRouter credits. See plans and wallet.
If you call OpenRouter's Anthropic-style Messages endpoint with an Anthropic SDK, read Migrate from Anthropic too. Tokens serves POST /v1/messages, with a base URL without /v1.
Test the switch safely#
- Create a second key at /dashboard/keys with a low monthly spend cap and an allowed-models list that holds only the models you are testing. Neither can be edited afterwards. Tokens counts the worst-case cost of each request against the cap.
- Read the base URL, key and model id from configuration, not from constants, so the switch is a config change.
- Run both for a while. Replay logged requests, or mirror a share of live traffic to Tokens and discard its answers. Compare the same prompts on both:
| Check | How |
|---|---|
| Quality | Run your own prompts or evals. The two models in a pair are rarely the same model. |
| Fields you stopped sending | Run once without provider, models, transforms and the headers. Did anything depend on them? |
| Tool calls | Valid JSON arguments for your schema, on the model you chose. |
finish_reason | More length than before means max_tokens is too low. |
| Cost per finished task | OpenRouter's usage.cost against the cost in usage. |
| Errors | Count them by error.code. |
- Ramp up with a feature flag or a percentage. Create the production key with the cap you want before you cut over.
Roll back#
- Keep your OpenRouter key and credits until Tokens has carried production traffic for a full billing cycle.
- Put the old base URL, key, model ids and headers back in the configuration and deploy or flip the flag.
- Revoke the Tokens key you no longer use at /dashboard/keys. Your wallet balance stays on your account; see the refund policy.
Because the id mapping lives in configuration, rolling back is the same change in reverse.
Where to go next#
- Chat Completions, Streaming and Tool calling.
- Errors and Rate limits.
- Migrate from OpenAI for the shared parts of an OpenAI-style switch.
Sources, checked October 2026: OpenRouter API overview, errors, limits, provider routing, model fallbacks, model variants, usage accounting and app attribution. Tokens behaviour is from the gateway code and the pages linked above. It has not been tested against a live OpenRouter account.