Skip to content

Migrate from OpenAI

Move an app that calls the OpenAI API to Tokens: the three settings that change, what stays the same, the endpoints and behaviours that differ, how to test the switch with a capped key and how to roll back.

On this page

If your app already calls the OpenAI API, moving it to Tokens takes three changes: the base URL, the API key and the model id. The request and response formats for chat completions, streaming and tool calling stay the same, so most code does not change. This page lists what does differ, so you find it in a test and not in production.

What changes and what stays the same#

SettingOpenAITokensWhere you set it
Base URLhttps://api.openai.com/v1 (the SDK default)https://tokens.bd/v1base_url (Python) or baseURL (Node.js), or the OPENAI_BASE_URL variable
API keysk-...tok_live_... from API keysapi_key or apiKey, or the OPENAI_API_KEY variable
Model idOpenAI's own idAn alias in provider/model form from /modelsThe model field of every request
Org and projectOpenAI-Organization, OpenAI-Project headersNot used. Remove them.Client options

The official OpenAI Python and Node.js SDKs read OPENAI_API_KEY and OPENAI_BASE_URL when you do not pass the values in code (checked in the SDK sources, October 2026). Setting both variables switches the endpoint without a code change. You still change the model id in code or config.

What stays the same:

  • The request and response JSON of POST /v1/chat/completions, including messages, tools, tool_choice, response_format, stream and the usage object. See Chat Completions.
  • Server-Sent Events streaming. The usage chunk at the end appears only when you send stream_options: {"include_usage": true}, as with OpenAI.
  • Authorization: Bearer <key> authentication.
  • The Responses API at POST /v1/responses (see below).
  • The OpenAI SDK classes and error types. An error status raises the same exception it would with OpenAI.

Before and after#

 import os
 from openai import OpenAI

-client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
+client = OpenAI(
+    base_url="https://tokens.bd/v1",
+    api_key=os.environ["TOKENS_API_KEY"],
+)

 resp = client.chat.completions.create(
-    model="your-openai-model",
+    model="deepseek/deepseek-v4.1-flash",
     messages=[{"role": "user", "content": "What does HTTP 429 mean?"}],
     max_tokens=300,
 )
 print(resp.choices[0].message.content)

Keep the Content-Type: application/json header when you use curl or a bare HTTP client. Without it Tokens does not read the body and answers 400 invalid_request ("must specify a 'model' field"). The SDKs set it for you.

Choose the model id#

Do not translate OpenAI's id by hand. Tokens ids are aliases that follow provider/model, and they do not always match the provider's own id. List what your key can call, then copy the id:

bash
curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY"

The list is filtered for the key: a key with an allowed-models list, or an account without a plan or wallet balance, sees fewer models. Prices, context windows and capabilities are on /models, not in the API response. Choosing a model helps you pick one. A different model gives different answers, so the switch is also a model change. Test your prompts, not just the connection.

Endpoints#

OpenAI endpointOn Tokens
POST /v1/chat/completionsSupported.
POST /v1/responsesSupported where the provider behind the model implements it. See Responses API.
POST /v1/completions (legacy)Supported where the provider implements it. Many chat models do not.
POST /v1/embeddingsOnly for catalog models that are embedding models.
GET /v1/modelsSupported, filtered for the calling key.
GET /v1/models/{id} (models.retrieve)Not supported: 404 unsupported_endpoint. Call the list and filter it.
Images, audio, files, uploads, batches, fine-tuning, moderationsNot supported: 404 unsupported_endpoint.
Assistants, threads, runsNot supported: 404 unsupported_endpoint. OpenAI retired the Assistants API on 26 August 2026 and points to the Responses API.
Realtime APINot supported.

Any path outside the supported list returns 404 with code unsupported_endpoint. If your app uses one of these endpoints, keep that part of the app on OpenAI and move only the text calls. Use two clients, each with its own base URL and key.

Tokens adds two endpoints OpenAI does not have: GET /v1/tokens/usage (plan windows, wallet balance and key limits, see Models and usage) and the Anthropic-style POST /v1/messages (Messages).

The Responses API#

If your code uses client.responses.create, it works against the Tokens base URL with the same change. Two cautions from the Responses API page:

  • Tokens does not store prompts or responses. Do not rely on store: true or previous_response_id to keep conversation state. Send the full conversation in input on every call.
  • Hosted tools such as web search and file search are provider features. Do not assume they work through the gateway. Test them first.

A model that is served only by a provider that speaks Anthropic's Messages protocol answers chat completions (Tokens translates the request) but not /v1/responses, /v1/completions or /v1/embeddings. Those calls fail with 400 endpoint_not_supported_for_model. Use chat completions for such a model.

Differences that can bite#

Rate limits are per account, and lower by default#

OpenAI limits requests and tokens per minute by organization and project, and returns x-ratelimit-* headers. Tokens limits requests per minute per account (60 by default, or your plan's value) and concurrent requests per account (10 with a plan, 3 without). The rate limits page lists no tokens-per-minute limit. Extra keys do not raise either limit, because both are per account.

  • A 429 carries Retry-After in seconds. There are no x-ratelimit-* headers, so code that reads them gets nothing. Poll GET /v1/tokens/usage to see what is left in a plan window.
  • A parallel job that was fine on OpenAI can hit concurrency_limit. Cap the parallelism on your side, for example with a semaphore.
  • window_exhausted and model_limit_reached can carry a Retry-After of hours or days. Do not retry those in a loop. The SDKs retry 429 twice by default; pass max_retries=0 (Python) or maxRetries: 0 (Node.js) if you handle retries yourself. The backoff example in rate limits does this.

Errors have the same shape and different codes#

Gateway errors use OpenAI's JSON shape: error.message, error.type, error.code, error.param and an added error.request_id. Branch on error.code. The ones that differ from what OpenAI code usually expects:

SituationOpenAITokens
Out of credit429 with a quota or spend-limit code402 insufficient_credits, no_funding or outstanding_debt. type is insufficient_quota.
Per-minute limit429 with x-ratelimit-* headers429 rate_limited with Retry-After
Provider failure500, or 503 server_is_overloaded502 upstream_unreachable, 504 upstream_timeout, or the provider's 5xx as upstream_error

Tokens also has codes that have no OpenAI counterpart: model_not_found (404, the id is not in the catalog), tier_permission_denied (403, your plan does not include the model), and the key limits model_not_allowed_on_key and monthly_spend_cap_exceeded (403).

If your code treats every 429 as "retry later" and every quota problem as a 429, it will retry 402s, which never succeed. The full list is in errors. Retry 429 (except window_exhausted and model_limit_reached) and 5xx with backoff. Do not retry 400, 401, 402, 403 or 404.

Error messages from the provider behind a model are replaced with a generic one, for example "The request was rejected by the upstream provider." A 400 from a model that does not accept a parameter therefore does not tell you which parameter. Check the request against that model's page in /models.

Request ids are different#

OpenAI returns x-request-id. Tokens returns x-tokens-request-id on every response, and echoes your own x-request-id in x-request-id if you send one. Keep x-tokens-request-id in your logs. Support searches by it. Tokens passes on only a short list of provider response headers (content-type, cache-control and retry-after), so provider-specific headers such as rate-limit headers do not reach you.

Parameters depend on the model#

Tokens does not validate the body beyond model and n (1 to 4). tools, response_format, reasoning_effort, seed, logprobs and similar fields go to the provider behind the model. A model that does not support one ignores it or answers 400. Features you took for granted on OpenAI's models, such as strict structured outputs or image input, depend on the model you choose, so check its page.

One rewrite does happen. For OpenAI-style reasoning models (the o-series and GPT-5 and later) on chat completions, the gateway renames max_tokens to max_completion_tokens and removes temperature and top_p unless they equal 1, because those models refuse them.

Output limits and the credit reservation#

Before it forwards a request, Tokens reserves the worst-case cost, using your max_tokens (or 8,192 output tokens if you set none). With a low balance or a key near its cap, a large max_tokens can be refused or, on a low balance, lowered to what you can afford (never below 16). Set max_tokens to what you need. You pay for the tokens used, not the reservation. See Chat Completions.

Browsers, size and privacy#

  • No CORS headers: calls from a browser fail. Call Tokens from a server. See authentication.
  • The request body can be up to 10 MB (413 request_entity_too_large). Large base64 images count toward it.
  • Tokens adds one network hop, so latency to the first token is not lower than calling the provider directly.
  • Prompts reach the provider that serves the model, and that provider's data policy applies. Tokens itself stores usage metadata, not prompt content. See security and privacy.

Billing#

You pay Tokens, in USD or BDT, from a plan or a wallet, at the catalog price of each model. Your OpenAI invoice stops for the traffic you move. See plans and wallet.

Test the switch safely#

  1. Create a second key at /dashboard/keys. Name it for the test, set a low monthly spend cap (for example a few dollars) and an allowed-models list of the one or two models you will try. The cap and the list cannot be edited later, so create a new key to change them. Tokens counts the worst-case cost of each request against the cap, so a very large max_tokens can be refused near the cap.
  2. Switch by configuration. Read the base URL, key and model id from environment variables or a config file, so a deploy does not need a code change to move between providers.
  3. Run both for a while. Send the same prompts to OpenAI and Tokens (a replay of logged requests, or a mirror of live traffic whose Tokens answer you discard) and compare. A small script is enough:
python
import os
import time

from openai import OpenAI

openai_client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
tokens_client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_TEST_KEY"])

CANDIDATES = [
    ("openai", openai_client, "your-openai-model"),
    ("tokens", tokens_client, "deepseek/deepseek-v4.1-flash"),
]

prompt = [{"role": "user", "content": "Write a Python function that parses an ISO 8601 date."}]

for name, client, model in CANDIDATES:
    start = time.perf_counter()
    raw = client.chat.completions.with_raw_response.create(
        model=model, messages=prompt, max_tokens=400
    )
    elapsed = time.perf_counter() - start
    resp = raw.parse()
    print(name, resp.choices[0].finish_reason, resp.usage.total_tokens, f"{elapsed:.2f}s")
    print("  request id:", raw.headers.get("x-tokens-request-id") or raw.headers.get("x-request-id"))
  1. Compare what matters to your app, not only the text:
CheckHow
Answer qualityRun your own test prompts or evals on both. Models differ.
Tool callsAre the arguments valid JSON for your schema? Does the model call the right tool?
finish_reasonMore length than before means max_tokens is too low for this model.
Token counts and costToken counts differ by model. Compare cost per finished task in usage.
LatencyTime to first token with stream: true, from where your app runs.
ErrorsCount them by error.code. A 402 or 429 concurrency_limit shows a sizing problem.
  1. Ramp up. Move a small share of traffic (a feature flag or a percentage), watch for a day or two, then increase it. Replace the test key with a production key that has the cap you want. Create it before you cut over, because rotating or revoking takes effect at once.

Roll back#

Rolling back is the reverse of the switch, if you kept the way back open:

  1. Keep your OpenAI key and its billing active until Tokens has carried production traffic for a full billing cycle.
  2. Point the base URL, key and model id back through the same configuration, and redeploy or flip the flag. If you used OPENAI_BASE_URL, unset it.
  3. Revoke or rotate the Tokens key you no longer use at /dashboard/keys.

Your wallet balance and plan on Tokens stay on your account. For refunds see the refund policy. Nothing needs to be exported, because Tokens does not keep prompt content.

Where to go next#

Sources, checked October 2026: OpenAI API reference overview, error codes, rate limits, Assistants migration, and the openai-python and openai-node sources. Tokens behaviour is from the gateway code and the pages linked above. It has not been tested against a live OpenAI account.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.