# Migrate from OpenAI

> Move an app that calls the OpenAI API to Tokens: the three settings that change, what stays the same, the endpoints and behaviours that differ, how to test the switch with a capped key and how to roll back.

If your app already calls the OpenAI API, moving it to Tokens takes three changes: the base URL, the API key and the model id. The request and response formats for chat completions, streaming and tool calling stay the same, so most code does not change. This page lists what does differ, so you find it in a test and not in production.

## What changes and what stays the same

| Setting      | OpenAI                                       | Tokens                                           | Where you set it                                                         |
| ------------ | -------------------------------------------- | ------------------------------------------------ | ------------------------------------------------------------------------ |
| Base URL     | `https://api.openai.com/v1` (the SDK default) | `https://tokens.bd/v1`                            | `base_url` (Python) or `baseURL` (Node.js), or the `OPENAI_BASE_URL` variable |
| API key      | `sk-...`                                     | `tok_live_...` from [API keys](/docs/api-keys)   | `api_key` or `apiKey`, or the `OPENAI_API_KEY` variable                  |
| Model id     | OpenAI's own id                              | An alias in `provider/model` form from `/models` | The `model` field of every request                                       |
| Org and project | `OpenAI-Organization`, `OpenAI-Project` headers | Not used. Remove them.                       | Client options                                                           |

The official OpenAI Python and Node.js SDKs read `OPENAI_API_KEY` and `OPENAI_BASE_URL` when you do not pass the values in code (checked in the SDK sources, October 2026). Setting both variables switches the endpoint without a code change. You still change the model id in code or config.

What stays the same:

- The request and response JSON of `POST /v1/chat/completions`, including `messages`, `tools`, `tool_choice`, `response_format`, `stream` and the `usage` object. See [Chat Completions](/docs/chat-completions).
- Server-Sent Events streaming. The usage chunk at the end appears only when you send `stream_options: {"include_usage": true}`, as with OpenAI.
- `Authorization: Bearer <key>` authentication.
- The Responses API at `POST /v1/responses` (see below).
- The OpenAI SDK classes and error types. An error status raises the same exception it would with OpenAI.

## Before and after

:::code-tabs

```diff title="Python"
 import os
 from openai import OpenAI

-client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
+client = OpenAI(
+    base_url="https://tokens.bd/v1",
+    api_key=os.environ["TOKENS_API_KEY"],
+)

 resp = client.chat.completions.create(
-    model="your-openai-model",
+    model="deepseek/deepseek-v4.1-flash",
     messages=[{"role": "user", "content": "What does HTTP 429 mean?"}],
     max_tokens=300,
 )
 print(resp.choices[0].message.content)
```

```diff title="Node.js"
 import OpenAI from "openai";

-const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
+const client = new OpenAI({
+  baseURL: "https://tokens.bd/v1",
+  apiKey: process.env.TOKENS_API_KEY,
+});

 const resp = await client.chat.completions.create({
-  model: "your-openai-model",
+  model: "deepseek/deepseek-v4.1-flash",
   messages: [{ role: "user", content: "What does HTTP 429 mean?" }],
   max_tokens: 300,
 });
 console.log(resp.choices[0].message.content);
```

```diff title="curl"
-curl https://api.openai.com/v1/chat/completions \
-  -H "Authorization: Bearer $OPENAI_API_KEY" \
+curl https://tokens.bd/v1/chat/completions \
+  -H "Authorization: Bearer $TOKENS_API_KEY" \
   -H "Content-Type: application/json" \
   -d '{
-    "model": "your-openai-model",
+    "model": "deepseek/deepseek-v4.1-flash",
     "messages": [{"role": "user", "content": "What does HTTP 429 mean?"}],
     "max_tokens": 300
   }'
```

:::

Keep the `Content-Type: application/json` header when you use curl or a bare HTTP client. Without it Tokens does not read the body and answers 400 `invalid_request` ("must specify a 'model' field"). The SDKs set it for you.

### Choose the model id

Do not translate OpenAI's id by hand. Tokens ids are aliases that follow `provider/model`, and they do not always match the provider's own id. List what your key can call, then copy the id:

```bash
curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY"
```

The list is filtered for the key: a key with an allowed-models list, or an account without a plan or wallet balance, sees fewer models. Prices, context windows and capabilities are on [/models](/models), not in the API response. [Choosing a model](/docs/choosing-a-model) helps you pick one. A different model gives different answers, so the switch is also a model change. Test your prompts, not just the connection.

## Endpoints

| OpenAI endpoint                                                                    | On Tokens                                                                                                                       |
| ---------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `POST /v1/chat/completions`                                                        | Supported.                                                                                                                      |
| `POST /v1/responses`                                                               | Supported where the provider behind the model implements it. See [Responses API](/docs/responses).                              |
| `POST /v1/completions` (legacy)                                                    | Supported where the provider implements it. Many chat models do not.                                                            |
| `POST /v1/embeddings`                                                              | Only for catalog models that are embedding models.                                                                              |
| `GET /v1/models`                                                                   | Supported, filtered for the calling key.                                                                                        |
| `GET /v1/models/{id}` (`models.retrieve`)                                          | Not supported: 404 `unsupported_endpoint`. Call the list and filter it.                                                         |
| Images, audio, files, uploads, batches, fine-tuning, moderations                   | Not supported: 404 `unsupported_endpoint`.                                                                                      |
| Assistants, threads, runs                                                          | Not supported: 404 `unsupported_endpoint`. OpenAI retired the Assistants API on 26 August 2026 and points to the Responses API. |
| Realtime API                                                                       | Not supported.                                                                                                                  |

Any path outside the supported list returns 404 with code `unsupported_endpoint`. If your app uses one of these endpoints, keep that part of the app on OpenAI and move only the text calls. Use two clients, each with its own base URL and key.

Tokens adds two endpoints OpenAI does not have: `GET /v1/tokens/usage` (plan windows, wallet balance and key limits, see [Models and usage](/docs/models-and-usage)) and the Anthropic-style `POST /v1/messages` ([Messages](/docs/messages)).

### The Responses API

If your code uses `client.responses.create`, it works against the Tokens base URL with the same change. Two cautions from the [Responses API page](/docs/responses):

- Tokens does not store prompts or responses. Do not rely on `store: true` or `previous_response_id` to keep conversation state. Send the full conversation in `input` on every call.
- Hosted tools such as web search and file search are provider features. Do not assume they work through the gateway. Test them first.

A model that is served only by a provider that speaks Anthropic's Messages protocol answers chat completions (Tokens translates the request) but not `/v1/responses`, `/v1/completions` or `/v1/embeddings`. Those calls fail with 400 `endpoint_not_supported_for_model`. Use chat completions for such a model.

## Differences that can bite

### Rate limits are per account, and lower by default

OpenAI limits requests and tokens per minute by organization and project, and returns `x-ratelimit-*` headers. Tokens limits requests per minute per account (60 by default, or your plan's value) and concurrent requests per account (10 with a plan, 3 without). The [rate limits](/docs/rate-limits) page lists no tokens-per-minute limit. Extra keys do not raise either limit, because both are per account.

- A 429 carries `Retry-After` in seconds. There are no `x-ratelimit-*` headers, so code that reads them gets nothing. Poll `GET /v1/tokens/usage` to see what is left in a plan window.
- A parallel job that was fine on OpenAI can hit `concurrency_limit`. Cap the parallelism on your side, for example with a semaphore.
- `window_exhausted` and `model_limit_reached` can carry a `Retry-After` of hours or days. Do not retry those in a loop. The SDKs retry 429 twice by default; pass `max_retries=0` (Python) or `maxRetries: 0` (Node.js) if you handle retries yourself. The backoff example in [rate limits](/docs/rate-limits) does this.

### Errors have the same shape and different codes

Gateway errors use OpenAI's JSON shape: `error.message`, `error.type`, `error.code`, `error.param` and an added `error.request_id`. Branch on `error.code`. The ones that differ from what OpenAI code usually expects:

| Situation        | OpenAI                                      | Tokens                                                                                          |
| ---------------- | ------------------------------------------- | ----------------------------------------------------------------------------------------------- |
| Out of credit    | 429 with a quota or spend-limit code        | 402 `insufficient_credits`, `no_funding` or `outstanding_debt`. `type` is `insufficient_quota`. |
| Per-minute limit | 429 with `x-ratelimit-*` headers            | 429 `rate_limited` with `Retry-After`                                                           |
| Provider failure | 500, or 503 `server_is_overloaded`          | 502 `upstream_unreachable`, 504 `upstream_timeout`, or the provider's 5xx as `upstream_error`   |

Tokens also has codes that have no OpenAI counterpart: `model_not_found` (404, the id is not in the catalog), `tier_permission_denied` (403, your plan does not include the model), and the key limits `model_not_allowed_on_key` and `monthly_spend_cap_exceeded` (403).

If your code treats every 429 as "retry later" and every quota problem as a 429, it will retry 402s, which never succeed. The full list is in [errors](/docs/errors). Retry 429 (except `window_exhausted` and `model_limit_reached`) and 5xx with backoff. Do not retry 400, 401, 402, 403 or 404.

Error messages from the provider behind a model are replaced with a generic one, for example "The request was rejected by the upstream provider." A 400 from a model that does not accept a parameter therefore does not tell you which parameter. Check the request against that model's page in [/models](/models).

### Request ids are different

OpenAI returns `x-request-id`. Tokens returns `x-tokens-request-id` on every response, and echoes your own `x-request-id` in `x-request-id` if you send one. Keep `x-tokens-request-id` in your logs. [Support](/docs/support) searches by it. Tokens passes on only a short list of provider response headers (`content-type`, `cache-control` and `retry-after`), so provider-specific headers such as rate-limit headers do not reach you.

### Parameters depend on the model

Tokens does not validate the body beyond `model` and `n` (1 to 4). `tools`, `response_format`, `reasoning_effort`, `seed`, `logprobs` and similar fields go to the provider behind the model. A model that does not support one ignores it or answers 400. Features you took for granted on OpenAI's models, such as strict structured outputs or image input, depend on the model you choose, so check its page.

One rewrite does happen. For OpenAI-style reasoning models (the o-series and GPT-5 and later) on chat completions, the gateway renames `max_tokens` to `max_completion_tokens` and removes `temperature` and `top_p` unless they equal 1, because those models refuse them.

### Output limits and the credit reservation

Before it forwards a request, Tokens reserves the worst-case cost, using your `max_tokens` (or 8,192 output tokens if you set none). With a low balance or a key near its cap, a large `max_tokens` can be refused or, on a low balance, lowered to what you can afford (never below 16). Set `max_tokens` to what you need. You pay for the tokens used, not the reservation. See [Chat Completions](/docs/chat-completions).

### Browsers, size and privacy

- No CORS headers: calls from a browser fail. Call Tokens from a server. See [authentication](/docs/authentication).
- The request body can be up to 10 MB (413 `request_entity_too_large`). Large base64 images count toward it.
- Tokens adds one network hop, so latency to the first token is not lower than calling the provider directly.
- Prompts reach the provider that serves the model, and that provider's data policy applies. Tokens itself stores usage metadata, not prompt content. See [security and privacy](/docs/security-and-privacy).

### Billing

You pay Tokens, in USD or BDT, from a plan or a wallet, at the catalog price of each model. Your OpenAI invoice stops for the traffic you move. See [plans and wallet](/docs/plans-and-wallet).

## Test the switch safely

1. **Create a second key** at [/dashboard/keys](/dashboard/keys). Name it for the test, set a **low monthly spend cap** (for example a few dollars) and an **allowed-models list** of the one or two models you will try. The cap and the list cannot be edited later, so create a new key to change them. Tokens counts the worst-case cost of each request against the cap, so a very large `max_tokens` can be refused near the cap.
2. **Switch by configuration.** Read the base URL, key and model id from environment variables or a config file, so a deploy does not need a code change to move between providers.
3. **Run both for a while.** Send the same prompts to OpenAI and Tokens (a replay of logged requests, or a mirror of live traffic whose Tokens answer you discard) and compare. A small script is enough:

```python
import os
import time

from openai import OpenAI

openai_client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
tokens_client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_TEST_KEY"])

CANDIDATES = [
    ("openai", openai_client, "your-openai-model"),
    ("tokens", tokens_client, "deepseek/deepseek-v4.1-flash"),
]

prompt = [{"role": "user", "content": "Write a Python function that parses an ISO 8601 date."}]

for name, client, model in CANDIDATES:
    start = time.perf_counter()
    raw = client.chat.completions.with_raw_response.create(
        model=model, messages=prompt, max_tokens=400
    )
    elapsed = time.perf_counter() - start
    resp = raw.parse()
    print(name, resp.choices[0].finish_reason, resp.usage.total_tokens, f"{elapsed:.2f}s")
    print("  request id:", raw.headers.get("x-tokens-request-id") or raw.headers.get("x-request-id"))
```

4. **Compare what matters to your app**, not only the text:

| Check                    | How                                                                                                  |
| ------------------------ | ---------------------------------------------------------------------------------------------------- |
| Answer quality           | Run your own test prompts or evals on both. Models differ.                                           |
| Tool calls               | Are the arguments valid JSON for your schema? Does the model call the right tool?                    |
| `finish_reason`          | More `length` than before means `max_tokens` is too low for this model.                              |
| Token counts and cost    | Token counts differ by model. Compare cost per finished task in [usage](/docs/usage-and-alerts).      |
| Latency                  | Time to first token with `stream: true`, from where your app runs.                                   |
| Errors                   | Count them by `error.code`. A 402 or 429 `concurrency_limit` shows a sizing problem.                 |

5. **Ramp up.** Move a small share of traffic (a feature flag or a percentage), watch for a day or two, then increase it. Replace the test key with a production key that has the cap you want. Create it before you cut over, because [rotating or revoking](/docs/api-keys) takes effect at once.

## Roll back

Rolling back is the reverse of the switch, if you kept the way back open:

1. Keep your OpenAI key and its billing active until Tokens has carried production traffic for a full billing cycle.
2. Point the base URL, key and model id back through the same configuration, and redeploy or flip the flag. If you used `OPENAI_BASE_URL`, unset it.
3. Revoke or rotate the Tokens key you no longer use at [/dashboard/keys](/dashboard/keys).

Your wallet balance and plan on Tokens stay on your account. For refunds see the [refund policy](/refund-policy). Nothing needs to be exported, because Tokens does not keep prompt content.

## Where to go next

- [Chat Completions](/docs/chat-completions), [Responses API](/docs/responses) and [Streaming](/docs/streaming) for the request formats.
- [Python](/docs/python) and [Node.js](/docs/nodejs) for SDK setup.
- [Errors](/docs/errors) and [Rate limits](/docs/rate-limits) for the full code list and a backoff example.
- [Migrate from OpenRouter](/docs/migrate-from-openrouter) and [Migrate from Anthropic](/docs/migrate-from-anthropic).

Sources, checked October 2026: OpenAI [API reference overview](https://developers.openai.com/api/reference/overview), [error codes](https://developers.openai.com/api/docs/guides/error-codes), [rate limits](https://developers.openai.com/api/docs/guides/rate-limits), [Assistants migration](https://developers.openai.com/api/docs/assistants/migration), and the [openai-python](https://github.com/openai/openai-python) and [openai-node](https://github.com/openai/openai-node) sources. Tokens behaviour is from the gateway code and the pages linked above. It has not been tested against a live OpenAI account.

---
Page: https://tokens.bd/docs/migrate-from-openai
