TL;DR: The OpenAI SDK sends every request to a configurable URL. Point
base_urlathttps://tokens.bd/v1and your existing Python or Node.js code talks to 62 models from 10 providers instead of OpenAI alone — same SDK, same code shape. The model is just a string: swapopenai/gpt-6-luna($0.11/$0.55 per 1M) formoonshotai/kimi-k3($3.30/$16.50) and the next request routes there. Prices checked 8 October 2026 on the Tokens catalog.
Every line of code written against the OpenAI SDK in the last four years has a hidden superpower: it was never really bound to OpenAI. The SDK is a thin HTTP client for the chat-completions format, and the one address it talks to — https://api.openai.com/v1 — is just a default you can override. Change that one line, and every tutorial, every agent framework, every snippet on Stack Overflow becomes portable across models and providers.
Why base_url exists#
The OpenAI SDK's job is to serialize your messages array into JSON, POST it to an endpoint, and parse the response. It doesn't care whose server answers, as long as the server speaks the chat-completions shape: POST /v1/chat/completions with a model string and a messages array, returning choices[].message.content.
AI gateways exploit exactly that. Tokens exposes an OpenAI-compatible base URL at https://tokens.bd/v1 (a separate Anthropic-compatible base at https://tokens.bd exists for Claude Code-style tools — OpenAI-style tools add /v1, Anthropic-style ones take the base without it, per the docs). One API key authenticates everything behind it: 62 models from 10 providers, each called by a stable provider/model id. Your code never learns a new API.

The one-line change#
Get a key from your Tokens dashboard, then:
Python
from openai import OpenAI
client = OpenAI(
base_url="https://tokens.bd/v1", # <- the whole change
api_key="tk-...",
)
resp = client.chat.completions.create(
model="anthropic/claude-sonnet-5.5",
messages=[{"role": "user", "content": "Explain recursion like I'm a junior dev."}],
)
print(resp.choices[0].message.content)Node.js
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://tokens.bd/v1", // <- the whole change
apiKey: process.env.TOKENS_API_KEY,
});
const resp = await client.chat.completions.create({
model: "openai/gpt-6-luna",
messages: [{ role: "user", content: "Summarize this diff." }],
});
console.log(resp.choices[0].message.content);Or zero code changes, with environment variables:
export OPENAI_BASE_URL="https://tokens.bd/v1"
export OPENAI_API_KEY="$TOKENS_API_KEY"Both official SDKs read OPENAI_BASE_URL / OPENAI_API_KEY automatically — so existing scripts and tools that read those variables just start routing through the gateway. (Caveat: if your shell also needs genuine OpenAI access in the same session, keep the Tokens key in a differently-named variable and pass it explicitly.)
cURL, for the terminal:
curl https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "moonshotai/kimi-k3",
"messages": [{"role": "user", "content": "Hello"}]
}'What you actually get: 62 models behind one string#
This is where the trick pays off. The model parameter becomes the only decision that matters. Catalog prices per million tokens, checked 8 October 2026:
| Model id | Provider | Input / 1M | Output / 1M | Best for |
|---|---|---|---|---|
openai/gpt-6-luna | OpenAI | $0.11 | $0.55 | High-volume, latency-sensitive work |
deepseek/deepseek-v4-flash-fast | DeepSeek | $0.31 | $0.62 | Cheap drafts and first passes |
xiaomi/mimo-v2.6-pro | Xiaomi | $0.48 | $0.96 | Budget reasoning |
openai/gpt-6-sol | OpenAI | $2.00 | $10.00 | General coding and agents |
anthropic/claude-sonnet-5.5 | Anthropic | $2.20 | $11.00 | Strong reasoning, long contexts |
moonshotai/kimi-k3 | Moonshot | $3.30 | $16.50 | Frontier-grade output |
x-ai/grok-4.7 | xAI | $2.20 | $6.60 | Fast frontier alternative |
google/gemini-3.8-flash | $1.65 | $8.25 | Google ecosystem tasks | |
inclusionai/ling-3.1-flash | InclusionAI | $0.00 | $0.00 | Free tier experiments |
The full list — with context windows and plan availability — is in the model catalog.
Switching models is one word#
This is the part that feels like cheating after years of per-provider SDKs. Because every model shares the same request shape, moving a workload from a $0.11-input model to a $3.30-input model is a single string change — no new import, no new client, no new key. That makes A/B testing models trivial: loop over a list of model ids, keep everything else constant, and compare outputs on your own benchmark before you commit spend.

Two practical patterns this unlocks:
- Tiered routing. Draft with
deepseek/deepseek-v4-flash-fastat $0.31/1M input, escalate toanthropic/claude-sonnet-5.5only when an objective check fails. The escalation logic we described in our frontier models guide needs no SDK changes — just a different string. - Cost-aware fallbacks. If your primary model degrades, your retry path just substitutes another id. No second integration to maintain.
What stays the same — and what doesn't#
Compatible: chat completions, message roles, temperature and sampling parameters, streaming (both SDKs support stream=True against the gateway), and JSON mode where the underlying model supports it. If your tool or agent framework speaks OpenAI's protocol, it plugs in.
Not the same: account-level endpoints (fine-tuning, organization billing, batch jobs) live at OpenAI and don't exist on the gateway — a gateway routes inference, not OpenAI's admin surface. Provider-specific features (Anthropic's prompt caching semantics, OpenAI's reasoning-effort dial on reasoning models) are exposed as documented on each model; behavior follows the underlying model, not the SDK you used to reach it. Our prompt caching and thinking budgets posts cover how those bill.
Four pitfalls, in order of how often they bite#
- Wrong base URL variant.
https://tokens.bd/v1for OpenAI-style tools;https://tokens.bd(no/v1) for Anthropic-style tools like Claude Code. The docs spell this out, but it's the number-one support question pattern — a 404 on/v1/v1/chat/completionsmeans you doubled the path. - Forgetting the provider prefix. The id is
anthropic/claude-sonnet-5.5, notclaude-sonnet-5.5. An unknown-model error almost always means the prefix is missing or misspelled. Copy ids from the catalog. - Hardcoding the key. One leaked key in a repo is how a $2,000 weekend happens. Keep keys in environment variables, and give each project its own key with a spend cap — Tokens supports per-key caps with
403 monthly_spend_cap_exceededpast the limit, plus usage alerts. See per-key spend caps. - Pricing by vibes. The price spread above is 30x from cheapest to priciest input. A model swap that "feels" equivalent can multiply a batch job's bill. Check the price table before you promote an experiment to production.
A note for developers in Bangladesh#
If you reached for a gateway partly for payments, the pairing is natural: OpenAI's API is card-and-USD-only, while Tokens takes invoice top-ups in taka. The details — what 1M tokens costs in ৳, the $5 prepaid minimum, credit expiry — are in our guide to using the OpenAI API from Bangladesh. How Tokens compares to other gateways on price is in the OpenRouter comparison.
OpenAI SDK base_url FAQ#
Does this really work with any OpenAI SDK version?
Any recent openai Python (v1.x) or Node.js SDK supports base_url/baseURL and the OPENAI_BASE_URL environment variable. Older pre-1.0 Python SDKs used different configuration; upgrade if you're on one.
Will my code's streaming and tool calls still work? Yes for chat completions through the gateway, including streaming. Fine-tuning, batches, and other OpenAI account endpoints don't exist on the gateway — those requests need genuine OpenAI access.
Why do model names have a provider prefix?
The catalog holds 62 models from 10 providers behind one API, so ids are namespaced for stability: anthropic/claude-sonnet-5.5, openai/gpt-6-luna, moonshotai/kimi-k3. The catalog is the authority on ids.
Can I point base_url at Tokens and still use OpenAI directly sometimes?
Yes — create two client instances, one with the default base and one with https://tokens.bd/v1, or toggle via environment variables per command. Don't mix one key across both; Tokens keys and OpenAI keys are not interchangeable.
Does switching models reset my prompt cache? Caches are per model, so yes — the first request to a new model pays full input price. That's usually still cheaper than running everything on the expensive model; see prompt caching explained.
Model ids, base URLs and prices checked 8 October 2026 against tokens.bd/models and tokens.bd/docs. Prices change; the catalog is the authority.
Sources: Tokens documentation · Tokens model catalog · OpenAI Python SDK (base_url)



