Skip to content

Python (OpenAI SDK)

Use the official openai Python package with Tokens: client setup, sync and async calls, streaming, tool calls, timeouts, retries and error handling.

Works withOpenAI Python SDK
On this page

The official openai Python package works with Tokens unchanged. You point it at https://tokens.bd/v1, pass your Tokens key, and use model IDs from the Tokens catalog. Everything else (streaming, tool calls, async) is the SDK you already know.

Install and configure the OpenAI Python SDK#

bash
pip install --upgrade openai
export TOKENS_API_KEY="tok_live_your_key"
hello.py
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
)

completion = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    messages=[
        {"role": "system", "content": "Answer in one short paragraph."},
        {"role": "user", "content": "When should I use a dataclass instead of a dict?"},
    ],
    max_tokens=400,
)

print(completion.choices[0].message.content)
print(completion.usage)

Use the exact model ID from /models or client.models.list(). IDs follow a provider/model pattern, and a typo returns 404 model_not_found.

Alternative: OPENAI_BASE_URL and OPENAI_API_KEY#

The SDK reads OPENAI_BASE_URL and OPENAI_API_KEY from the environment when you don't pass them, so existing code can switch to Tokens with no edits:

bash
export OPENAI_BASE_URL="https://tokens.bd/v1"
export OPENAI_API_KEY="$TOKENS_API_KEY"
python
from openai import OpenAI

client = OpenAI()  # picks up OPENAI_BASE_URL and OPENAI_API_KEY

This is convenient, but those variables are global. Every tool in that shell that reads them (Aider, some agents, other scripts) will also send traffic to Tokens. If you want that scoped, pass base_url and api_key explicitly as in the first example.

Stream responses#

python
stream = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Write a haiku about merge conflicts."}],
    stream=True,
    stream_options={"include_usage": True},
)

for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:
        print(f"\n\n{chunk.usage.prompt_tokens} in, {chunk.usage.completion_tokens} out")

The if chunk.choices guard matters. With include_usage on, the final chunk has an empty choices list and carries only usage. Without include_usage, streams carry no token counts at all. See Streaming.

Use the async client#

AsyncOpenAI takes the same arguments. Use it in FastAPI, aiohttp or anything else running an event loop.

python
import asyncio
import os
from openai import AsyncOpenAI

client = AsyncOpenAI(
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
)

async def main() -> None:
    stream = await client.chat.completions.create(
        model="deepseek/deepseek-v4.1-flash",
        messages=[{"role": "user", "content": "Name three uses for asyncio.Semaphore."}],
        stream=True,
    )
    async for chunk in stream:
        if chunk.choices and chunk.choices[0].delta.content:
            print(chunk.choices[0].delta.content, end="", flush=True)

asyncio.run(main())

If you fan out many requests with asyncio.gather, cap them with a semaphore. Tokens limits concurrent requests per account (the limit comes from your plan), and going over it returns 429 concurrency_limit. See Rate limits.

Tool calls#

Tool calling works on models that support it. Check the model's page in /models before you rely on it.

python
import json

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]

messages = [{"role": "user", "content": "Is it raining in Dhaka?"}]
first = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash", messages=messages, tools=tools
)
msg = first.choices[0].message

if msg.tool_calls:
    messages.append(msg)
    for call in msg.tool_calls:
        args = json.loads(call.function.arguments)
        result = {"city": args["city"], "condition": "light rain", "temp_c": 29}  # your real lookup here
        messages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)})
    final = client.chat.completions.create(
        model="deepseek/deepseek-v4.1-flash", messages=messages, tools=tools
    )
    print(final.choices[0].message.content)

The full request and response shapes are in Tool calling.

Configure timeouts and retries#

The SDK retries some failures (connection errors, 408, 409, 429 and 5xx) twice by default with backoff. Its default timeout is 10 minutes. Both can be set per client or per call:

python
import os
import httpx
from openai import OpenAI

client = OpenAI(
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    timeout=httpx.Timeout(120.0, connect=10.0),
    max_retries=3,
)

# Override for one call
client.with_options(timeout=30.0, max_retries=0).chat.completions.create(...)

Some notes specific to Tokens:

  • The gateway already fails over to another upstream source on 429, 502, 503, 504 and connection errors before it replies, so a 5xx you see means that failover didn't help. A couple of SDK retries is plenty.
  • Long generations are fine. The gateway waits up to 600 seconds for response headers. For long outputs, stream so you aren't holding an idle connection.
  • Retrying 429 window_exhausted won't help until the window resets. Retry-After gives the seconds until it does, and that can be hours.

Handle Tokens error codes#

Errors use the OpenAI shape, so openai raises its usual exception classes. The useful part for Tokens is error.code.

python
import openai

try:
    client.chat.completions.create(
        model="deepseek/deepseek-v4.1-flash",
        messages=[{"role": "user", "content": "hi"}],
    )
except openai.APIStatusError as e:
    request_id = e.response.headers.get("x-tokens-request-id")
    print(e.status_code, e.code, e.message, request_id)
    if e.code == "insufficient_credits":
        print("Top up at https://tokens.bd/dashboard/billing")
except openai.APIConnectionError as e:
    print("Network problem:", e)

Log x-tokens-request-id with every failure. Support can trace a request from that ID. The codes you're likely to meet are invalid_api_key (401), model_not_allowed_on_key and tier_permission_denied (403), insufficient_credits (402), and rate_limited, concurrency_limit and window_exhausted (429). Errors lists them all, and Troubleshooting gives a fix for each.

Warning

Keep the key in an environment variable or a secrets manager. If it ends up in a committed file or a notebook you've shared, rotate it in /dashboard/keys. The old secret stops working immediately.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.