Skip to content

Anthropic SDK (Python and TypeScript)

Point the official Anthropic Python and TypeScript SDKs at Tokens with base URL https://tokens.bd, then call messages.create and stream with any model in the Tokens catalog.

Works withAnthropic Python SDKAnthropic TypeScript SDK
On this page

Tokens exposes an Anthropic-compatible Messages API, so the official Anthropic SDKs work with one change: the base URL. If your code is already written against messages.create, you can keep it and switch providers by changing the model string.

Use the Anthropic SDK base URL: https://tokens.bd#

Set the base URL to https://tokens.bd, without /v1. The SDKs append /v1/messages themselves. If you set https://tokens.bd/v1, requests go to /v1/v1/messages and fail with a 404.

SettingValue
Base URLhttps://tokens.bd
API keyyour Tokens key (tok_live_...)
Modelany ID from /models, e.g. deepseek/deepseek-v4.1-flash

The SDK sends the key as x-api-key. Tokens accepts that or Authorization: Bearer.

Which models work#

Any model in the Tokens catalog, not just Claude models. When you send an Anthropic-format request for a model from another provider, the gateway translates the request and the response, so your code receives normal Anthropic Message objects and stream events either way. Check which IDs your key can use with GET https://tokens.bd/v1/models (see cURL).

Features beyond plain text depend on the model you pick. Tool use, extended thinking and prompt caching only work where the underlying model and provider support them. Check the model's page in /models before you rely on one of them.

Python#

bash
pip install --upgrade anthropic
export TOKENS_API_KEY="tok_live_your_key"
hello_anthropic.py
import os
import anthropic

client = anthropic.Anthropic(
    base_url="https://tokens.bd",
    api_key=os.environ["TOKENS_API_KEY"],
)

message = client.messages.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=512,
    system="You are a concise senior engineer.",
    messages=[{"role": "user", "content": "When is a B-tree index the wrong choice?"}],
)

for block in message.content:
    if block.type == "text":
        print(block.text)

print(message.usage)

max_tokens is required by the Messages API, as it is with Anthropic directly.

Stream with the Python SDK#

python
with client.messages.stream(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Write a bash one-liner that finds the 10 largest files."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
    final = stream.get_final_message()

print("\n", final.usage)

The async client is anthropic.AsyncAnthropic with the same arguments. Use async with client.messages.stream(...) and async for text in stream.text_stream.

TypeScript#

The Anthropic TypeScript SDK needs Node.js 20 or later.

bash
npm install @anthropic-ai/sdk
hello-anthropic.ts
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  baseURL: "https://tokens.bd",
  apiKey: process.env.TOKENS_API_KEY,
});

const message = await client.messages.create({
  model: "deepseek/deepseek-v4.1-flash",
  max_tokens: 512,
  messages: [{ role: "user", content: "Explain optimistic locking in three sentences." }],
});

for (const block of message.content) {
  if (block.type === "text") console.log(block.text);
}

Stream with the TypeScript SDK#

ts
const stream = client.messages
  .stream({
    model: "deepseek/deepseek-v4.1-flash",
    max_tokens: 1024,
    messages: [{ role: "user", content: "Draft a short commit message for a null-check fix." }],
  })
  .on("text", (text) => process.stdout.write(text));

const final = await stream.finalMessage();
console.log("\n", final.usage);

If you only need the raw events and want to use less memory, client.messages.create({ ..., stream: true }) returns an async iterable instead.

Environment variables instead of code#

Both SDKs read ANTHROPIC_API_KEY and ANTHROPIC_BASE_URL when you construct the client with no arguments:

bash
export ANTHROPIC_BASE_URL="https://tokens.bd"
export ANTHROPIC_API_KEY="$TOKENS_API_KEY"

Be careful with this in a shell where you also run Claude Code or other Anthropic-based tools, because they read the same variables. Claude Code has its own setup with ANTHROPIC_AUTH_TOKEN in ~/.claude/settings.json, covered in Claude Code.

Errors and retries#

The SDKs raise their usual APIError subclasses (AuthenticationError for 401, PermissionDeniedError for 403, RateLimitError for 429, and so on). On /v1/messages, errors coming from the upstream provider use the Anthropic error shape. Errors from Tokens itself (bad key, credits, plan limits) use the OpenAI-style body, {"error": {"message", "type", "code", ...}}, so read code from the error body when you need to tell insufficient_credits from window_exhausted.

python
try:
    client.messages.create(model="deepseek/deepseek-v4.1-flash", max_tokens=64,
                           messages=[{"role": "user", "content": "hi"}])
except anthropic.APIStatusError as e:
    print(e.status_code, e.response.headers.get("x-tokens-request-id"), e.body)

Include the x-tokens-request-id header value in any support ticket. Both SDKs retry 429 and 5xx twice by default. That's reasonable here, but retrying 429 window_exhausted won't succeed until the time in Retry-After has passed. Codes and fixes are in Errors, and the request format is in Messages.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.