Tokens documentation
One API key for coding models, behind OpenAI- and Anthropic-compatible endpoints. Connect your coding agent in a few minutes, call the API from your own code, and keep spending under control with per-key caps.
Endpoints and authentication
- OpenAI-compatible base URL
https://tokens.bd/v1- Anthropic-compatible base URL
https://tokens.bd- Authentication
- Authorization: Bearer tok_live_...
OpenAI-style tools add /v1; Anthropic-style tools such as Claude Code take the base without it. Authentication details
Start here
- Platform overviewStart hereWhat Tokens is, how a request travels from your tool to the model provider, what gets metered, and the handful of concepts you need before your first call.Read guide
- QuickstartPopularCreate an account, add credit, create an API key and make your first request to Tokens with cURL, Python or Node.js, then connect your coding agent.Read guide
- Claude CodePopularConnect Claude Code to Tokens through the Anthropic-compatible endpoint, map the opus, sonnet and haiku aliases to a Tokens model, and fix the usual gateway errors.Read guide
- OpenCodePopularAdd Tokens to OpenCode as an OpenAI-compatible provider in opencode.json, keep the key in TOKENS_API_KEY, set context limits so compaction works, and switch models.Read guide
- OpenClawPopularConnect OpenClaw to Tokens as a custom provider, either with non-interactive onboarding or by editing ~/.openclaw/openclaw.json, then set real context limits and switch models.Read guide
- Hermes AgentNewPoint Nous Research's Hermes Agent at Tokens as a custom chat completions endpoint, keep the key in ~/.hermes/.env, set a context length of at least 64K, and switch models.Read guide
Connect your agents in one command
The Tokens CLI signs you in through the browser, picks a default model and configures OpenCode, Claude Code, Codex CLI and Crush, backing up each file first. CLI reference
curl -fsSL https://tokens.bd/cli/tokens.mjs -o tokens.mjs
node tokens.mjs setup --base-url https://tokens.bdAll documentation
Getting Started
Create a key, make your first request and pick a model.
- Platform overviewWhat Tokens is, how a request travels from your tool to the model provider, what gets metered, and the handful of concepts you need before your first call.
- QuickstartCreate an account, add credit, create an API key and make your first request to Tokens with cURL, Python or Node.js, then connect your coding agent.
- API keysCreate, limit, store, rotate and revoke Tokens API keys, and understand the key-related errors: invalid_api_key, key_inactive, model_not_allowed_on_key and monthly_spend_cap_exceeded.
- Choosing a modelPick a model by the job: long agentic coding sessions, cheap high-volume edits, big-context repo work, vision input or fast interactive use. Includes a comparison of representative models and how model ids work on Tokens.
- Tokens CLI (one-command agent setup)Use the Tokens CLI to sign in through your browser, pick a default model and configure OpenCode, Claude Code, Codex CLI and Crush in one command, with backups of every file it changes.
Coding Agents
Connect Claude Code, Codex CLI, OpenCode, OpenClaw, Hermes and other agents.
- Claude CodeConnect Claude Code to Tokens through the Anthropic-compatible endpoint, map the opus, sonnet and haiku aliases to a Tokens model, and fix the usual gateway errors.
- Codex CLIAdd Tokens as a custom model provider in Codex CLI's config.toml, keep the key in an environment variable, and handle models whose upstream doesn't support the Responses API.
- OpenCodeAdd Tokens to OpenCode as an OpenAI-compatible provider in opencode.json, keep the key in TOKENS_API_KEY, set context limits so compaction works, and switch models.
- OpenClawConnect OpenClaw to Tokens as a custom provider, either with non-interactive onboarding or by editing ~/.openclaw/openclaw.json, then set real context limits and switch models.
- Hermes AgentPoint Nous Research's Hermes Agent at Tokens as a custom chat completions endpoint, keep the key in ~/.hermes/.env, set a context length of at least 64K, and switch models.
- CrushAdd Tokens to Charm's Crush as an openai-compat provider, using the new crushrc format or the older crush.json the Tokens CLI writes, then pick large and small models.
- Connect Cursor to TokensPoint Cursor's chat at the Tokens OpenAI-compatible endpoint with a custom OpenAI base URL, and know which Cursor features will keep using Cursor's own models.
- Connect Cline to TokensSet up the Cline coding agent in VS Code with the OpenAI Compatible provider, the Tokens base URL, your key and a model ID, plus the model settings that matter.
- Connect Roo Code to TokensConfigure Roo Code's OpenAI Compatible provider for Tokens. The Roo Code repository was archived in May 2026, so this page also points to maintained alternatives.
- Connect Kilo Code to TokensAdd Tokens to Kilo Code as a custom OpenAI Compatible provider, from the settings UI or a kilo.json file, with model limits set so context management works.
- Connect Continue to TokensAdd Tokens models to Continue in VS Code or JetBrains with a config.yaml entry: provider openai, apiBase, key, roles, tool use and context length.
- Connect Zed to TokensUse Tokens models in Zed's Agent Panel through an OpenAI-compatible provider in settings.json, with the API key kept out of the file.
- Connect GitHub Copilot to TokensUse Tokens models in GitHub Copilot Chat in VS Code with the bring-your-own-key Custom Endpoint provider and a chatLanguageModels.json entry.
- AiderRun Aider against Tokens with OPENAI_API_BASE and the openai/ model prefix, silence unknown-model warnings with a metadata file, and switch models per session.
- GooseAdd Tokens to Goose as a custom OpenAI-compatible provider through goose configure or a JSON file in custom_providers, keep the key in TOKENS_API_KEY, and switch models.
- Qwen CodeConnect Qwen Code to Tokens through its OpenAI protocol in ~/.qwen/settings.json or with three environment variables, then switch between Tokens models with /model.
- Kimi Code CLIAdd Tokens to Moonshot's Kimi Code CLI as an openai provider in ~/.kimi-code/config.toml, read the key from TOKENS_API_KEY, and map local model aliases to Tokens model ids.
- Connect Factory Droid to TokensAdd Tokens models to Factory's Droid CLI as custom models in ~/.factory/settings.json, using the Chat Completions or Anthropic Messages protocol.
- Connect Warp to TokensPoint Warp's agents at Tokens with a custom inference endpoint that speaks OpenAI Chat Completions, and know where Warp will not use it.
- Connect Amp to TokensRoute models in Amp's own catalog through Tokens with a Model Routing Custom URL connection. Amp cannot add arbitrary model IDs, so read the limits first.
- Other tools and compatibilityWhich coding agents and editors can use a custom OpenAI or Anthropic endpoint like Tokens, which cannot, and a generic recipe for connecting any OpenAI-compatible tool.
SDKs & Libraries
Call Tokens from cURL, Python, Node.js and popular AI frameworks.
- cURLCall the Tokens API with cURL: chat completions, streaming, Anthropic Messages, listing models, checking usage and reading request IDs from error responses.
- Python (OpenAI SDK)Use the official openai Python package with Tokens: client setup, sync and async calls, streaming, tool calls, timeouts, retries and error handling.
- Node.js and TypeScript (OpenAI SDK)Use the official openai npm package with Tokens from Node.js and TypeScript: client setup, streaming, error handling with APIError, and a Next.js route handler that keeps your key on the server.
- Anthropic SDK (Python and TypeScript)Point the official Anthropic Python and TypeScript SDKs at Tokens with base URL https://tokens.bd, then call messages.create and stream with any model in the Tokens catalog.
- Vercel AI SDKUse Tokens as a provider in the Vercel AI SDK with createOpenAICompatible from @ai-sdk/openai-compatible, then call generateText and streamText, including a Next.js route.
- LangChain and LiteLLMUse Tokens from LangChain (Python and JavaScript) with ChatOpenAI and a custom base URL, and from LiteLLM as an SDK or proxy with the openai/ model prefix.
API Reference
Endpoints, authentication, streaming, tool calling, errors and limits.
- AuthenticationBase URLs, the two supported auth headers, the tok_live_ key format, and what the 401 and 403 error codes mean.
- Chat CompletionsPOST /v1/chat/completions: request fields, a full request and response, the usage object, and how the gateway treats max_tokens, n and streaming.
- Messages (Anthropic API)Call POST /v1/messages with Anthropic SDKs or curl: base URL, headers, a full example, streaming events, error shapes, and which headers are not forwarded.
- Responses APIPOST /v1/responses, the OpenAI Responses API: when to use it instead of chat completions, request and response examples, max_output_tokens and streaming events.
- Models and Usage EndpointsGET /v1/models lists the models your key can call. GET /v1/tokens/usage returns your plan, usage windows, wallet balance and key limits. Plus notes on embeddings and legacy completions.
- StreamingHow Server-Sent Events streaming works through the Tokens gateway for chat completions and messages: the event format, usage chunks, timeouts, and handling disconnects.
- Tool CallingUse OpenAI-style function tools and Anthropic-style tools through the Tokens gateway, with a complete runnable tool loop in Python and tips on model support.
- ErrorsEvery error code the Tokens API returns, what it means and what to do, plus the error JSON shape, request ids for support tickets, and which errors to retry.
- Rate LimitsRequests per minute, concurrency, plan usage windows and per-key monthly spend caps: how each limit works, the errors they return, and how to back off correctly.
Account & Billing
Plans, wallet, paying in BDT, usage alerts, security and support.
- Plans, credits and walletHow subscription plans, credits, usage windows and the pay-as-you-go wallet work on Tokens, what happens when a window or your balance runs out, and how renewal, cancellation and coupons work.
- Paying in BDTPay for Tokens plans and wallet top-ups in Bangladeshi taka: available payment methods, how the exchange rate is locked at checkout, the minimum top-up, receipts and refunds.
- Usage, limits and alertsTrack spend and remaining allowance on Tokens from the dashboard, a CSV export, the GET /v1/tokens/usage endpoint or the CLI, set up usage alerts, and handle rate limits and Retry-After correctly.
- Security and data privacyWhat Tokens stores and doesn't store about your requests, how your prompts reach model providers, and how to secure your account and API keys with MFA, session control and key hygiene.
- Getting helpHow to get help with Tokens: check the status page, find the request id, and open a support ticket in the dashboard with the details that get it solved fast.
Troubleshooting
Fix common errors and find answers to frequent questions.
- TroubleshootingFix common Tokens API errors by symptom: 401 invalid key, 403 model not allowed, 402 insufficient credits, 429 limits, 404 model not found, 5xx, wrong base URL, stuck streams and Windows env vars.
- FAQShort answers to common questions about Tokens: OpenAI and Anthropic compatibility, models, privacy, paying in BDT, balances, keys, receipts, account deletion and uptime.
Generate a config file
Pick a tool or language to get a ready-to-paste configuration with your base URL.
{
"provider": {
"tokens": {
"npm": "@ai-sdk/openai-compatible",
"options": {
"baseURL": "https://tokens.bd/v1",
"apiKey": "tok_live_your_key"
},
"models": {
"deepseek/deepseek-v4.1-flash": {}
}
}
}
}Project or global configuration for OpenCode CLI / IDE agent.
{
"env": {
"ANTHROPIC_BASE_URL": "https://tokens.bd",
"ANTHROPIC_AUTH_TOKEN": "tok_live_your_key",
"ANTHROPIC_MODEL": "deepseek/deepseek-v4.1-flash"
}
}Configure Claude Code CLI via ~/.claude/settings.json or environment exports.
# ~/.codex/config.toml model = "deepseek/deepseek-v4.1-flash" model_provider = "tokens" [model_providers.tokens] name = "Tokens" base_url = "https://tokens.bd/v1" env_key = "TOKENS_API_KEY" wire_api = "responses" # In your shell profile (~/.bashrc, ~/.zshrc): export TOKENS_API_KEY="tok_live_your_key"
Add a custom model provider to ~/.codex/config.toml and export the key in your shell.
export OPENAI_API_BASE="https://tokens.bd/v1" export OPENAI_API_KEY="tok_live_your_key" aider --model openai/deepseek/deepseek-v4.1-flash
Run Aider terminal pair programmer against the Tokens gateway.
from openai import OpenAI
client = OpenAI(
base_url="https://tokens.bd/v1",
api_key="tok_live_your_key",
)
response = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
messages=[
{"role": "system", "content": "You are an expert coding assistant."},
{"role": "user", "content": "Write a quick hello world in Python"},
],
temperature=0.7,
)
print(response.choices[0].message.content)Official OpenAI Python library pointing to Tokens endpoint.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://tokens.bd/v1",
apiKey: "tok_live_your_key",
});
async function main() {
const response = await client.chat.completions.create({
model: "deepseek/deepseek-v4.1-flash",
messages: [
{ role: "system", content: "You are an expert coding assistant." },
{ role: "user", content: "Write a quick hello world in Python" },
],
temperature: 0.7,
});
console.log(response.choices[0]?.message?.content);
}
main().catch(console.error);Official OpenAI Node.js SDK pointing to Tokens endpoint.
curl -X POST "https://tokens.bd/v1/chat/completions" \
-H "Authorization: Bearer tok_live_your_key" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"messages": [
{"role": "system", "content": "You are an expert coding assistant."},
{"role": "user", "content": "Write a quick hello world in Python"}
],
"temperature": 0.7
}'Raw HTTP POST request to OpenAI-compatible chat completions.
Test your key
Send a small real request with your key to confirm everything works before you configure a tool. It uses a few tokens from your balance.
Live connection test
Check latency and token generation before you start coding.