# Platform overview > What Tokens is, how a request travels from your tool to the model provider, what gets metered, and the handful of concepts you need before your first call. Section: Getting Started. Page: https://tokens.bd/docs/overview Tokens is an AI model gateway: one API key gives you OpenAI-compatible and Anthropic-compatible endpoints for models from many providers, and you pay in USD or BDT. This page explains how the platform works so the rest of the docs make sense. ## What Tokens is Most coding agents and SDKs already speak one of two dialects: the OpenAI API or the Anthropic Messages API. Tokens speaks both. You point your tool at Tokens instead of at a single provider, use one `tok_live_` key, and pick any model your plan or wallet covers by changing the `model` field. | You use | Base URL | Typical clients | | ------------------------ | ---------------------- | ----------------------------------------------------------------- | | OpenAI-compatible API | `https://tokens.bd/v1` | OpenAI SDKs, OpenCode, Codex CLI, Cursor, Cline, Aider, Continue | | Anthropic-compatible API | `https://tokens.bd` | Claude Code, the Anthropic SDK (it appends `/v1/messages` itself) | The full model list with prices is the [model catalog](/models). Plans and top-ups are on [pricing](/pricing). ## How a request flows 1. **Your tool sends a request** to `https://tokens.bd/v1/...` with your key in `Authorization: Bearer` or `x-api-key`. 2. **Tokens checks it**: the key is valid and active, the model is allowed on that key and on your plan, you are inside your rate limits and usage windows, and you have credit to pay for it. 3. **Tokens forwards it to an upstream source** for that model. If that source answers with 429, 502, 503, 504 or drops the connection, Tokens retries on another source for the same model automatically. 4. **The response streams back** to your tool unchanged (SSE for streaming requests), and Tokens meters the tokens as they pass. 5. **The cost is settled** against your plan credits or wallet when the response finishes. Every response carries an `x-tokens-request-id` header. Keep it when something goes wrong; it is the fastest way for [support](/docs/support) to find your request. :::note Tokens adds one network hop between you and the provider. Long requests are fine (response headers can take up to 600 seconds), but don't expect lower latency than calling the provider directly. ::: ## What the API covers Supported endpoints: `POST /v1/chat/completions`, `POST /v1/messages`, `POST /v1/responses`, `POST /v1/completions` (legacy), `POST /v1/embeddings` (only for models that are embedding models), `GET /v1/models` and `GET /v1/tokens/usage`. Not supported: image generation, audio, files, batches, assistants, fine-tuning and moderations. Calls to those return `404 unsupported_endpoint`. Browser calls are not supported either, because responses carry no CORS headers. Call Tokens from a server, a script or a CLI tool. ## What is metered Tokens meters every inference request, from `/v1` and from the dashboard playground. For each request it records: - the model, input tokens, output tokens and cache-read tokens - the cost at that model's catalog price - latency, timestamps, the key used and the request id That metadata is what you see in the usage dashboard and the CSV export. Prompt and response content is not stored. See [security and data privacy](/docs/security-and-privacy) for the details. Read-only calls such as `GET /v1/tokens/usage` are not metered. ## Key concepts ### API key A secret that starts with `tok_live_`. It is shown once when you create it in [/dashboard/keys](/dashboard/keys). A key can carry an optional monthly spend cap and an optional list of allowed models. See [API keys](/docs/api-keys). ### Model id Models are addressed by an alias in `provider/model` form, for example `deepseek/deepseek-v4.1-flash`. The alias is what goes in the `model` field. It does not always match the provider's own id, so copy it from [/models](/models) or from `GET /v1/models`, which returns exactly the models your key can use. [Choosing a model](/docs/choosing-a-model) helps you pick one. ### Plan vs wallet There are two ways to pay for usage, and you can have both: - **A plan** is a weekly or monthly subscription that gives you a credit allowance (100 credits = 1 USD by default) and may include usage windows. - **The wallet** is a prepaid USD balance for pay-as-you-go use. You can top it up in USD or BDT. Details are in [plans, credits and wallet](/docs/plans-and-wallet). ### Usage windows Some plans spread their allowance over time with windows: a rolling 5-hour session, a weekly ceiling and a monthly ceiling. When a window is used up, requests return `429 window_exhausted` with a `Retry-After` header that says how long until it resets. You can see every window in the dashboard, with `GET /v1/tokens/usage`, or with the CLI's `usage` command. See [usage, limits and alerts](/docs/usage-and-alerts). ## Where to go next - [Quickstart](/docs/quickstart): create a key and make your first request in a few minutes. - [Tokens CLI](/docs/tokens-cli): configure OpenCode, Claude Code, Codex CLI and Crush with one command. - [Claude Code](/docs/claude-code), [Codex CLI](/docs/codex-cli), [OpenCode](/docs/opencode), [Cursor](/docs/cursor): per-agent setup guides. - [Errors](/docs/errors): every error code and what to do about it. --- # Quickstart > Create an account, add credit, create an API key and make your first request to Tokens with cURL, Python or Node.js, then connect your coding agent. Section: Getting Started. Page: https://tokens.bd/docs/quickstart This quickstart takes you from a new account to a working API call, then to a connected coding agent. You need a terminal and, for the SDK examples, a recent Python 3 or Node.js 18+. ## 1. Create an account Sign up at [tokens.bd/register](/register) with email or Google. Verify your email address: you can't create API keys until it is verified. ## 2. Add credit: a plan or a top-up Keys only work when there is something to pay for usage with. Pick one (or both) in [/dashboard/billing](/dashboard/billing): - **Subscribe to a plan** for a weekly or monthly credit allowance. Compare them on [pricing](/pricing). - **Top up your wallet** for pay-as-you-go use. The minimum top-up is $5 or ৳500. Prices are shown in USD, BDT. Available payment methods: Manual Payment (bKash, Nagad). If you want to pay in taka, read [paying in BDT](/docs/paying-in-bdt) first. ## 3. Create an API key Open [/dashboard/keys](/dashboard/keys) and create a key. Give it a name you will recognise later ("laptop", "ci", "staging-bot"). You can also set a monthly spend cap and limit the key to certain models; both are optional, and both are fixed once the key exists. [API keys](/docs/api-keys) explains when to use them. The secret starts with `tok_live_` and is shown **once**. Copy it now and put it in an environment variable: :::code-tabs ```bash title="macOS / Linux" export TOKENS_API_KEY="tok_live_your_key" ``` ```powershell title="Windows PowerShell" $env:TOKENS_API_KEY = "tok_live_your_key" ``` ::: :::warning Never paste the key into code you commit, and never ship it in a web page. Browser calls to Tokens are not supported; call it from a server, script or CLI. ::: ## 4. Make your first request The OpenAI-compatible base URL is `https://tokens.bd/v1`. The examples use `deepseek/deepseek-v4.1-flash`; swap in any model id from [/models](/models) that your plan or wallet covers. :::code-tabs ```bash title="cURL" curl https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "messages": [{"role": "user", "content": "Say hello in one short sentence."}] }' ``` ```python title="Python" # pip install openai import os from openai import OpenAI client = OpenAI( base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], ) response = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": "Say hello in one short sentence."}], ) print(response.choices[0].message.content) ``` ```js title="Node.js" // npm install openai (save as hello.mjs, run: node hello.mjs) import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY, }); const response = await client.chat.completions.create({ model: "deepseek/deepseek-v4.1-flash", messages: [{ role: "user", content: "Say hello in one short sentence." }], }); console.log(response.choices[0].message.content); ``` ::: A successful response is a normal OpenAI chat completion. If you get an error instead, the `error.code` field tells you why: | Code | Usual cause | Fix | | ------------------------------------- | ------------------------------------------------------------ | --------------------------------------------------------------- | | `invalid_api_key` | Typo, or the variable isn't set in this shell | `echo $TOKENS_API_KEY` and re-export | | `model_not_found` | The model id is wrong | Copy the id from `GET /v1/models` | | `tier_permission_denied` | Your plan doesn't include that model and the wallet is empty | Pick another model or top up | | `no_funding` / `insufficient_credits` | No plan and no wallet balance | Subscribe or top up in [/dashboard/billing](/dashboard/billing) | The full list is in [errors](/docs/errors). ### List the models your key can use ```bash curl https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` The `id` of each entry in `data` is exactly what you put in `model`. ## 5. Connect your coding agent The quickest route for OpenCode, Claude Code, Codex CLI and Crush is the Tokens CLI. It signs you in through the browser, lets you pick a default model, and edits the agents' config files with a backup of each: :::code-tabs ```bash title="macOS / Linux" curl -fsSL https://tokens.bd/cli/tokens.mjs -o tokens.mjs && node tokens.mjs setup --base-url https://tokens.bd ``` ```powershell title="Windows PowerShell" iwr https://tokens.bd/cli/tokens.mjs -OutFile tokens.mjs; node tokens.mjs setup --base-url https://tokens.bd ``` ::: See [Tokens CLI](/docs/tokens-cli) for what it changes. If you prefer to configure things by hand, or use another tool, [/dashboard/connect](/dashboard/connect) has copy-and-paste settings for each agent, and the guides cover them in detail: - [Claude Code](/docs/claude-code) (Anthropic-compatible, base URL `https://tokens.bd`) - [Codex CLI](/docs/codex-cli), [OpenCode](/docs/opencode), [Crush](/docs/crush) - [Cursor](/docs/cursor), [Cline](/docs/cline), [Roo Code](/docs/roo-code), [Kilo Code](/docs/kilo-code), [Continue](/docs/continue), [Zed](/docs/zed), [Aider](/docs/aider) ## Test your key from this page Paste your key below to send a short test request through the gateway. It costs a few tokens. ::connection-tester ## Next steps - [Choosing a model](/docs/choosing-a-model): which model for which job. - [Usage, limits and alerts](/docs/usage-and-alerts): watch spend and set up alerts before you hand a key to an agent. - [Streaming](/docs/streaming) and [tool calling](/docs/tool-calling) for the API details agents rely on. --- # API keys > Create, limit, store, rotate and revoke Tokens API keys, and understand the key-related errors: invalid_api_key, key_inactive, model_not_allowed_on_key and monthly_spend_cap_exceeded. Section: Getting Started. Page: https://tokens.bd/docs/api-keys A Tokens API key authenticates every call to `/v1`. This page covers creating keys, the two limits you can put on them, storing them safely, rotating and revoking, and what each key-related error means. ## Create an API key Open [/dashboard/keys](/dashboard/keys) and create a key. Two things must be true first: - your email address is verified, and - you have an active plan or a positive wallet balance. The form asks for: | Field | Required | What it does | | ----------------------- | -------- | ------------------------------------------------------------------------------------ | | Name | Yes | A label for you. Use where the key lives: "laptop", "github-actions", "support-bot". | | Monthly spend cap (USD) | No | The most this key may spend in a calendar month. | | Allowed models | No | The only model ids this key may call. Empty means every model your account can use. | The secret looks like `tok_live_` followed by 48 hex characters. It is shown **once**, right after creation. Tokens stores only a hash, so nobody can show it to you again. If you lose it, rotate the key or create a new one. Send it as either header: ```bash curl https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" curl https://tokens.bd/v1/models -H "x-api-key: $TOKENS_API_KEY" ``` ## Set a monthly spend cap and allowed models If you only change one setting, make it the spend cap. Coding agents run long loops, and a cap turns a runaway session into an error instead of a bill. - **Monthly spend cap.** Tokens adds up what the key has spent since the start of the calendar month. Before each request it also counts the most that request could cost, so a request that might push the key over the cap is refused with `monthly_spend_cap_exceeded` slightly before the exact figure is reached. If you leave the field blank and your plan defines a default cap, the key gets that default. - **Allowed models.** Restrict a key to the models a tool actually needs. A CI job that runs one cheap model shouldn't be able to call an expensive one. `GET /v1/models` with a restricted key lists only the allowed models. :::warning[Limits are fixed at creation] The spend cap and the allowed-models list can't be edited after the key is created. To change either, create a new key with the new limits, move your tools to it, then revoke the old one. ::: ## How many keys can I have? Each plan sets a maximum number of active keys; the default is 3, and pay-as-you-go accounts without a plan also default to 3. When you reach it, creating another returns `key_limit_reached` ("Your ... plan allows N active keys"). Revoke a key you no longer use, or move to a plan with a higher limit on [pricing](/pricing). One key per tool or machine is a good habit. It makes the usage dashboard readable and lets you revoke one leaked key without breaking everything else. ## Store API keys safely Treat a key like a password. It spends your money. - **Use an environment variable.** All examples in these docs read `TOKENS_API_KEY`. - **For projects, use a `.env` file and ignore it in git:** ```bash title=".env" TOKENS_API_KEY=tok_live_your_key ``` ```bash title=".gitignore" .env .env.* ``` - **In CI**, store the key as a secret in your CI provider and expose it as an environment variable at run time. - **Never put a key in client-side code.** Browser and mobile apps ship their source to users. Tokens doesn't support browser calls anyway (responses have no CORS headers), so put a small server or serverless function in between and keep the key there. - **Agent config files.** Claude Code, OpenCode and Crush keep the key in their config files under your home directory. That's fine on your own machine, but don't commit those files to a dotfiles repository. :::danger[If a key leaks] Rotate or revoke it in [/dashboard/keys](/dashboard/keys) immediately, then check [usage](/dashboard/usage) for requests you don't recognise. Deleting a commit doesn't help: assume anything pushed to a remote has been copied. ::: ## Rotate an API key Rotation replaces the secret and keeps everything else about the key: its name, spend cap, allowed models and usage history. **The old secret stops working immediately.** There is no grace period. Any tool still using it gets `401 invalid_api_key` on its next request. So rotate in this order: 1. Have the places that use the key ready to update (env vars, CI secrets, agent configs). 2. Click **Rotate** on the key and copy the new secret. 3. Update every place that used the old one. If you can't afford a short outage, create a second key first, switch your tools to it, and then revoke the old one. ## Revoke an API key Revoking disables a key permanently, effective immediately. Use it for keys you no longer need and for any key you think has leaked. Revoked keys no longer count towards your active-key limit. Signing out of the [Tokens CLI](/docs/tokens-cli) with `logout` only deletes the copy on your computer; the key itself stays valid until you revoke it here. ## Key error codes Errors on the OpenAI-style endpoints come back in the OpenAI shape (on `/v1/messages` they use Anthropic's, see [Errors](/docs/errors)), with a `request_id` you can quote to [support](/docs/support): ```json { "error": { "message": "Model 'example/model' is not permitted on this API key. Permitted models: deepseek/deepseek-v4.1-flash.", "type": "permission_denied_error", "code": "model_not_allowed_on_key", "param": null, "request_id": "..." } } ``` | HTTP | Code | Meaning | What to do | | ---- | ---------------------------- | ----------------------------------------------------------------------------------------- | ------------------------------------------------------------------- | | 401 | `missing_api_key` | No `Authorization` or `x-api-key` header | Check the variable is set in the shell that runs your tool | | 401 | `invalid_api_key` | The key doesn't exist, was mistyped, or was rotated | Copy the current secret, or create a new key | | 403 | `key_inactive` | The key is no longer active (for example, suspended) | Use another key, or ask support why it was suspended | | 403 | `key_expired` | The key was provisioned by an administrator with an expiry date, and that date has passed | Create a new key | | 403 | `model_not_allowed_on_key` | The model isn't on this key's allowed list; the message lists the allowed ones | Use an allowed model, or create a key that includes this one | | 403 | `monthly_spend_cap_exceeded` | This key has hit its monthly cap | Wait for the next calendar month, or create a key with a higher cap | | 403 | `account_suspended` | The whole account is suspended | Contact [support](/docs/support) | `monthly_spend_cap_exceeded` is about the key. If your plan or wallet has run out instead, you will see `window_exhausted`, `insufficient_credits` or `tier_permission_denied`; those are covered in [plans, credits and wallet](/docs/plans-and-wallet) and [errors](/docs/errors). --- # Choosing a model > Pick a model by the job: long agentic coding sessions, cheap high-volume edits, big-context repo work, vision input or fast interactive use. Includes a comparison of representative models and how model ids work on Tokens. Section: Getting Started. Page: https://tokens.bd/docs/choosing-a-model Choosing a model is mostly about the job you give it, not a leaderboard. This page groups models by the work they suit, compares a representative set, and explains how model ids work on Tokens. Everything you can buy is in [the catalog at /models](/models); the exact list for one key comes from `GET /v1/models`. :::note[About the prices on this page] Prices below are the providers' own list prices per 1M tokens in USD, checked October 2026. They show how models compare to each other. Your price on Tokens is on [/models](/models), and it can differ from these. ::: ## Model ids on Tokens Every model has an id in `provider/model` form, for example `deepseek/deepseek-v4.1-flash`. That id goes in the `model` field of every request and in your agent's config. - The id on Tokens is an alias and doesn't always match the provider's own id. DeepSeek's API, for example, calls its current Flash model `deepseek-flash`. Always copy ids from [/models](/models) or from the API. - The catalog shows everything Tokens offers. Your key may see fewer models: plans include different models, and a key can have an allowed-models list. ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` Each `data[].id` in the response is a model this key can call right now. ## Compare representative models | Model (maker) | Context | Known for | List price in / out | | ------------------------------ | --------------- | -------------------------------------------------------------------------- | ----------------------------------------------------------------------- | | Claude Sonnet 5.5 (Anthropic) | 1M, 128K out | Agentic coding, tool use, image input | $2 / $10 | | GPT-5.6 Sol (OpenAI) | 1.05M, 128K out | Complex, long-horizon agentic work | $4 / $20 (promotional until at least 2026-11-21; launch price $5 / $30) | | GPT-6 Luna (OpenAI) | 1.05M, 128K out | Focused, high-volume tasks | $0.10 / $0.50 | | Gemini 3.8 Flash (Google) | 1M in, 65K out | Long-horizon software engineering; text, image, video, audio and PDF input | $0.75 / $3.75 until 2026-12-31, then $1.50 / $7.50 | | Grok 4.7 (xAI) | 500K | Coding and knowledge work | $2 / $6 (prompts under 200K) | | Kimi K3 (Moonshot AI) | 1M | Long-horizon coding and deep reasoning; thinking always on | $3 / $15 | | GLM-5.3 (Z.ai) | 1M, 128K out | Coding and agents; text only; reasoning always on | $1.40 / $4.40 | | Qwen 3.8 Max (Alibaba) | 1M, 131K out | Coding, multi-tool orchestration, image and video input | $2 / $6 | | DeepSeek V4.1 Flash (DeepSeek) | 1M, 384K out | Fast and cheap, image input | $0.30 / $1.20 at peak hours, half off-peak | | MiniMax M3 (MiniMax) | 1M | Multimodal coding | $0.30 / $1.20 (input up to 512K) | | Qwen 3.8 Flash (Alibaba) | 1M, 131K out | Cheap, fast coding and agents | $0.15 / $0.47 | | MiMo V2.6 Flash (Xiaomi) | 1M, 128K out | Low-cost reasoning, vision input | $0.14 / $0.28 | The "known for" column paraphrases each maker's own positioning. It isn't a benchmark result. ## Best model for agentic coding with long sessions Agents like Claude Code, Codex CLI and OpenCode call the model dozens of times per task, resending the conversation and file contents each turn. Three things matter more than raw intelligence scores: 1. **Reliable tool calling.** A model that fumbles tool arguments wastes turns. 2. **Input price, and cached-input price.** Long sessions are dominated by input tokens. Caching is done by the upstream provider, and Tokens bills cached input at the model's cached-input price when one is set. The catalog does not show cached prices today: see [Prompt caching](/docs/prompt-caching). 3. **Context that holds up.** Most current flagships offer around 1M tokens. Good starting points: Claude Sonnet 5.5, GPT-5.6 Sol, Kimi K3, GLM-5.3, Gemini 3.8 Flash and Qwen 3.8 Max. Kimi K3 and GLM-5.3 always think before answering, which helps on hard problems but adds output tokens to every call. :::tip Use a strong model for the main agent and a cheap one for side work. Several agents let you set a separate small or fast model for titles, summaries and quick edits. See the guide for your agent, such as [Claude Code](/docs/claude-code) or [OpenCode](/docs/opencode). ::: ## Cheap models for high-volume edits For bulk refactors, commit messages, classification, test scaffolding or anything you run thousands of times, price per token decides. From the table: GPT-6 Luna ($0.10 / $0.50), MiMo V2.6 Flash ($0.14 / $0.28), Qwen 3.8 Flash ($0.15 / $0.47) and DeepSeek V4.1 Flash. Z.ai's GLM-5.3 Flash ($0.15 / $0.50) is in the same range and accepts image, video and file input. One detail on GPT-6 Luna: OpenAI supports its tool calling on Chat Completions only with `reasoning_effort` set to `none`; full tool support is on the Responses API. Tokens supports both `POST /v1/chat/completions` and `POST /v1/responses`. ## Big-context models for repo-wide work A 1M-token window is common now, but filling it is not free. At Claude Sonnet 5.5's list price, an 800K-token prompt costs $1.60 in input alone, every time you send it. Some providers also charge more past a threshold: - GPT-6 Luna and GPT-5.6 models: higher rates above 272K input tokens. - Grok 4.5 to 4.7: $4 / $12 at 200K prompt tokens and above. - MiniMax M3: double rates above 512K input. For most repo questions, a good agent that searches and reads selectively beats pasting the whole tree. When you do need the full window, check the price on [/models](/models) first. Also note that a request body to Tokens can be at most 10 MB. ## Vision and multimodal input If your prompts include screenshots, diagrams, PDFs or recordings, check input types before anything else: - **Broadest input:** Gemini 3.8 Flash (text, image, video, audio, PDF), Qwen 3.8 Omni Flash (text, image, audio, video), MiMo V2.6 Pro (text, image, video, audio). - **Image input with strong coding:** Claude Sonnet 5.5, Qwen 3.8 Max, Kimi K3, MiniMax M3, DeepSeek V4.1 Flash. - **Text only:** GLM-5.3, DeepSeek V4 Pro, LongCat 2.0 and Tencent Hy4 Preview. Images sent to these will fail or be ignored. Images go inside the chat message, in the format your SDK uses. Tokens doesn't offer image generation, audio or file-upload endpoints. ## Fast models for interactive use Several makers sell a faster-serving version of a model at a higher price. Their speed figures, as published by the makers: - Kimi K2.7 Code HighSpeed: about 180 tokens/s, at twice the price of K2.7 Code ($1.90 / $8). - GLM-5.3 FlashX: about 200 tokens/s, $0.37 / $1.25. - MiMo V2.6 Pro UltraSpeed: "up to 20x faster output" than MiMo V2.6 Pro, $4.35 / $8.70. - Step 3.5 Flash: 100 to 350 tokens/s, $0.10 / $0.30. Tokens adds a network hop, so time to first token through Tokens won't beat calling the provider directly. Turn on streaming (`"stream": true`) so text appears as it is generated; see [streaming](/docs/streaming). ## Try two models before you commit Pick two candidates and give them the same real task from your codebase. The dashboard Playground is the quickest way to compare single answers. For agent work, point your agent at each model for one session and compare cost per finished task in [usage](/docs/usage-and-alerts), not price per token. A model with a higher token price that finishes in fewer turns can be the cheaper one. When you've chosen, set it as the default in your agent: the [Tokens CLI](/docs/tokens-cli) asks for a default model during setup, or use the per-agent guides. Sources (list prices and specs, checked October 2026): [Anthropic pricing](https://platform.claude.com/docs/en/about-claude/pricing), [OpenAI pricing](https://developers.openai.com/api/docs/pricing), [Google Gemini pricing](https://ai.google.dev/gemini-api/docs/pricing), [xAI models](https://docs.x.ai/docs/models), [DeepSeek pricing](https://api-docs.deepseek.com/quick_start/pricing), [Kimi pricing](https://platform.kimi.ai/docs/pricing/chat), [Z.ai pricing](https://docs.z.ai/guides/overview/pricing), [Alibaba Model Studio pricing](https://www.alibabacloud.com/help/en/model-studio/model-pricing), [MiniMax pricing](https://platform.minimax.io/docs/guides/pricing-paygo.md), [Xiaomi MiMo](https://mimo.mi.com), [StepFun pricing](https://platform.stepfun.ai/docs/en/guides/pricing/details). --- # Tokens CLI (one-command agent setup) > Use the Tokens CLI to sign in through your browser, pick a default model and configure OpenCode, Claude Code, Codex CLI and Crush in one command, with backups of every file it changes. Section: Getting Started. Page: https://tokens.bd/docs/tokens-cli The Tokens CLI is a single JavaScript file that connects your coding agents to Tokens. One `setup` command signs you in through the browser, creates a key, lets you choose a default model and writes the config for OpenCode, Claude Code, Codex CLI and Crush. It needs Node.js 18 or newer and has no dependencies. ## Install and run the Tokens CLI There's nothing to install globally. Download the file and run it with Node: :::code-tabs ```bash title="macOS / Linux" curl -fsSL https://tokens.bd/cli/tokens.mjs -o tokens.mjs && node tokens.mjs setup --base-url https://tokens.bd ``` ```powershell title="Windows PowerShell" iwr https://tokens.bd/cli/tokens.mjs -OutFile tokens.mjs; node tokens.mjs setup --base-url https://tokens.bd ``` ::: It's plain, readable JavaScript (about 650 lines). If you're cautious about running downloaded scripts, and you should be, open `tokens.mjs` and read it before the second half of that command runs. Before you start, make sure your account has a plan or wallet balance. The CLI can't do anything useful on an empty account. ## What `setup` does, step by step ### 1. Signs you in through the browser Unless you pass `--key`, the CLI starts a device login. It prints a link and a code like `ABCD-EFGH`, and opens the link in your browser. On the page ([/dashboard/connect/cli](/dashboard/connect/cli)), check that the code matches the one in your terminal and approve it. The CLI waits and, once approved, receives a new API key named `CLI ()`. It's a normal key: you'll see it in [/dashboard/keys](/dashboard/keys) and can revoke it there. On a machine without a browser (SSH, a container), add `--no-browser` and open the printed link on any device where you're signed in. ### 2. Lists your models and asks for a default The CLI reads the gateway's base URLs from `https://tokens.bd/api/gateway/config`, then calls `GET /v1/models` with the key. It shows the first 10 models and asks you to pick one; Enter picks the first. To use a model that isn't in the first 10, pass it with `--model`. ### 3. Saves your credentials The base URL and key are saved to: - macOS / Linux: `~/.config/tokens/credentials.json` (or `$XDG_CONFIG_HOME/tokens/credentials.json`), readable only by you (mode 600) - Windows: `%APPDATA%\tokens\credentials.json` `usage` and `models` use this file later, so you don't need to pass the key again. ### 4. Finds your agents and shows what it will change It looks for OpenCode, Claude Code, Codex CLI and Crush, either by their config folder or by the program on your `PATH`. It prints the exact files it will update and asks `Continue? [Y/n]`. Nothing is written until you answer. | Agent | File | What is written | | ----------- | ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- | | OpenCode | `~/.config/opencode/opencode.json` | A `tokens` provider (`@ai-sdk/openai-compatible`) with your key, and `model` set to `tokens/` | | Claude Code | `~/.claude/settings.json` | `env.ANTHROPIC_BASE_URL`, `env.ANTHROPIC_AUTH_TOKEN` and `env.ANTHROPIC_MODEL` | | Codex CLI | `~/.codex/config.toml` (or `$CODEX_HOME/config.toml`) | A marked block with `[model_providers.tokens]` (`env_key = "TOKENS_API_KEY"`, `wire_api = "responses"`) and `[profiles.tokens]` | | Crush | `~/.config/crush/crush.json` | A `tokens` provider of type `openai-compat` with your key and model | JSON files are merged: your other settings stay. The Codex block sits between `# >>> tokens` and `# <<< tokens` markers and is replaced in place if you run setup again. ### 5. Backs up every file first Before changing a file, the CLI copies it to `.tokens-backup-`, for example `settings.json.tokens-backup-2026-10-03T09-15-00-000Z`. To undo, copy the backup over the file. If a file can't be edited safely (JSON with comments, a syntax error, or a Codex config that already has its own `[model_providers.tokens]` table), the CLI leaves it alone, marks it `skipped`, and points you to the paste-in settings at `/dashboard/connect?agent=`. ### 6. Codex CLI: one manual step Codex reads the key from an environment variable, so its config file holds no secret. After setup, add the variable to your shell profile and start Codex with the profile: :::code-tabs ```bash title="macOS / Linux" export TOKENS_API_KEY="tok_live_your_key" codex --profile tokens ``` ```powershell title="Windows PowerShell" setx TOKENS_API_KEY "tok_live_your_key" # open a new terminal, then: codex --profile tokens ``` ::: The CLI prints this line with your real key filled in. See [Codex CLI](/docs/codex-cli) for more. :::note OpenCode, Claude Code and Crush store the key in their config files in your home directory. Don't commit those files to a dotfiles repository. ::: ## Commands and flags ```text node tokens.mjs setup [--base-url URL] [--key KEY] [--model ID] [--agents opencode,claude,codex,crush] [--yes] [--dry-run] [--no-browser] node tokens.mjs usage [--json] node tokens.mjs models node tokens.mjs logout node tokens.mjs version node tokens.mjs help ``` | Flag | What it does | | ---------------- | -------------------------------------------------------------------------------------------------------------- | | `--base-url URL` | The Tokens site, `https://tokens.bd`. Needed the first time unless `TOKENS_BASE_URL` is set; saved afterwards. | | `--key KEY` | Use an existing key instead of the browser login. | | `--model ID` | Set the default model without the picker. Must be a model the key can use. | | `--agents LIST` | Configure only these agents (`opencode`, `claude`, `codex`, `crush`), even if they weren't detected. | | `--yes` | Skip the confirmation. Needed in scripts, where the prompt defaults to "no". | | `--dry-run` | Show which files would be written, without writing them or saving credentials. | | `--no-browser` | Print the sign-in link instead of opening it. | `usage` prints your plan, each usage window with a progress bar and reset time, your wallet balance and the key's cap and allowed models. `usage --json` prints the raw response of `GET /v1/tokens/usage` (see [usage, limits and alerts](/docs/usage-and-alerts)). `models` prints the model ids your key can use, one per line. :::tip[Dry runs still sign in] `--dry-run` doesn't write files, but without `--key` it still runs the browser login, which creates a key. To preview with no side effects at all, combine it with an existing key: `node tokens.mjs setup --base-url https://tokens.bd --dry-run --key "$TOKENS_API_KEY"`. ::: ## Environment variables | Variable | Used for | | ----------------- | ----------------------------------------------------------------------- | | `TOKENS_BASE_URL` | Base URL when `--base-url` isn't given | | `TOKENS_API_KEY` | Key when `--key` isn't given; takes priority over the saved credentials | The CLI picks the key in this order: `--key`, then `TOKENS_API_KEY`, then the saved credentials file. ## Troubleshooting **"Your account has no models available yet. Subscribe or add funds first."** The key works but can't use any model. Subscribe to a plan or top up your wallet in [/dashboard/billing](/dashboard/billing), then run `setup` again. **"That API key was not accepted."** The key was rotated, revoked or mistyped. Run `setup` without `--key` to get a fresh one, or check that an old `TOKENS_API_KEY` in your environment isn't overriding the saved key. **A file shows as `skipped`.** It has comments or an unexpected format. Configure that agent by hand from [/dashboard/connect](/dashboard/connect) or its guide: [OpenCode](/docs/opencode), [Claude Code](/docs/claude-code), [Codex CLI](/docs/codex-cli), [Crush](/docs/crush). **"No supported agents were found."** Your key is saved anyway. Install the agent first, or force one with `--agents claude`. For Cursor, Cline, Zed, Aider and others, use [/dashboard/connect](/dashboard/connect). **"The code expired" or "Timed out waiting for approval."** Approve the code within its time limit (up to 15 minutes). Run the command again for a new code. **"Could not reach https://tokens.bd."** Check your network or proxy, and the [status page](/status). ## Log out and uninstall ```bash node tokens.mjs logout ``` `logout` deletes the saved credentials file on this computer and nothing else. **The key stays valid.** To disable it, revoke it in [/dashboard/keys](/dashboard/keys). To remove Tokens from your agents, restore the `.tokens-backup-` copies or delete the settings the CLI added (the `tokens` provider in OpenCode and Crush, the three `ANTHROPIC_*` entries in Claude Code's `env`, the marked block in Codex's `config.toml`). Then delete `tokens.mjs` and, once you're happy, the backup files. The CLI leaves nothing else on disk. --- # Connect your tool > Find the guide for your tool: coding agents, editors, chat apps, automation tools, SDKs and agent frameworks that work with Tokens, and which protocol each one uses. Section: Getting Started. Page: https://tokens.bd/docs/integrations Tokens speaks two protocols, so most tools that let you set a custom endpoint work with it. This page is the map: find your tool, open its guide, and use the table at the end if yours is not listed. ## Pick by what you are doing | You want to... | Start here | | ---------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | Code with an AI agent in the terminal | [Claude Code](/docs/claude-code), [Codex CLI](/docs/codex-cli), [OpenCode](/docs/opencode), [Aider](/docs/aider), [OpenHands](/docs/openhands) | | Code inside your editor | [Cursor](/docs/cursor), [Cline](/docs/cline), [Continue](/docs/continue), [Zed](/docs/zed), [JetBrains AI Assistant](/docs/jetbrains-ai-assistant), [Xcode](/docs/xcode) | | Chat with models in a web interface you host | [Open WebUI](/docs/open-webui), [LibreChat](/docs/librechat) | | Build a workflow or an AI app without much code | [n8n](/docs/n8n), [Dify](/docs/dify) | | Call models from your own code | [cURL](/docs/curl), [Python](/docs/python), [Node.js](/docs/nodejs), [PHP](/docs/php), [Laravel AI SDK](/docs/laravel-ai-sdk) | | Build an agent with a framework | [OpenAI Agents SDK](/docs/openai-agents-sdk), [Claude Agent SDK](/docs/claude-agent-sdk), [LlamaIndex](/docs/llamaindex), [CrewAI](/docs/crewai), [LangChain](/docs/langchain), [Vercel AI SDK](/docs/vercel-ai-sdk) | | Move an app over from another provider | [From OpenAI](/docs/migrate-from-openai), [from OpenRouter](/docs/migrate-from-openrouter), [from Anthropic](/docs/migrate-from-anthropic) | ## Which protocol and base URL Which one a tool needs is nearly always on its settings screen, under a name such as "OpenAI compatible" or "Anthropic". | The tool speaks | Base URL | The tool then calls | | ----------------------- | ------------------------- | ---------------------- | | OpenAI Chat Completions | `https://tokens.bd/v1` | `/v1/chat/completions` | | OpenAI Responses | `https://tokens.bd/v1` | `/v1/responses` | | Anthropic Messages | `https://tokens.bd` | `/v1/messages` | Two tools in the lists above are different: [Xcode](/docs/xcode) wants the address without `/v1` and adds the path itself, and tools built on LiteLLM ([Aider](/docs/aider), [OpenHands](/docs/openhands), [CrewAI](/docs/crewai)) want an `openai/` prefix on the model id. Each guide says exactly what to type. ## Every guide, with what it needs | Tool | How it connects | | ------------------------------------------------ | ------------------------------------------------------------- | | [Claude Code](/docs/claude-code) | Anthropic Messages | | [Codex CLI](/docs/codex-cli) | Responses API | | [OpenCode](/docs/opencode) | Chat Completions | | [OpenClaw](/docs/openclaw) | Chat Completions or Anthropic Messages | | [Hermes Agent](/docs/hermes-agent) | Chat Completions | | [Crush](/docs/crush) | Chat Completions | | [Aider](/docs/aider) | Chat Completions, `openai/` prefix | | [Goose](/docs/goose) | Chat Completions, full endpoint URL | | [Qwen Code](/docs/qwen-code) | Chat Completions | | [Kimi Code CLI](/docs/kimi-code) | Chat Completions or Anthropic Messages | | [Factory Droid](/docs/factory-droid) | Chat Completions or Anthropic Messages | | [GitHub Copilot CLI](/docs/github-copilot-cli) | Chat Completions, set with environment variables | | [OpenHands](/docs/openhands) | Chat Completions through LiteLLM, `openai/` prefix | | [Cline](/docs/cline) | Chat Completions | | [Kilo Code](/docs/kilo-code) | Chat Completions | | [Roo Code](/docs/roo-code) | Chat Completions (the project is archived) | | [Continue](/docs/continue) | Chat Completions | | [Zed](/docs/zed) | Chat Completions or Anthropic Messages | | [Cursor](/docs/cursor) | Chat Completions, partly | | [GitHub Copilot (VS Code)](/docs/github-copilot) | Chat Completions, chat only | | [JetBrains AI Assistant](/docs/jetbrains-ai-assistant) | OpenAI-compatible provider | | [Xcode](/docs/xcode) | Chat Completions, address without `/v1` | | [Warp](/docs/warp) | Chat Completions, partly | | [Amp](/docs/amp) | Only models in Amp's own catalog | | [Open WebUI](/docs/open-webui) | OpenAI connection | | [LibreChat](/docs/librechat) | Custom endpoint in `librechat.yaml` | | [n8n](/docs/n8n) | OpenAI credential and Chat Model node | | [Dify](/docs/dify) | OpenAI-API-compatible provider | SDKs and frameworks are in the [SDKs and libraries](/docs/curl) section: cURL, Python, Node.js, the Anthropic SDK, the Vercel AI SDK, LangChain and LiteLLM, the OpenAI Agents SDK, the Claude Agent SDK, LlamaIndex, CrewAI, PHP and the Laravel AI SDK. ## If your tool is not here Any tool with a setting called "OpenAI compatible", "custom provider", "API base" or "base URL" can use Tokens. [Other tools and compatibility](/docs/other-tools) lists what does not work (Gemini CLI, which only accepts Google's format) and gives the four things every tool needs: the base URL, a key, the exact model id and the model's limits. Tools that only run in a web page cannot call Tokens directly, because Tokens' responses carry no CORS headers and a key must never be in code the browser downloads. [Browser and mobile apps](/docs/browser-and-mobile) shows the pattern that works. ## Before you connect anything 1. Create an API key with a monthly spend cap: [API keys](/docs/api-keys). 2. Pick a model your plan or wallet can use: [Choosing a model](/docs/choosing-a-model). 3. Check the key and model work with one request before you debug a tool's settings: the connection tester on the [docs home](/docs), or the curl command in [Quickstart](/docs/quickstart). --- # Claude Code > Connect Claude Code to Tokens through the Anthropic-compatible endpoint, map the opus, sonnet and haiku aliases to a Tokens model, and fix the usual gateway errors. Section: Coding Agents. Page: https://tokens.bd/docs/claude-code Claude Code is Anthropic's agentic coding CLI, which also runs inside VS Code and JetBrains. It speaks the Anthropic Messages API, so you point it at `https://tokens.bd` (no `/v1`) and it sends requests to [/v1/messages](/docs/messages). :::warning[Non-Claude models are best effort] Anthropic's gateway docs say it "doesn't support routing Claude Code to non-Claude models through any gateway". In practice it works when the model handles tool calls well, but that is per model and nobody guarantees it. If one struggles in long agent sessions, switch to another rather than fighting it. ::: ## Quick setup with the Tokens CLI The [Tokens CLI](/docs/tokens-cli) signs you in through the browser, creates a key, lets you pick a model and writes Claude Code's settings for you. It needs Node 18 or newer. :::code-tabs ```bash title="macOS / Linux" curl -fsSL https://tokens.bd/cli/tokens.mjs -o tokens.mjs && node tokens.mjs setup --base-url https://tokens.bd --agents claude ``` ```powershell title="Windows PowerShell" iwr https://tokens.bd/cli/tokens.mjs -OutFile tokens.mjs; node tokens.mjs setup --base-url https://tokens.bd --agents claude ``` ::: It merges three values into the `env` block of `~/.claude/settings.json`: `ANTHROPIC_BASE_URL`, `ANTHROPIC_AUTH_TOKEN` and `ANTHROPIC_MODEL`. The old file is kept as `settings.json.tokens-backup-`. If the file has comments or a syntax error, the CLI leaves it alone and points you to [/dashboard/connect](/dashboard/connect) instead. The CLI does not set the alias and beta variables described below. Add them afterwards; they prevent most of the odd behavior people hit with non-Claude models. ## Set ANTHROPIC_BASE_URL for Claude Code manually Install Claude Code first if you haven't (`curl -fsSL https://claude.ai/install.sh | bash` on macOS/Linux, `irm https://claude.ai/install.ps1 | iex` in PowerShell, or `npm install -g @anthropic-ai/claude-code`). Then create a key at [/dashboard/keys](/dashboard/keys) (see [API keys](/docs/api-keys)). ### Option A: environment variables Good for trying things out, or if you never want the key in a file. :::code-tabs ```bash title="macOS / Linux" export TOKENS_API_KEY="tok_live_your_key" export ANTHROPIC_BASE_URL="https://tokens.bd" export ANTHROPIC_AUTH_TOKEN="$TOKENS_API_KEY" export ANTHROPIC_MODEL="deepseek/deepseek-v4.1-flash" export ANTHROPIC_DEFAULT_OPUS_MODEL="deepseek/deepseek-v4.1-flash" export ANTHROPIC_DEFAULT_SONNET_MODEL="deepseek/deepseek-v4.1-flash" export ANTHROPIC_DEFAULT_HAIKU_MODEL="deepseek/deepseek-v4.1-flash" export CLAUDE_CODE_SUBAGENT_MODEL="deepseek/deepseek-v4.1-flash" export CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1 ``` ```powershell title="Windows PowerShell" $env:TOKENS_API_KEY = "tok_live_your_key" $env:ANTHROPIC_BASE_URL = "https://tokens.bd" $env:ANTHROPIC_AUTH_TOKEN = $env:TOKENS_API_KEY $env:ANTHROPIC_MODEL = "deepseek/deepseek-v4.1-flash" $env:ANTHROPIC_DEFAULT_OPUS_MODEL = "deepseek/deepseek-v4.1-flash" $env:ANTHROPIC_DEFAULT_SONNET_MODEL = "deepseek/deepseek-v4.1-flash" $env:ANTHROPIC_DEFAULT_HAIKU_MODEL = "deepseek/deepseek-v4.1-flash" $env:CLAUDE_CODE_SUBAGENT_MODEL = "deepseek/deepseek-v4.1-flash" $env:CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS = "1" ``` ::: `$env:` lasts for the current PowerShell window. Use `setx NAME "value"` for each variable to keep it in new windows. ### Option B: ~/.claude/settings.json Values in the settings file's `env` block win over shell exports. On Windows the file is `%USERPROFILE%\.claude\settings.json`. ```json title="~/.claude/settings.json" { "env": { "ANTHROPIC_BASE_URL": "https://tokens.bd", "ANTHROPIC_AUTH_TOKEN": "tok_live_your_key", "ANTHROPIC_MODEL": "deepseek/deepseek-v4.1-flash", "ANTHROPIC_DEFAULT_OPUS_MODEL": "deepseek/deepseek-v4.1-flash", "ANTHROPIC_DEFAULT_SONNET_MODEL": "deepseek/deepseek-v4.1-flash", "ANTHROPIC_DEFAULT_HAIKU_MODEL": "deepseek/deepseek-v4.1-flash", "CLAUDE_CODE_SUBAGENT_MODEL": "deepseek/deepseek-v4.1-flash", "CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS": "1", "CLAUDE_CODE_MAX_CONTEXT_TOKENS": "128000", "CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1" } } ``` :::warning[Keep the key out of project files] Only put the key in your user-level `~/.claude/settings.json`, never in a repository's `.claude/settings.json`. To keep it out of files entirely, drop `ANTHROPIC_AUTH_TOKEN` from the JSON and export it in your shell profile as in Option A. ::: `128000` is a placeholder. Set `CLAUDE_CODE_MAX_CONTEXT_TOKENS` to the real context window shown on the model's page in [/models](/models). What the variables do: | Variable | Purpose | | ------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | | `ANTHROPIC_AUTH_TOKEN` | Sends `Authorization: Bearer`. `ANTHROPIC_API_KEY` also works (sends `x-api-key`) but asks for a one-time approval in interactive mode. | | `ANTHROPIC_MODEL` | Model for the session. | | `ANTHROPIC_DEFAULT_OPUS/SONNET/HAIKU_MODEL` | Back the `opus`, `sonnet` and `haiku` aliases. Haiku also runs background tasks. | | `CLAUDE_CODE_SUBAGENT_MODEL` | Model for subagents. | | `CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS` | Strips most Claude-only beta fields from requests. | ### VS Code extension Set the same variables in your VS Code user settings: ```json title="settings.json (VS Code)" { "claudeCode.environmentVariables": [ { "name": "ANTHROPIC_BASE_URL", "value": "https://tokens.bd" }, { "name": "ANTHROPIC_AUTH_TOKEN", "value": "tok_live_your_key" }, { "name": "ANTHROPIC_MODEL", "value": "deepseek/deepseek-v4.1-flash" } ] } ``` ## Switch models Change `ANTHROPIC_MODEL` and the four alias variables to another id and restart Claude Code. Keep them in sync, or the `opus` alias and background tasks keep running the old model. Copy exact ids from [/models](/models), from `node tokens.mjs models`, or from `GET /v1/models`. [Choosing a model](/docs/choosing-a-model) covers which models suit agent work. The `/model` picker won't list Tokens models on its own: Claude Code's gateway model discovery keeps only ids containing `claude` or `anthropic`. To add one extra entry to the picker, set `ANTHROPIC_CUSTOM_MODEL_OPTION` to a model id. ## Verify it works Test the endpoint directly first (macOS, Linux, WSL or Git Bash): ```bash curl -sS -w '\n%{http_code}\n' -X POST "https://tokens.bd/v1/messages" \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "anthropic-version: 2023-06-01" -H "content-type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 1, "messages": [{"role": "user", "content": "."}]}' ``` A `200` means the key and model are fine. Then run `claude`, type `/status`, and check that the base URL line shows `https://tokens.bd` and the auth token line names `ANTHROPIC_AUTH_TOKEN`. `claude doctor` validates your settings files. The request should also appear in Usage analytics in the dashboard. ## Troubleshooting **404 on every request.** The base URL ends in `/v1`. Claude Code appends `/v1/messages` itself, so use `https://tokens.bd`. **401 `invalid_api_key` even though you exported a new key.** The `env` block in `~/.claude/settings.json` overrides your shell. Check which value `/status` reports. **400 errors about `thinking`, effort or unknown fields.** For model ids it doesn't recognize, Claude Code sends adaptive thinking, effort and beta tool fields. `CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1` removes most of them, though not thinking or effort; Claude Code retries without those when the upstream rejects them. If a model still fails, try another one. **The session never compacts, or hits "prompt too long".** Claude Code assumes 200K tokens of context for unknown ids. Set `CLAUDE_CODE_MAX_CONTEXT_TOKENS` to the model's real window. `CLAUDE_CODE_AUTO_COMPACT_WINDOW` (minimum 100,000) and `CLAUDE_CODE_MAX_OUTPUT_TOKENS` give finer control. **Unexpected model in your usage log.** Background tasks use `ANTHROPIC_DEFAULT_HAIKU_MODEL`, or the main model if it's unset. **Features missing.** Remote Control and voice dictation are disabled while a gateway credential is set, and fast mode checks Anthropic's API directly, so it won't work through Tokens. **403 `model_not_allowed_on_key` or `tier_permission_denied`, 402, or 429 `window_exhausted`.** These are account limits, not Claude Code problems. See [Troubleshooting](/docs/troubleshooting) and include the `x-tokens-request-id` header value if you open a ticket. --- # Codex CLI > Add Tokens as a custom model provider in Codex CLI's config.toml, keep the key in an environment variable, and handle models whose upstream doesn't support the Responses API. Section: Coding Agents. Page: https://tokens.bd/docs/codex-cli Codex CLI is OpenAI's open-source terminal coding agent. It only speaks the OpenAI Responses API, so with Tokens it uses `https://tokens.bd/v1` and sends every request to [/v1/responses](/docs/responses). :::warning[Responses support depends on the model] Tokens forwards `/v1/responses` as is. Whether a request succeeds depends on whether the provider serving that model implements the Responses API, including streaming events and tool calls. Codex works best with models whose upstream supports Responses. If requests fail with one model, try another before changing anything else. ::: ## Quick setup with the Tokens CLI The [Tokens CLI](/docs/tokens-cli) handles sign-in, key creation and model choice, then writes the Codex provider for you. It needs Node 18 or newer. :::code-tabs ```bash title="macOS / Linux" curl -fsSL https://tokens.bd/cli/tokens.mjs -o tokens.mjs && node tokens.mjs setup --base-url https://tokens.bd --agents codex ``` ```powershell title="Windows PowerShell" iwr https://tokens.bd/cli/tokens.mjs -OutFile tokens.mjs; node tokens.mjs setup --base-url https://tokens.bd --agents codex ``` ::: It adds a marked block to `~/.codex/config.toml` (or `$CODEX_HOME/config.toml`) and leaves the rest of the file alone. Running setup again replaces the block in place: ```toml title="~/.codex/config.toml (written by the Tokens CLI)" [model_providers.tokens] name = "Tokens" base_url = "https://tokens.bd/v1" env_key = "TOKENS_API_KEY" wire_api = "responses" [profiles.tokens] model = "deepseek/deepseek-v4.1-flash" model_provider = "tokens" ``` The file holds no secret. Codex reads the key from `TOKENS_API_KEY`, so export it (next section) and start Codex with `codex --profile tokens`. :::note Current Codex docs describe profiles as separate files (`~/.codex/.config.toml`). If `codex --profile tokens` reports that the profile doesn't exist, use the manual setup below, which makes Tokens the default provider and needs no profile. ::: ## Export TOKENS_API_KEY Create a key at [/dashboard/keys](/dashboard/keys) if the CLI didn't make one for you ([API keys](/docs/api-keys) explains caps and allow-lists). :::code-tabs ```bash title="macOS / Linux" # add to ~/.bashrc or ~/.zshrc to keep it export TOKENS_API_KEY="tok_live_your_key" ``` ```powershell title="Windows PowerShell" $env:TOKENS_API_KEY = "tok_live_your_key" # this window setx TOKENS_API_KEY "tok_live_your_key" # new windows ``` ::: `setx` only affects terminals opened afterwards, so open a new one before running `codex`. ## Configure ~/.codex/config.toml manually Install Codex if needed: ```bash npm install -g @openai/codex ``` On Windows you can also use `powershell -ExecutionPolicy ByPass -c "irm https://chatgpt.com/codex/install.ps1 | iex"`. Codex reads `config.toml` from `CODEX_HOME`, which defaults to `~/.codex` (`%USERPROFILE%\.codex` on Windows). Add the provider and select it with the top-level `model_provider` and `model` keys: ```toml title="~/.codex/config.toml" model = "deepseek/deepseek-v4.1-flash" model_provider = "tokens" # model_context_window = 128000 # optional; placeholder, use the real value [model_providers.tokens] name = "Tokens" base_url = "https://tokens.bd/v1" env_key = "TOKENS_API_KEY" wire_api = "responses" ``` Top-level keys must come before any `[table]` in TOML, so put `model` and `model_provider` near the top of the file. A few rules from the Codex docs: - The provider id can't be one of the built-ins (`openai`, `ollama`, `lmstudio`). `tokens` is fine. - `wire_api` must be `"responses"`. The old `"chat"` value is no longer accepted. - If you set `model_context_window`, take the real number from the model's page in [/models](/models). `128000` above is a placeholder. ## Switch models Change `model` in `config.toml` (or under `[profiles.tokens]` if you use the CLI's block) and restart Codex. Get exact ids from [/models](/models), `node tokens.mjs models`, or `GET /v1/models`. Because Responses support varies by upstream, test a new model with the curl call below before committing to it. [Choosing a model](/docs/choosing-a-model) has guidance on picking one for agent work. ## Verify it works Check that the model answers on the Responses endpoint first: ```bash curl https://tokens.bd/v1/responses \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "input": "Reply with OK", "max_output_tokens": 20}' ``` Then run a one-shot task through Codex: ```bash codex exec "Reply with OK" ``` Interactively, start `codex` (or `codex --profile tokens`) and check the model and provider in the status line. ## Troubleshooting **`Error loading config.toml: wire_api = "chat" is no longer supported`.** Change `wire_api` to `"responses"`. Codex removed Chat Completions support. **400 or 404 from upstream on `/v1/responses`.** The provider behind that model doesn't implement the Responses API, or doesn't implement part of it that Codex uses. Pick another model. The same model may still work fine from tools that use chat completions. **401 `missing_api_key`.** `TOKENS_API_KEY` isn't set in the shell that launched Codex. After `setx`, open a new terminal. On macOS/Linux, check that your profile file exports it. **403 `model_not_allowed_on_key`.** The key has an allow-list that doesn't include this model. Allow-lists can't be edited, so create a new key. **TOML parse error after editing.** A top-level key placed after a `[table]` header belongs to that table. Move `model` and `model_provider` above the first table. **The Tokens CLI skipped Codex.** It refuses to touch a file that already has its own `[model_providers.tokens]` or `[profiles.tokens]` table outside its marked block. Remove or rename yours, then run setup again. For 402, 429 and 5xx codes, see [Troubleshooting](/docs/troubleshooting). Include the `x-tokens-request-id` response header in support tickets. --- # OpenCode > Add Tokens to OpenCode as an OpenAI-compatible provider in opencode.json, keep the key in TOKENS_API_KEY, set context limits so compaction works, and switch models. Section: Coding Agents. Page: https://tokens.bd/docs/opencode OpenCode is an open-source terminal coding agent built on the AI SDK provider packages. It connects to Tokens through `@ai-sdk/openai-compatible`, which sends chat completions to `https://tokens.bd/v1`. ## Quick setup with the Tokens CLI The [Tokens CLI](/docs/tokens-cli) signs you in through the browser, creates a key, lets you pick a default model and adds a `tokens` provider to OpenCode. It needs Node 18 or newer. :::code-tabs ```bash title="macOS / Linux" curl -fsSL https://tokens.bd/cli/tokens.mjs -o tokens.mjs && node tokens.mjs setup --base-url https://tokens.bd --agents opencode ``` ```powershell title="Windows PowerShell" iwr https://tokens.bd/cli/tokens.mjs -OutFile tokens.mjs; node tokens.mjs setup --base-url https://tokens.bd --agents opencode ``` ::: It merges the provider into `~/.config/opencode/opencode.json`, sets `"model": "tokens/"`, and backs up the old file as `opencode.json.tokens-backup-`. Two things to know: - The key is written into the file as `options.apiKey`. That's fine for a file in your home folder; don't copy it into a repository. - It doesn't set `limit` on the model. Add real context and output limits afterwards (see below), or OpenCode can't compact long sessions properly. If your `opencode.json` contains comments, the CLI won't risk rewriting it and points you to [/dashboard/connect](/dashboard/connect) for the snippet instead. ## Configure opencode.json manually Install OpenCode if you haven't: ```bash curl -fsSL https://opencode.ai/install | bash # or npm install -g opencode-ai ``` Export your key ([API keys](/docs/api-keys) explains how to create one with a spend cap): :::code-tabs ```bash title="macOS / Linux" export TOKENS_API_KEY="tok_live_your_key" ``` ```powershell title="Windows PowerShell" $env:TOKENS_API_KEY = "tok_live_your_key" # this window setx TOKENS_API_KEY "tok_live_your_key" # new windows ``` ::: Then add the provider. The global file is `~/.config/opencode/opencode.json`; a project can also have its own `opencode.json` in the repo root, and `OPENCODE_CONFIG` points at a custom path. OpenCode merges these files rather than replacing one with another. ```json title="~/.config/opencode/opencode.json" { "$schema": "https://opencode.ai/config.json", "provider": { "tokens": { "npm": "@ai-sdk/openai-compatible", "name": "Tokens", "options": { "baseURL": "https://tokens.bd/v1", "apiKey": "{env:TOKENS_API_KEY}" }, "models": { "deepseek/deepseek-v4.1-flash": { "name": "DeepSeek V4.1 Flash", "limit": { "context": 128000, "output": 8192 } } } } }, "model": "tokens/deepseek/deepseek-v4.1-flash" } ``` `{env:TOKENS_API_KEY}` makes OpenCode read the key from the environment, so the file is safe to commit. The `limit` numbers are placeholders: copy the real context window and max output for the model from its page in [/models](/models). The model reference is `provider_id/model_id`. Our model ids already contain a slash, so the full reference becomes `tokens/deepseek/deepseek-v4.1-flash`. ### Generate the config Pick a model and this generator fills in the provider block for you: ::config-generator{format="opencode"} ### Store the key with /connect instead You can also keep the key out of the environment: run `/connect` in the OpenCode TUI, choose "Other", enter the provider id `tokens` and paste the key. OpenCode stores it in `~/.local/share/opencode/auth.json`. The provider id must match the one in your config, and you can then drop `apiKey` from `options`. ## Switch models Add more entries under `provider.tokens.models`, one per model id from [/models](/models) or `GET /v1/models`: ```json "models": { "deepseek/deepseek-v4.1-flash": { "name": "DeepSeek V4.1 Flash", "limit": { "context": 128000, "output": 8192 } }, "moonshotai/kimi-k3": { "name": "Kimi K3", "limit": { "context": 128000, "output": 8192 } } } ``` The second id is an example; check the exact id in [/models](/models). Then switch in the TUI with `/models`, change the top-level `model`, or pass `--model` for a single run. `small_model` takes the same `tokens/...` form if you want a cheaper model for OpenCode's lightweight tasks. [Choosing a model](/docs/choosing-a-model) helps with the choice. ## Verify it works ```bash opencode models tokens opencode run --model tokens/deepseek/deepseek-v4.1-flash "Reply with OK" ``` The first command should list your Tokens models; the second should print a short answer. In the TUI, `/models` shows the same list. ## Troubleshooting **401 `missing_api_key`.** `{env:TOKENS_API_KEY}` resolved to an empty string because the variable isn't set in the shell that started OpenCode. After `setx` on Windows, open a new terminal. **The model isn't listed, or you get 404 `model_not_found`.** Check that the key under `models` matches the Tokens id exactly and that `model` is `tokens/` followed by that id. OpenCode's docs don't explicitly cover model ids that contain a slash; the pattern is the same one OpenRouter-style ids use, so a typo is the more likely cause. **Wrong package.** Use `@ai-sdk/openai-compatible`, which speaks chat completions. `@ai-sdk/openai` targets the Responses API, which not every upstream supports (see [Responses API](/docs/responses)). **Long sessions fail instead of compacting.** Set `limit.context` and `limit.output` to the model's real values. **`/connect` key ignored.** The provider id you typed in `/connect` must be exactly `tokens`. **The CLI skipped OpenCode.** Your file has comments or invalid JSON. Copy the block above in by hand. For 402, 403 and 429 errors, see [Troubleshooting](/docs/troubleshooting). --- # OpenClaw > Connect OpenClaw to Tokens as a custom provider, either with non-interactive onboarding or by editing ~/.openclaw/openclaw.json, then set real context limits and switch models. Section: Coding Agents. Page: https://tokens.bd/docs/openclaw OpenClaw is an open-source personal AI assistant that runs locally as a daemon and talks to you through chat apps and its own UI. With Tokens it uses the OpenAI Chat Completions protocol (`openai-completions`) against `https://tokens.bd/v1`, or the Anthropic Messages protocol against `https://tokens.bd` if you prefer. The [Tokens CLI](/docs/tokens-cli) doesn't configure OpenClaw, so use OpenClaw's own onboarding or edit its config file. Both take a couple of minutes. ## Install OpenClaw :::code-tabs ```bash title="macOS / Linux / WSL2" curl -fsSL https://openclaw.ai/install.sh | bash ``` ```powershell title="Windows PowerShell" iwr -useb https://openclaw.ai/install.ps1 | iex ``` ::: With npm (Node 24.16+ or 26.1+), run `npm install -g openclaw@latest --allow-scripts=openclaw` and then `openclaw onboard --install-daemon`. ## Export TOKENS_API_KEY Create a key at [/dashboard/keys](/dashboard/keys) ([API keys](/docs/api-keys) covers spend caps), then export it: :::code-tabs ```bash title="macOS / Linux" export TOKENS_API_KEY="tok_live_your_key" ``` ```powershell title="Windows PowerShell" $env:TOKENS_API_KEY = "tok_live_your_key" # this window setx TOKENS_API_KEY "tok_live_your_key" # new windows ``` ::: OpenClaw runs as a background daemon, so make sure the variable is visible to the daemon, not only to your current terminal. Putting it in your shell profile (or using `setx` on Windows) and restarting the daemon is the simplest way. ## Option A: non-interactive onboarding One command registers Tokens as a custom provider and sets the default model: :::code-tabs ```bash title="macOS / Linux" openclaw onboard --non-interactive --accept-risk --skip-health \ --mode local \ --auth-choice custom-api-key \ --custom-base-url "https://tokens.bd/v1" \ --custom-model-id "deepseek/deepseek-v4.1-flash" \ --custom-api-key "$TOKENS_API_KEY" \ --custom-provider-id "tokens" \ --custom-compatibility openai ``` ```powershell title="Windows PowerShell" openclaw onboard --non-interactive --accept-risk --skip-health ` --mode local ` --auth-choice custom-api-key ` --custom-base-url "https://tokens.bd/v1" ` --custom-model-id "deepseek/deepseek-v4.1-flash" ` --custom-api-key "$env:TOKENS_API_KEY" ` --custom-provider-id "tokens" ` --custom-compatibility openai ``` ::: `--custom-compatibility openai` means chat completions, which is what you want. The other values are `openai-responses` and `anthropic`. Running `openclaw onboard` without flags offers the same custom-provider option interactively. ## Option B: edit ~/.openclaw/openclaw.json OpenClaw reads a JSON5 config from `~/.openclaw/openclaw.json` (comments and trailing commas are allowed). `OPENCLAW_CONFIG_PATH` points it elsewhere. The gateway watches the file and reloads changes without a restart. ```jsonc title="~/.openclaw/openclaw.json" { "agents": { "defaults": { "model": { "primary": "tokens/deepseek/deepseek-v4.1-flash" }, "models": { "tokens/deepseek/deepseek-v4.1-flash": { "alias": "Tokens Flash" } }, }, }, "models": { "mode": "merge", "providers": { "tokens": { "baseUrl": "https://tokens.bd/v1", "apiKey": "${TOKENS_API_KEY}", "api": "openai-completions", "timeoutSeconds": 300, "models": [ { "id": "deepseek/deepseek-v4.1-flash", "name": "DeepSeek V4.1 Flash (Tokens)", "reasoning": false, "input": ["text"], "cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 }, "contextWindow": 128000, // placeholder: use the real value "maxTokens": 8192, // placeholder: use the real value }, ], }, }, }, } ``` Replace `contextWindow` and `maxTokens` with the real limits from the model's page in [/models](/models). The `cost` fields are only OpenClaw's local estimate; Tokens bills from its own meter, so zeros are fine. The default model is written as `provider/model-id`. Our ids already contain a slash, which OpenClaw handles: its own docs use `lmstudio/openai/gpt-oss-20b` as an example. To add the provider without rewriting the file, use `openclaw config set models.providers.tokens '' --strict-json --merge`. ### Anthropic-compatible variant If you'd rather use the Messages API, set `api: "anthropic-messages"` and `baseUrl: "https://tokens.bd"` (no `/v1`). OpenClaw doesn't send its implicit `anthropic-beta` headers to non-Anthropic hosts; if you need them, set `models.providers.tokens.headers["anthropic-beta"]`. ## Switch models Add another object to the provider's `models` array for each Tokens model you want, using exact ids from [/models](/models) or `GET /v1/models`. Then change the default: ```bash openclaw models set tokens/deepseek/deepseek-v4.1-flash ``` or edit `agents.defaults.model.primary`. [Choosing a model](/docs/choosing-a-model) covers which models suit agent work. ## Verify it works ```bash openclaw models list openclaw models status openclaw models status --probe ``` `--probe` sends a real request, so it uses a few tokens. A successful probe should also show in your Usage analytics on the dashboard. ## Troubleshooting **Requests go to the wrong endpoint.** Leaving `api` out on a provider with a `baseUrl` defaults to `openai-completions`, which is correct. Only use `openai-responses` if you know the model's upstream supports [/v1/responses](/docs/responses). **Long conversations fail.** If `contextWindow` is missing, OpenClaw assumes 200,000 tokens. If `maxTokens` is missing, it sends no output limit at all. Set both to the model's real values. **401 `missing_api_key`.** The daemon can't see `TOKENS_API_KEY`. Export it where the daemon starts, or put the key directly in `apiKey` in your local config (never in a shared or committed file). **Requests look different from OpenAI's.** On hosts other than api.openai.com, OpenClaw turns off the `developer` role (`compat.supportsDeveloperRole: false`) and skips OpenAI-only fields like `service_tier`, `store` and prompt-cache hints. That's expected. **Need vendor-specific request fields.** Put them under `agents.defaults.models["tokens/"].params.extra_body`. **Missing usage numbers in streams.** Only set `compat.supportsUsageInStreaming: true` if you've confirmed the stream includes usage; Tokens sends a usage chunk on OpenAI-style streams only when the client asks for it with `stream_options.include_usage`. For 402, 403 and 429 errors, see [Troubleshooting](/docs/troubleshooting). --- # Hermes Agent > Point Nous Research's Hermes Agent at Tokens as a custom chat completions endpoint, keep the key in ~/.hermes/.env, set a context length of at least 64K, and switch models. Section: Coding Agents. Page: https://tokens.bd/docs/hermes-agent Hermes Agent is Nous Research's open-source, self-improving agent with a CLI, a TUI and a messaging gateway for Telegram, Discord and others. It connects to Tokens as a custom endpoint over OpenAI-style chat completions at `https://tokens.bd/v1`. :::warning[Hermes needs at least 64K tokens of context] Hermes refuses to start with a model whose context window is under 64,000 tokens, and it needs a model with OpenAI-style tool calling. Check the context window on the model's page in [/models](/models) before you pick one. ::: The [Tokens CLI](/docs/tokens-cli) doesn't configure Hermes, so use Hermes's own wizard or edit its config file. ## Install Hermes Agent :::code-tabs ```bash title="Linux / macOS / WSL2" curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash source ~/.bashrc # or ~/.zshrc ``` ```powershell title="Windows PowerShell" iex (irm https://hermes-agent.nousresearch.com/install.ps1) ``` ::: Create a key at [/dashboard/keys](/dashboard/keys) if you don't have one. [API keys](/docs/api-keys) explains spend caps and model allow-lists. ## Option A: the hermes model wizard The Hermes docs recommend the interactive wizard. Run it from your terminal, outside a chat session: ```bash hermes model ``` Choose **Custom endpoint (self-hosted / VLLM / etc.)** and enter: | Prompt | Value | | ------------ | ------------------------------ | | API base URL | `https://tokens.bd/v1` | | API key | your `tok_live_...` key | | Model name | `deepseek/deepseek-v4.1-flash` | Afterwards, open `~/.hermes/config.yaml` and add `context_length` as shown below, so Hermes doesn't have to guess. ## Option B: edit ~/.hermes/config.yaml The Hermes docs call `config.yaml` the single source of truth. Put the key in `~/.hermes/.env` and reference it by name: ```ini title="~/.hermes/.env" TOKENS_API_KEY=tok_live_your_key ``` If you'd rather export it in your shell, that works too: :::code-tabs ```bash title="macOS / Linux" export TOKENS_API_KEY="tok_live_your_key" ``` ```powershell title="Windows PowerShell" $env:TOKENS_API_KEY = "tok_live_your_key" # this window setx TOKENS_API_KEY "tok_live_your_key" # new windows ``` ::: Then set the model block: ```yaml title="~/.hermes/config.yaml" model: default: deepseek/deepseek-v4.1-flash provider: custom base_url: https://tokens.bd/v1 key_env: TOKENS_API_KEY context_length: 128000 # placeholder: use the model's real window, at least 64000 ``` `context_length` is a placeholder. Copy the real value from [/models](/models). Hermes looks for the context size in your config first, then the endpoint's model list, then its own defaults, so setting it explicitly removes the guesswork. ### Keep Tokens as one named provider If you use Hermes with several providers, declare Tokens under `providers:` instead. The older `custom_providers:` list migrates to this form automatically. ```yaml title="~/.hermes/config.yaml" providers: tokens: api: https://tokens.bd/v1 key_env: TOKENS_API_KEY transport: chat_completions default_model: deepseek/deepseek-v4.1-flash context_length: 128000 # placeholder ``` Hermes also has an `anthropic_messages` transport. We haven't confirmed whether it appends `/v1/messages` to the base URL itself, so stick with `chat_completions` unless you have a reason to change it. ## Switch models Change `model.default` (or `default_model` under your named provider) to another id from [/models](/models) or `GET /v1/models`, and update `context_length` to match the new model. Run `hermes model` again if you prefer the wizard. Inside a session, the documented syntax for a named provider is: ```text /model custom:tokens:deepseek/deepseek-v4.1-flash ``` The format is `/model custom::`. Hermes's docs don't say how it parses a model id that itself contains a slash, so if this doesn't switch, change the model in `config.yaml` instead. `/model` inside a session can only switch between providers that are already configured. [Choosing a model](/docs/choosing-a-model) helps you pick one. ## Verify it works ```bash hermes status hermes doctor hermes chat --oneshot -q "Reply with OK" ``` The last command answers once and exits. When Hermes starts a session it prints a `Context limit: X tokens` line; check that it matches the model's real window. ## Troubleshooting **Hermes refuses to start because of context size.** The configured or detected context is under 64,000 tokens. Pick a model with a larger window, and set `context_length` explicitly. **`OPENAI_BASE_URL` or `LLM_MODEL` has no effect.** `LLM_MODEL` was removed. `OPENAI_BASE_URL` is only read for the `openai-api` provider, not `custom`. `CUSTOM_BASE_URL` survives only as a legacy fallback. Set everything in `config.yaml`. **401 `missing_api_key`.** `key_env` names a variable Hermes can't see. Check the spelling in `~/.hermes/.env`, or that it's exported in the shell that started Hermes. **Tool calls fail or the agent loops.** Hermes depends on OpenAI tool calling. Switch to another model; [Choosing a model](/docs/choosing-a-model) covers what to look for in an agent model. **Unexpected usage from side tasks.** Vision, web summarization, compression and title generation use the main model by default. You can point them at a different model under `auxiliary.*` in `config.yaml`. **403 `model_not_allowed_on_key` or 429 `window_exhausted`.** These are key and plan limits. See [Troubleshooting](/docs/troubleshooting) for every error code. --- # Crush > Add Tokens to Charm's Crush as an openai-compat provider, using the new crushrc format or the older crush.json the Tokens CLI writes, then pick large and small models. Section: Coding Agents. Page: https://tokens.bd/docs/crush Crush is Charm's open-source terminal coding agent. It connects to Tokens as an `openai-compat` provider, sending chat completions to `https://tokens.bd/v1`. ## Quick setup with the Tokens CLI The [Tokens CLI](/docs/tokens-cli) signs you in through the browser, creates a key, lets you choose a default model and registers Tokens in Crush. It needs Node 18 or newer. :::code-tabs ```bash title="macOS / Linux" curl -fsSL https://tokens.bd/cli/tokens.mjs -o tokens.mjs && node tokens.mjs setup --base-url https://tokens.bd --agents crush ``` ```powershell title="Windows PowerShell" iwr https://tokens.bd/cli/tokens.mjs -OutFile tokens.mjs; node tokens.mjs setup --base-url https://tokens.bd --agents crush ``` ::: The CLI merges a `tokens` provider into `~/.config/crush/crush.json` and backs up the old file as `crush.json.tokens-backup-`. What it writes: ```json title="~/.config/crush/crush.json (written by the Tokens CLI)" { "$schema": "https://charm.land/crush.json", "providers": { "tokens": { "name": "Tokens", "type": "openai-compat", "base_url": "https://tokens.bd/v1", "api_key": "tok_live_your_key", "models": [ { "id": "deepseek/deepseek-v4.1-flash", "name": "deepseek/deepseek-v4.1-flash", "context_window": 128000, "default_max_tokens": 8192 } ] } } } ``` Three things to know about it: - The key is stored in the file. Keep this file out of any repository. - `context_window` and `default_max_tokens` are placeholders. Edit them to the real values from the model's page in [/models](/models). - Crush now calls `crush.json` deprecated, though still supported. It works today; if you'd rather use the current format, set Tokens up in `crushrc` as below and remove the `tokens` provider from `crush.json` so you don't maintain it in two places. After setup, open Crush and choose the Tokens model from the model menu (Ctrl+P). ## Configure crushrc manually Install Crush if you haven't: ```bash brew install charmbracelet/tap/crush # or npm install -g @charmland/crush ``` On Windows, `winget install charmbracelet.crush` or `scoop install crush` also work. Export your key ([API keys](/docs/api-keys) explains how to create one with a spend cap): :::code-tabs ```bash title="macOS / Linux" export TOKENS_API_KEY="tok_live_your_key" ``` ```powershell title="Windows PowerShell" $env:TOKENS_API_KEY = "tok_live_your_key" # this window setx TOKENS_API_KEY "tok_live_your_key" # new windows ``` ::: A `crushrc` is Bash with Crush builtins. Crush looks for one in this order, highest priority first: 1. `./.crushrc` in the project 2. `./crushrc` in the project 3. `~/.config/crush/crushrc` (`%USERPROFILE%\.config\crush\crushrc` on Windows) `CRUSH_GLOBAL_CONFIG` and `CRUSH_GLOBAL_DATA` override the global locations. ```bash title="~/.config/crush/crushrc" provider add tokens --type openai-compat \ --name "Tokens" \ --base-url "https://tokens.bd/v1" \ --api-key "${TOKENS_API_KEY:?set TOKENS_API_KEY}" model add tokens/deepseek/deepseek-v4.1-flash \ --name "DeepSeek V4.1 Flash" \ --context-window 128000 \ --default-max-tokens 8192 model large tokens/deepseek/deepseek-v4.1-flash model small tokens/deepseek/deepseek-v4.1-flash ``` `${TOKENS_API_KEY:?...}` makes Crush stop with a clear message if the variable isn't set, instead of sending an empty key. Replace the two numbers with the model's real limits. Use `openai-compat`, not `openai`. Crush's docs reserve the `openai` type for OpenAI itself. ### Let Crush discover models Instead of `model add`, you can add `--discover-models true` to `provider add`. Crush then merges in the models your key can use from `GET /v1/models`. It also discovers automatically when an `openai-compat` provider has no models defined. ## Switch models `model large` sets the main model and `model small` the one Crush uses for lighter work. Point them at any id you've added, then restart Crush. Inside a session, Ctrl+P opens the menu with model switching. To add a model, repeat `model add tokens/` with an exact id from [/models](/models) or `node tokens.mjs models`. In `crush.json`, add another object to the `models` array. [Choosing a model](/docs/choosing-a-model) can help you pick a large and a small model. ## Verify it works ```bash crush models ``` Your Tokens models should be listed in `/` form. Then start `crush`, send a short prompt, and check that it appears in Usage analytics on the dashboard. Logs are in `./.crush/logs/crush.log` in the project folder. ## Troubleshooting **The model id gets split wrongly.** `model add /` separates the provider from the id, and Crush's docs don't say how it handles an id that itself contains a slash, like `tokens/deepseek/deepseek-v4.1-flash`. If the model doesn't show in `crush models`, remove the `model add` lines and use `--discover-models true` instead, then reference the id exactly as `crush models` prints it. **401 `missing_api_key`.** The variable isn't set where Crush runs. With the `:?` form above, Crush stops with "set TOKENS_API_KEY" instead. **Changes don't take effect.** A project `.crushrc` or `crushrc` outranks your global file. Check the project folder. **404 `model_not_found` or 403 `model_not_allowed_on_key`.** The id is misspelled, or the key's allow-list excludes the model. Allow-lists can't be edited, so create a new key if you need a different set. **The Tokens CLI skipped Crush.** Your `crush.json` has comments or invalid JSON, so the CLI didn't touch it. Use the `crushrc` setup above, or copy the snippet from [/dashboard/connect](/dashboard/connect). :::warning[Config files run as code] A `crushrc` is executed as shell, and `$(...)` inside `crush.json` runs when Crush loads it. Only use config files you wrote or have read, especially project-level ones in cloned repositories. ::: For other error codes, see [Troubleshooting](/docs/troubleshooting). --- # Connect Cursor to Tokens > Point Cursor's chat at the Tokens OpenAI-compatible endpoint with a custom OpenAI base URL, and know which Cursor features will keep using Cursor's own models. Section: Coding Agents. Page: https://tokens.bd/docs/cursor Cursor is an AI code editor built on VS Code. To use Tokens models in Cursor chat, you give Cursor a Tokens key in its OpenAI API key field and override the OpenAI base URL to `https://tokens.bd/v1`. Cursor then speaks the OpenAI Chat Completions protocol to Tokens. :::warning[Not officially documented by Cursor] Cursor's current "Bring your own API key" help page (checked October 2026) covers OpenAI, Anthropic, Google, Azure OpenAI and AWS Bedrock keys only. It does not mention an "Override OpenAI Base URL" setting. The steps below come from third-party reports. The toggle may be missing on your Cursor build or plan, and where it exists it may only apply to some features. Treat this setup as best effort. ::: ## What Tokens can and cannot do inside Cursor Before you spend time on settings, know the limits. These come from Cursor's own docs unless marked otherwise. | Cursor feature | Uses your Tokens key? | | ------------------------ | --------------------------------------------------------------------------- | | Chat with a custom model | Yes, when the override is available | | Tab completion | No. Cursor docs: "Tab completion continues using Cursor's built-in models." | | Agent mode | Limited or ignored, according to third-party reports | Two more points from Cursor's docs: - Cursor routes BYOK requests through its own servers for prompt building. Your endpoint must be reachable from the public internet, which `tokens.bd` is. - On Teams and Enterprise plans, BYOK requests still incur the Cursor Token Rate of $0.25 per million tokens (list price, checked October 2026). That charge is Cursor's, on top of what Tokens bills. If you need an agent that reliably uses your key for everything, install [Cline](/docs/cline), [Kilo Code](/docs/kilo-code) or [Continue](/docs/continue) inside Cursor. Cursor runs VS Code extensions, and those three document custom endpoints officially. ## Set a custom OpenAI base URL in Cursor ### 1. Create a key Create a key in the dashboard. A dedicated key with a monthly spend cap is a good idea for an editor, because you can revoke it without touching your other tools. See [API keys](/docs/api-keys) for the options. ### 2. Enter the settings 1. Open Cursor Settings with `Ctrl+Shift+J` (Windows, Linux) or `Cmd+Shift+J` (macOS), then go to **Models**. 2. Under **OpenAI API Key**, paste your key. 3. Turn on **Override OpenAI Base URL** and enter `https://tokens.bd/v1`. 4. Click **Add Custom Model** and enter the model ID exactly. 5. Save, and make sure the new model is enabled in the model list. The values, in one place: ```text title="Cursor Settings > Models" OpenAI API Key: tok_live_your_key Override OpenAI Base URL: https://tokens.bd/v1 Custom model: deepseek/deepseek-v4.1-flash ``` Cursor stores the key in its own settings. It does not read `TOKENS_API_KEY` from your environment, so there is nothing to export for Cursor itself. The same values appear, with your real key, under [Connect Your Agent](/dashboard/connect). The base URL needs the `/v1` suffix. Without it, Cursor calls a path that the gateway does not serve. ## Switch models Add another custom model for each Tokens model you want, using the exact ID with its provider prefix, for example `deepseek/deepseek-v4.1-flash`. Then pick it from the model dropdown in chat. Find IDs in the [model catalog](/models), or list the ones your key can use: ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` If you are unsure which model fits a coding workload, [Choosing a model](/docs/choosing-a-model) covers the trade-offs. ## Verify it works Test the key outside Cursor first. If this fails, no Cursor setting will fix it. ```bash export TOKENS_API_KEY=tok_live_your_key curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` Then open Cursor chat, select `deepseek/deepseek-v4.1-flash`, and send a short prompt. Within a minute or so the request should show in your dashboard usage analytics. If chat answers but nothing appears in usage, Cursor answered with a different model. ## Troubleshooting **There is no "Override OpenAI Base URL" toggle.** Third-party guides report it is missing on some builds and plans. Cursor's docs don't promise it, so there is no setting to hunt for. Use Cline, Kilo Code or Continue inside Cursor instead. **Claude models in Cursor fail with 422 after enabling the override.** Third-party guides report this. Turn the override off when you want Cursor's built-in Anthropic models, and back on for Tokens. **Agent mode or Tab ignores Tokens.** Expected, per the table above. **401 `invalid_api_key` or `missing_api_key`.** The key field is empty, has extra whitespace, or holds a revoked or rotated key. Rotating stops the old secret immediately. **404 `model_not_found`.** The model ID doesn't match a Tokens ID. With the override on, any model that uses the OpenAI key may be sent to Tokens, including names like `gpt-*` that Tokens may not carry under that ID. Disable those models in Cursor's list, or add the exact Tokens ID. **403 `model_not_allowed_on_key` or `tier_permission_denied`.** The key has an allow-list that excludes the model, or your plan doesn't include it and the wallet is empty. **402 `insufficient_credits`.** Top up or renew in [billing](/dashboard/billing). For rate limits and upstream errors, see [Troubleshooting](/docs/troubleshooting). The request format Cursor uses is documented in [Chat Completions](/docs/chat-completions). Sources: [Cursor API keys help page](https://cursor.com/help/models-and-usage/api-keys), [Cursor API keys docs](https://cursor.com/docs/settings/api-keys), checked October 2026. Base URL override steps from third-party guides ([coderouter.io](https://www.coderouter.io/blog/cursor-override-openai-base-url-claude-fix), [routerplex.com](https://routerplex.com/blog/cursor-override-openai-base-url)), not confirmed by Cursor. --- # Connect Cline to Tokens > Set up the Cline coding agent in VS Code with the OpenAI Compatible provider, the Tokens base URL, your key and a model ID, plus the model settings that matter. Section: Coding Agents. Page: https://tokens.bd/docs/cline Cline is an open-source autonomous coding agent. It runs as a VS Code extension and also ships a CLI and a desktop app. Cline connects to Tokens through its **OpenAI Compatible** provider, which speaks the OpenAI Chat Completions protocol to `https://tokens.bd/v1`. ## Before you start You need: - Cline installed. In VS Code, open the Extensions view and search for "Cline". - A Tokens key. Create one in the dashboard; [API keys](/docs/api-keys) explains spend caps and model allow-lists. Coding agents send many requests per task, so a monthly spend cap on this key is cheap insurance. - A model ID. This guide uses `deepseek/deepseek-v4.1-flash`. Browse others in the [model catalog](/models). ## Configure the OpenAI Compatible provider in Cline Open the Cline panel, then its settings, and fill in these fields: | Field | Value | | --------------------------------- | ------------------------------ | | API Provider | OpenAI Compatible | | Base URL | `https://tokens.bd/v1` | | API Key | `tok_live_your_key` | | Model ID | `deepseek/deepseek-v4.1-flash` | | Use Azure Identity Authentication | Off | The same values, with your real key filled in, are on [Connect Your Agent](/dashboard/connect). Cline keeps the key in its own settings field. It does not read a `TOKENS_API_KEY` environment variable, so paste the key directly. ### Model Configuration Under **Model Configuration**, Cline asks for details it can't discover from the endpoint: - **Context Window size.** Set the model's real context window from the [model catalog](/models). Cline uses this to decide when the conversation is too long. A number that's too high leads to context-length errors from the upstream; too low wastes capacity. - **Max Output Tokens.** The model's output limit, or a lower number if you want shorter replies. - **Computer Use.** This is Cline's tool use switch. Cline is an agent that edits files and runs commands through tool calls, so leave it on for models that support tool calling. - **Image Support.** Turn on only if the model accepts image input. The catalog lists this per model. - **Input/Output Price.** Cline uses these only to show its own cost estimate in the panel. Tokens bills from its own usage meter, and your dashboard is the number that counts. You can copy prices from [the catalog](/models) to make Cline's estimate closer, or leave them at zero. ## Switch models Change the **Model ID** field to another Tokens ID and update the Model Configuration values to match the new model. The context window in particular differs a lot between models. To see which IDs your key can use: ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` If you're choosing between fast, cheap models and stronger reasoning models for agent work, [Choosing a model](/docs/choosing-a-model) has the reasoning. ## Verify it works First confirm the key and model outside the editor: ```bash export TOKENS_API_KEY=tok_live_your_key curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` Then send a simple chat message in the Cline panel, such as "List the files in this folder." A working setup answers and, for that prompt, makes a tool call. The request appears in your dashboard usage analytics shortly after. ## Troubleshooting Cline's own troubleshooting points at three things: the key, the model ID, and the base URL. Here is what each looks like with Tokens. **401 `invalid_api_key`.** The key is wrong, revoked, or was rotated. Rotation stops the old secret immediately, so paste the new one into Cline. **"Model Not Found" or 404 `model_not_found`.** The Model ID must match exactly, including the provider prefix. `deepseek-v4.1-flash` without `deepseek/` will not resolve. **404 or connection errors on every request.** Check the Base URL. It must be `https://tokens.bd/v1` with `/v1` and no trailing `/chat/completions`. **403 `model_not_allowed_on_key`.** The key was created with an allow-list that doesn't include this model. Allow-lists can't be edited after creation, so create a new key. **429 `rate_limited` or `concurrency_limit`.** Agents can burst past the per-minute limit during long tasks. The response carries `Retry-After`; Cline retries, or you can wait and resume. A 429 `window_exhausted` means your plan's usage window is used up until it resets. **402 `insufficient_credits`.** Add funds or renew in [billing](/dashboard/billing). **The agent stalls or loops without editing files.** The model may handle tool calls poorly. Try a model that the catalog lists with tool support. More error codes are in [Troubleshooting](/docs/troubleshooting), and the request shape is in [Chat Completions](/docs/chat-completions). For a terminal agent instead of an editor extension, see [OpenCode](/docs/opencode) or [Claude Code](/docs/claude-code). Source: [Cline docs, OpenAI Compatible provider](https://docs.cline.bot/provider-config/openai-compatible), checked October 2026. --- # Connect Roo Code to Tokens > Configure Roo Code's OpenAI Compatible provider for Tokens. The Roo Code repository was archived in May 2026, so this page also points to maintained alternatives. Section: Coding Agents. Page: https://tokens.bd/docs/roo-code Roo Code is a VS Code extension for agentic coding. It connects to Tokens through its **OpenAI Compatible** provider, which uses the OpenAI Chat Completions protocol against `https://tokens.bd/v1`. :::warning[Roo Code is archived] The `RooCodeInc/Roo-Code` GitHub repository was archived on 2026-05-15. The last release is v3.54.0 from that date, and the docs moved to `roocodeinc.github.io/Roo-Code`. The extension still works, but it gets no fixes. If a VS Code update or an API change breaks it, nobody will patch it. For a new setup, use [Kilo Code](/docs/kilo-code) or [Cline](/docs/cline), which are maintained and connect to Tokens the same way. ::: If you already have Roo Code set up and want to keep it, the steps below work with Tokens. ## Configure the OpenAI Compatible provider in Roo Code ### 1. Get a key and a model ID Create a key in the dashboard; [API keys](/docs/api-keys) covers spend caps and allow-lists. This guide uses `deepseek/deepseek-v4.1-flash`. Other IDs are in the [model catalog](/models). ### 2. Enter the provider settings Open the Roo Code panel, go to its settings, and set: | Field | Value | | ------------ | ------------------------------ | | API Provider | OpenAI Compatible | | Base URL | `https://tokens.bd/v1` | | API Key | `tok_live_your_key` | | Model | `deepseek/deepseek-v4.1-flash` | These match the Roo Code snippet on [Connect Your Agent](/dashboard/connect), which fills in your real key. Roo Code takes the key in its settings field. It doesn't read `TOKENS_API_KEY` from the environment. ### 3. Fill in the model configuration Roo Code can't learn model limits from the endpoint, so it asks: - **Context Window.** Use the model's context window from the [catalog](/models). Roo Code manages conversation length based on this number. - **Max Output Tokens.** The model's output limit or less. - **Image Support.** On only if the model accepts images. - **Computer Use.** Tool use. Leave it on for models that support tool calling. - **Input/Output Price.** Only feeds Roo Code's cost display. Tokens bills from its own meter, and your dashboard usage is the real figure. ## Tool calling is required Per Roo Code's docs, it uses native tool calling exclusively, with no XML fallback. The model you pick must support OpenAI-style function calling. A model without it will answer in chat but fail as soon as Roo Code tries to read a file or run a command. Check the model's capabilities in the [catalog](/models) before switching. ## Switch models Change the **Model** field to another exact Tokens ID and update the context window and output values. List the IDs your key can use with: ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` [Choosing a model](/docs/choosing-a-model) explains how to pick one for agent work. ## Verify it works Test the key and model directly: ```bash export TOKENS_API_KEY=tok_live_your_key curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` Then ask Roo Code to do something that needs a tool, such as "Read package.json and tell me the project name." A successful tool call confirms both the connection and function calling. The request should appear in your dashboard usage analytics. ## Troubleshooting **Errors about tools or tool calls.** The model doesn't support native tool calling, or handles it poorly. Pick another model with tool support. **401 `invalid_api_key`.** Wrong, revoked or rotated key. Rotation stops the old secret immediately. **404 `model_not_found`.** The model ID needs the full provider prefix, for example `deepseek/deepseek-v4.1-flash`. **404 on every request.** The base URL must be exactly `https://tokens.bd/v1`. **403 `model_not_allowed_on_key`.** The key's allow-list excludes the model. Allow-lists are fixed at creation, so create a new key. **429 `rate_limited`, `concurrency_limit` or `window_exhausted`.** Per-minute limit, concurrent request limit, or plan usage window. Each response includes `Retry-After`. **Something broke after a VS Code update.** Roo Code won't receive a fix. Move to Kilo Code or Cline; the Tokens settings carry over almost one to one. See [Troubleshooting](/docs/troubleshooting) for every error code and [Chat Completions](/docs/chat-completions) for the request format. Sources: [Roo Code docs, OpenAI Compatible provider](https://roocodeinc.github.io/Roo-Code/providers/openai-compatible) (last updated 2026-05-15), [RooCodeInc/Roo-Code repository](https://github.com/RooCodeInc/Roo-Code) (archived), checked October 2026. --- # Connect Kilo Code to Tokens > Add Tokens to Kilo Code as a custom OpenAI Compatible provider, from the settings UI or a kilo.json file, with model limits set so context management works. Section: Coding Agents. Page: https://tokens.bd/docs/kilo-code Kilo Code is an open-source agentic coding extension for VS Code and JetBrains, with a CLI as well. It connects to Tokens as a **custom provider** using the OpenAI Compatible API, which talks Chat Completions to `https://tokens.bd/v1`. You can set it up in the settings UI or in a `kilo.json` file. This guide was checked against Kilo Code v7.8.3 (October 2026). ## Add Tokens as a custom provider in the UI 1. Click the gear icon, open the **Providers** tab, then click **Custom provider**. 2. Fill in the fields: | Field | Value | | ------------ | ------------------------------------------------------------------------ | | Provider ID | `tokens` | | Display name | `Tokens` | | Provider API | OpenAI Compatible | | Base URL | `https://tokens.bd/v1` | | API key | `tok_live_your_key` | | Models | Fetched from `/v1/models`, or add `deepseek/deepseek-v4.1-flash` by hand | | Headers | Leave empty | 3. Save, then pick the model from Kilo Code's model selector. Because Kilo Code reads `GET /v1/models`, the model list shows only the models your key is allowed to use. If the list comes back empty, your account has no plan or wallet balance yet, or the key's allow-list is narrow. The [Connect Your Agent](/dashboard/connect) page shows the same fields with your key filled in. Create a key first if you don't have one; see [API keys](/docs/api-keys). ## Configure Kilo Code with kilo.json For a setup you can copy between machines, use a config file. Kilo Code reads `~/.config/kilo/kilo.json` globally, or `./kilo.json` in a project. `kilo.jsonc` also works if you want comments. The format follows OpenCode's. ```json title="~/.config/kilo/kilo.json" { "provider": { "tokens": { "npm": "@ai-sdk/openai-compatible", "env": ["TOKENS_API_KEY"], "options": { "baseURL": "https://tokens.bd/v1" }, "models": { "deepseek/deepseek-v4.1-flash": { "name": "DeepSeek V4.1 Flash", "limit": { "context": 128000, "output": 8192 } } } } }, "model": "tokens/deepseek/deepseek-v4.1-flash" } ``` Then export the key in the shell or profile Kilo Code starts from: ```bash export TOKENS_API_KEY=tok_live_your_key ``` The `env` entry keeps the key out of the file, which matters for a project `kilo.json` you commit. Don't put a literal key in a committed config. ### Set real limits The `limit` values above are examples. Replace them with the model's context window and output limit from the [model catalog](/models). Kilo Code's docs say that an omitted `limit.context` or `limit.output` defaults to 0, which limits context management. In practice that means the agent can't tell when to compact a long conversation. ### Which npm package to use The `npm` field selects the protocol: | Package | Protocol | Tokens endpoint | | --------------------------- | ------------------ | ---------------------- | | `@ai-sdk/openai-compatible` | Chat Completions | `/v1/chat/completions` | | `@ai-sdk/openai` | Responses | `/v1/responses` | | `@ai-sdk/anthropic` | Anthropic Messages | `/v1/messages` | Stick with `@ai-sdk/openai-compatible`. Tokens serves `/v1/responses` too, but whether it works for a given model depends on the upstream provider behind it, and Chat Completions is the path every model supports. ## Switch models Add more entries under `models`, each keyed by the exact Tokens ID, and change the top-level `model` to `tokens/`. The model reference is the provider ID, a slash, then the full model ID, so `tokens/deepseek/deepseek-v4.1-flash` contains two slashes. That follows the same pattern as other gateway IDs, though Kilo Code's docs don't state it explicitly. In the UI, pick a different model from the selector. To list the IDs your key can use: ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` [Choosing a model](/docs/choosing-a-model) covers which models suit agent work. ## Verify it works ```bash curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` If that returns a completion, open Kilo Code, select the Tokens model, and ask it to read a file in your project. The request appears in your dashboard usage analytics shortly after. ## Troubleshooting **Model list is empty.** `GET /v1/models` returns nothing for keys without a plan or wallet balance. Subscribe or add funds in [billing](/dashboard/billing), or add the model by hand. **401 `missing_api_key` with kilo.json.** `TOKENS_API_KEY` isn't set in the environment Kilo Code was launched from. Editors started from a desktop launcher may not see variables exported in your shell profile. Restart the editor from a terminal, or enter the key in the UI. **404 `model_not_found`.** The model key under `models` must be the exact Tokens ID with its provider prefix. **Base URL format.** Kilo Code also accepts a full endpoint URL in the Base URL field, but `https://tokens.bd/v1` is the form to use. Don't add `/chat/completions` twice. **Context errors on long tasks.** Check that `limit.context` is set and matches the model. **429 errors.** See `Retry-After`; [Troubleshooting](/docs/troubleshooting) explains `rate_limited`, `concurrency_limit` and `window_exhausted`. Kilo Code's file format is close to [OpenCode](/docs/opencode), so that guide is useful if you run both. Source: [Kilo Code docs, OpenAI Compatible provider](https://kilo.ai/docs/providers/openai-compatible), checked October 2026. --- # Connect Continue to Tokens > Add Tokens models to Continue in VS Code or JetBrains with a config.yaml entry: provider openai, apiBase, key, roles, tool use and context length. Section: Coding Agents. Page: https://tokens.bd/docs/continue Continue is an open-source AI code assistant for VS Code and JetBrains, with a CLI as well. It connects to Tokens with its `openai` provider pointed at a custom `apiBase`, which sends OpenAI Chat Completions requests to `https://tokens.bd/v1`. Everything is set in one file, `config.yaml`. This guide was checked against Continue v2.0.0 for VS Code (October 2026). ## Before you start - Install Continue from the VS Code or JetBrains marketplace (search "Continue"). - Create a Tokens key in the dashboard. [API keys](/docs/api-keys) covers spend caps and allow-lists. - Pick a model ID from the [model catalog](/models). This guide uses `deepseek/deepseek-v4.1-flash`. ## Add Tokens to Continue's config.yaml Open `~/.continue/config.yaml` and add a model entry. The top-level `name`, `version` and `schema` keys are required. ```yaml title="~/.continue/config.yaml" name: Tokens version: 0.0.1 schema: v1 models: - name: DeepSeek V4.1 Flash (Tokens) provider: openai model: deepseek/deepseek-v4.1-flash apiBase: https://tokens.bd/v1 apiKey: ${{ secrets.TOKENS_API_KEY }} roles: [chat, edit, apply] capabilities: [tool_use] defaultCompletionOptions: contextLength: 128000 maxTokens: 8192 ``` What each part does: - `provider: openai` with `apiBase` makes Continue send OpenAI-format requests to Tokens instead of OpenAI. Keep `/v1` on the end of `apiBase`. - `model` is the exact Tokens ID, sent upstream as is. The slash in the ID is fine because Continue treats it as a plain string. - `apiKey` uses Continue's `${{ secrets.NAME }}` syntax, so the key isn't written into the file. If you prefer, you can put the literal `tok_live_your_key` here, but then keep `config.yaml` out of any repository or dotfiles you share. - `roles` decides where the model appears. The allowed values are `chat`, `autocomplete`, `embed`, `rerank`, `edit`, `apply` and `summarize`. - `capabilities: [tool_use]` tells Continue the model can call tools. Agent mode needs it when Continue can't infer tool support for an unfamiliar model ID, which is likely for gateway IDs. - `defaultCompletionOptions` holds the context window and output limit. The numbers above are examples; use the model's limits from the [catalog](/models). The [Connect Your Agent](/dashboard/connect) page generates a shorter version of this entry with your key filled in. The fields it uses are the same. ### Roles to avoid Don't give a Tokens chat model the `embed` or `rerank` roles. Tokens' `/v1/embeddings` endpoint only works for models that are embedding models, so keep whatever embedding provider Continue uses today. Autocomplete is possible with a fast model, but it sends a request on almost every pause in typing, which adds up on a metered key and against the per-minute rate limit. ## Switch models Add one entry per model under `models`, each with its own `name` and `model`. Continue shows every entry in the model dropdown, so you can switch per chat: ```yaml models: - name: DeepSeek V4.1 Flash (Tokens) provider: openai model: deepseek/deepseek-v4.1-flash apiBase: https://tokens.bd/v1 apiKey: ${{ secrets.TOKENS_API_KEY }} roles: [chat, edit, apply] capabilities: [tool_use] - name: Second model (Tokens) provider: openai model: provider/model-id apiBase: https://tokens.bd/v1 apiKey: ${{ secrets.TOKENS_API_KEY }} roles: [chat] ``` Replace `provider/model-id` with a real ID. Check the exact ID in the [catalog](/models) or with `GET /v1/models`. [Choosing a model](/docs/choosing-a-model) helps with the choice. ## Verify it works Check the key and model directly first: ```bash export TOKENS_API_KEY=tok_live_your_key curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` Then pick "DeepSeek V4.1 Flash (Tokens)" in Continue's chat model dropdown and send a message. To check Agent mode, switch to it and ask Continue to read a file. The request shows up in your dashboard usage analytics. ## Troubleshooting **The model doesn't appear in the dropdown.** Continue rejects a `config.yaml` that is missing `name`, `version` or `schema`, or has a YAML indentation error. Fix the file and reload. **The secret doesn't resolve, or you get 401 `missing_api_key`.** Continue couldn't find a value for `secrets.TOKENS_API_KEY`. Check Continue's documentation for how secrets are supplied in your setup, or test with the literal key to rule out everything else. **401 `invalid_api_key`.** The key is wrong, revoked or rotated. Rotation stops the old secret immediately. **404 `model_not_found`.** The `model` value must be the full Tokens ID, for example `deepseek/deepseek-v4.1-flash`. **404 on every request.** Check that `apiBase` is `https://tokens.bd/v1`. Don't set `useLegacyCompletionsEndpoint`; it's only for servers that lack Chat Completions. **Agent mode is unavailable for the model.** Add `capabilities: [tool_use]`. **429 `rate_limited`.** Usually autocomplete. Remove the `autocomplete` role or use a separate key. More codes are in [Troubleshooting](/docs/troubleshooting), and the request format is in [Chat Completions](/docs/chat-completions). If you'd rather work in a terminal, [Aider](/docs/aider) and [OpenCode](/docs/opencode) use the same endpoint. Sources: [Continue docs, OpenAI provider](https://docs.continue.dev/customize/model-providers/top-level/openai), [config.yaml reference](https://docs.continue.dev/reference), checked October 2026. --- # Connect Zed to Tokens > Use Tokens models in Zed's Agent Panel through an OpenAI-compatible provider in settings.json, with the API key kept out of the file. Section: Coding Agents. Page: https://tokens.bd/docs/zed Zed is a fast code editor with a built-in agent, the Zed Agent. It connects to Tokens through an **OpenAI-compatible provider** declared under `language_models.openai_compatible`, which sends Chat Completions requests to `https://tokens.bd/v1`. Zed can also use an Anthropic-compatible provider; that variant is covered near the end. This guide was checked against Zed v1.22.0 (October 2026). :::warning[Keys never go in settings.json] Zed does not read API keys from `settings.json`. Enter the key in the Agent Panel settings, or set it as an environment variable. Anything you put in `settings.json` may end up in a dotfiles repo; keep the key out. ::: ## Add the provider from the Agent Panel The quickest route is the UI: 1. Open Agent Settings. From the command palette, run `agent: open settings` (action `agent::OpenSettings`). 2. Under **LLM Providers**, click **Add Provider**. 3. Enter the provider name `Tokens`, the API URL `https://tokens.bd/v1`, the model ID `deepseek/deepseek-v4.1-flash`, and the context window from the [model catalog](/models). 4. Enter your key, `tok_live_your_key`, in the provider's key field. If you don't have a key yet, create one in the dashboard. See [API keys](/docs/api-keys). ## Declare Tokens in settings.json For a reproducible setup, declare the provider in your Zed `settings.json` (on macOS and Linux, `~/.config/zed/settings.json`): ```json title="~/.config/zed/settings.json" { "language_models": { "openai_compatible": { "Tokens": { "api_url": "https://tokens.bd/v1", "available_models": [ { "name": "deepseek/deepseek-v4.1-flash", "display_name": "DeepSeek V4.1 Flash", "max_tokens": 128000, "max_output_tokens": 8192, "capabilities": { "tools": true, "images": false, "parallel_tool_calls": false, "prompt_cache_key": false, "chat_completions": true, "interleaved_reasoning": false, "max_tokens_parameter": true } } ] } } } } ``` `max_tokens` here is the context window and `max_output_tokens` the output limit. The values shown are examples; use the numbers listed for the model in the [catalog](/models). ### Provide the key Zed builds the environment variable name from the provider name: `_API_KEY`. For a provider called `Tokens`, that's `TOKENS_API_KEY`. ```bash export TOKENS_API_KEY=tok_live_your_key ``` Zed must be started from an environment that has this variable. If you launch Zed from a dock or start menu, it may not see variables exported in your shell profile. Entering the key in the Agent Panel avoids that problem. ### Capability flags Zed's defaults are `tools: true`, `images: false`, `parallel_tool_calls: false`, `prompt_cache_key: false`, `chat_completions: true`, `interleaved_reasoning: false` and `max_tokens_parameter: false`. Three are worth thinking about: - `max_tokens_parameter: true` makes Zed send `max_tokens` instead of `max_completion_tokens`. `max_tokens` is the classic Chat Completions field and the safer choice for a gateway serving many providers, which is why the example turns it on. - `interleaved_reasoning: true` sends earlier thinking back in `reasoning_content`, which DeepSeek-style reasoning models use. Leave it off unless the model you chose is a reasoning model that needs it. - `chat_completions: false` switches Zed to the Responses API. Keep it `true`. Tokens serves `/v1/responses`, but support depends on the upstream for each model, while Chat Completions works for all of them. Set `images: true` only for models the catalog lists with image input. ## Switch models Add more objects to `available_models`, one per Tokens model, then pick between them in the Agent Panel model picker. Use exact IDs from the [catalog](/models) or from: ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` [Choosing a model](/docs/choosing-a-model) explains the trade-offs for agent work. ## Anthropic-compatible variant Zed also supports `language_models.anthropic_compatible`. Set `api_url` to `https://tokens.bd` (no `/v1`; the Anthropic protocol adds the path itself). The capability flags for that provider are `tools`, `images` and `prompt_caching`. Zed sends the key in an `X-Api-Key` header, which Tokens accepts. The OpenAI-compatible setup above is the simpler choice for most models; the Messages API is described in [Messages](/docs/messages). ## Verify it works ```bash curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` If that works, select "DeepSeek V4.1 Flash" in the Agent Panel model picker and send a prompt. Ask the agent to read a file to confirm tool calls work. The request appears in your dashboard usage analytics. ## Troubleshooting **The model is listed but every request fails with 401 `missing_api_key`.** Zed has no key for the provider. Enter it in the Agent Panel, or check that `TOKENS_API_KEY` is set in the environment Zed started from. A key added to `settings.json` is ignored. **401 `invalid_api_key`.** The key is wrong, revoked or rotated. Rotation stops the old secret immediately. **404 `model_not_found`.** `name` must be the exact Tokens ID, including `deepseek/`. `display_name` is only a label. **400 `invalid_request` mentioning a token parameter.** Try toggling `max_tokens_parameter`. **The agent never edits files.** Make sure `tools` is `true` and the model supports tool calling. **429 errors.** Read `Retry-After`. [Troubleshooting](/docs/troubleshooting) explains each rate limit code. Request details are in [Chat Completions](/docs/chat-completions). For a terminal agent alongside Zed, see [Crush](/docs/crush) or [OpenCode](/docs/opencode). Sources: [Zed docs, use your own API access](https://zed.dev/docs/ai/use-api-access), [Zed LLM providers](https://zed.dev/docs/ai/llm-providers), checked October 2026. --- # Connect GitHub Copilot to Tokens > Use Tokens models in GitHub Copilot Chat in VS Code with the bring-your-own-key Custom Endpoint provider and a chatLanguageModels.json entry. Section: Coding Agents. Page: https://tokens.bd/docs/github-copilot GitHub Copilot in VS Code can use models from outside GitHub through its bring-your-own-key (BYOK) **Custom Endpoint** provider. With Tokens, the Custom Endpoint sends OpenAI Chat Completions requests to `https://tokens.bd/v1/chat/completions`, and the models show up in Copilot Chat's model picker. The Custom Endpoint provider replaces the older "OpenAI Compatible" provider and the `github.copilot.chat.customOAIModels` setting, both deprecated. If you configured Tokens that way before, move to the steps below. ## What BYOK covers in Copilot Per VS Code's docs, BYOK models are used for chat and utility tasks only. Inline suggestions, semantic search and embeddings still go through GitHub. Your Copilot subscription still matters for those. On Copilot Business and Enterprise, admins can disable BYOK by policy. If the **Custom Endpoint** option is missing, ask whoever manages your organization's Copilot settings. ## Add a Custom Endpoint for Tokens in VS Code ### 1. Create a key Create a key in the dashboard. A key used only by Copilot, with its own monthly spend cap, is easy to track and revoke. [API keys](/docs/api-keys) explains the options. ### 2. Run the wizard 1. Open the model picker in Copilot Chat, click the gear, and choose **Manage Language Models**. You can also run **Chat: Manage Language Models** from the command palette. 2. Choose **Add Models**, then **Custom Endpoint**. 3. Enter a group name (`Tokens`), a display name, and your API key. 4. Choose the API type **Chat Completions**. VS Code then opens `chatLanguageModels.json`. ### 3. Edit chatLanguageModels.json Make the Tokens entry look like this: ```json title="chatLanguageModels.json" [ { "name": "Tokens", "vendor": "customendpoint", "apiKey": "${input:tokensApiKey}", "apiType": "chat-completions", "models": [ { "id": "deepseek/deepseek-v4.1-flash", "name": "DeepSeek V4.1 Flash (Tokens)", "url": "https://tokens.bd/v1/chat/completions", "toolCalling": true, "vision": false, "maxInputTokens": 120000, "maxOutputTokens": 8192 } ] } ] ``` Notes on the fields: - `id` is the exact Tokens model ID, sent upstream as is. - `url` can be the full endpoint, as here. VS Code uses a URL that already ends in `/chat/completions`, `/responses` or `/messages` without changes. Otherwise it appends the path for the API type, and inserts `/v1` if it's missing. Writing the full URL removes any guesswork. - `apiKey` uses an input variable, so the raw key isn't stored in the JSON file. If you paste a literal key instead, treat this file as a secret. - `toolCalling: true` is needed for Copilot's agent mode. Set it only for models that support tool calling. - `vision` should be `true` only for models with image input. - `maxInputTokens` and `maxOutputTokens` are example values. Use the model's limits from the [model catalog](/models). The key is sent as `Authorization: Bearer`, which Tokens accepts. ## Switch models Add one object to `models` per Tokens model, each with its own `id`, `name` and limits. They appear under the Tokens group in the Copilot Chat model picker. Get exact IDs from the [catalog](/models) or: ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` For help picking, see [Choosing a model](/docs/choosing-a-model). ### Anthropic Messages variant The Custom Endpoint also supports the Anthropic Messages API. Set `"apiType": "messages"` and `"url": "https://tokens.bd/v1/messages"`. With this type VS Code sends the key in `x-api-key`, which Tokens also accepts. The request format is in [Messages](/docs/messages). Chat Completions is the simpler default. ## Verify it works Test the key outside VS Code first: ```bash export TOKENS_API_KEY=tok_live_your_key curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` Then open Copilot Chat, pick "DeepSeek V4.1 Flash (Tokens)" in the model picker, and ask a question. In agent mode, ask it to read a file to confirm tool calling. The request shows in your dashboard usage analytics. ## Troubleshooting **No Custom Endpoint option.** Your VS Code or Copilot Chat version is older, or an organization policy disables BYOK. **401 `invalid_api_key` or `missing_api_key`.** The key wasn't entered or is wrong, revoked or rotated. Rotating a key stops the old secret immediately. Re-run the wizard to enter the new one. **404 `model_not_found`.** The `id` must be the full Tokens ID, such as `deepseek/deepseek-v4.1-flash`. **404 with an odd path.** Check `url`. A doubled path such as `/v1/chat/completions/chat/completions` means you combined a full URL with something that appends a path. Use the full URL exactly as shown. **The model isn't offered in agent mode.** Set `toolCalling: true`. **Inline suggestions don't use Tokens.** Expected; BYOK doesn't cover them. **402 `insufficient_credits` or 429 errors.** Top up in [billing](/dashboard/billing), or see [Troubleshooting](/docs/troubleshooting) for rate limits. Copilot's requests follow [Chat Completions](/docs/chat-completions). If you'd rather have an agent that runs everything through your key, see [Cline](/docs/cline) or [Continue](/docs/continue). Source: [VS Code docs, language models and Custom Endpoint](https://code.visualstudio.com/docs/agent-customization/language-models), checked October 2026. --- # Aider > Run Aider against Tokens with OPENAI_API_BASE and the openai/ model prefix, silence unknown-model warnings with a metadata file, and switch models per session. Section: Coding Agents. Page: https://tokens.bd/docs/aider Aider is an open-source AI pair-programming CLI that edits files in your git repository and commits as it goes. It talks to Tokens through its OpenAI-compatible provider, sending chat completions to `https://tokens.bd/v1`. The [Tokens CLI](/docs/tokens-cli) doesn't configure Aider; it only takes two environment variables and one flag, so there's little to automate. ## Install Aider ```bash python -m pip install aider-install aider-install ``` On macOS and Linux, `curl -LsSf https://aider.chat/install.sh | sh` does the same in one step. Aider's development has slowed: the last release we saw was v0.86.0 from August 2025. It still works with OpenAI-compatible endpoints like Tokens. ## Set OPENAI_API_BASE for Aider Create a key at [/dashboard/keys](/dashboard/keys) ([API keys](/docs/api-keys) covers spend caps). Keep it in `TOKENS_API_KEY` and hand it to Aider through the variables it reads: :::code-tabs ```bash title="macOS / Linux" export TOKENS_API_KEY="tok_live_your_key" export OPENAI_API_BASE="https://tokens.bd/v1" export OPENAI_API_KEY="$TOKENS_API_KEY" ``` ```powershell title="Windows PowerShell" $env:TOKENS_API_KEY = "tok_live_your_key" $env:OPENAI_API_BASE = "https://tokens.bd/v1" $env:OPENAI_API_KEY = $env:TOKENS_API_KEY ``` ::: The PowerShell lines last for the current window. To keep them, run `setx` for each (`setx OPENAI_API_BASE https://tokens.bd/v1`, and so on) and open a new terminal. :::warning[These variables are global] Other tools also read `OPENAI_API_BASE` and `OPENAI_API_KEY`. Setting them in your shell profile sends those tools to Tokens too. If that's not what you want, set them only in the terminal where you run Aider, or wrap Aider in a small script that exports them first. ::: Then start Aider in your project: ```bash cd /path/to/your/project aider --model openai/deepseek/deepseek-v4.1-flash ``` The `openai/` prefix tells Aider to use its OpenAI-compatible provider; Aider's docs say to always add it. Aider runs on LiteLLM, which strips that first `openai/` and sends `deepseek/deepseek-v4.1-flash` to Tokens. That stripping is standard LiteLLM behavior rather than something Aider's page spells out, but it's why the full id still reaches us intact. ### Generate the command Pick a model and this generator builds the exports and the `aider` command for you: ::config-generator{format="aider"} ## Add model metadata Aider doesn't know Tokens' model ids, so it prints a model warning at startup. The warning is harmless. To give Aider real limits and silence it, create `.aider.model.metadata.json` in your home folder, the repo root or the current directory (or pass `--model-metadata-file `): ```json title=".aider.model.metadata.json" { "openai/deepseek/deepseek-v4.1-flash": { "max_input_tokens": 128000, "max_output_tokens": 8192, "input_cost_per_token": 0, "output_cost_per_token": 0, "litellm_provider": "openai", "mode": "chat" } } ``` The token numbers are placeholders: copy the real context window and max output from the model's page in [/models](/models). Leave the costs at zero. Aider's cost display would only be an estimate anyway; your actual spend is in the dashboard and in `node tokens.mjs usage`. ## Switch models Pass a different id after `openai/`: ```bash aider --model openai/ ``` Copy the exact id from [/models](/models) or `GET /v1/models`, and add a matching entry to the metadata file if you want the warning gone for that model too. [Choosing a model](/docs/choosing-a-model) covers which models suit editing work. ## Verify it works Run the `aider --model ...` command above. Aider prints the model it's using on startup. Ask it something small ("What files are in this repo?") and check that it answers and that the request appears in Usage analytics on the dashboard. To rule out Aider itself, test the endpoint directly: ```bash curl https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "messages": [{"role": "user", "content": "Reply with OK"}], "max_tokens": 10}' ``` ## Troubleshooting **Requests never show up in your Tokens usage.** The `openai/` prefix is missing. Without it, Aider (through LiteLLM) may pick a different built-in provider from the first part of the id, and the request goes there instead of Tokens. The command needs `openai/` plus the full Tokens id: `openai/deepseek/deepseek-v4.1-flash`. **404 `model_not_found`.** The id after `openai/` doesn't match a Tokens model. Copy it again from [/models](/models). **404 on every request.** `OPENAI_API_BASE` must end in `/v1`. Without it the path doesn't exist on Tokens. **401 `invalid_api_key` or `missing_api_key`.** `OPENAI_API_KEY` is empty or still holds an OpenAI key. Run `echo $OPENAI_API_KEY` (or `$env:OPENAI_API_KEY` in PowerShell); a Tokens key starts with `tok_live_`. **Context errors on large repos.** Aider reports token limits but never enforces them. Keep the files you add to the chat small, or pick a model with a larger window. **"Model warnings" at startup.** Expected for ids Aider doesn't know. Add the metadata file above. For 402, 403 and 429 errors, see [Troubleshooting](/docs/troubleshooting). --- # Goose > Add Tokens to Goose as a custom OpenAI-compatible provider through goose configure or a JSON file in custom_providers, keep the key in TOKENS_API_KEY, and switch models. Section: Coding Agents. Page: https://tokens.bd/docs/goose Goose is an open-source AI agent that runs on your machine as a CLI and a desktop app; the project now lives at `aaif-goose/goose` (formerly `block/goose`). It connects to Tokens as a custom provider with the `openai` engine, which sends chat completions to `https://tokens.bd/v1/chat/completions`. The [Tokens CLI](/docs/tokens-cli) doesn't configure Goose. Use Goose's own wizard or drop a provider file in place. ## Install Goose ```bash curl -fsSL https://github.com/aaif-goose/goose/releases/download/stable/download_cli.sh | bash # or with Homebrew brew install block-goose-cli # CLI brew install --cask block-goose # desktop app ``` Create a key at [/dashboard/keys](/dashboard/keys). [API keys](/docs/api-keys) explains spend caps and allow-lists. ## Option A: goose configure 1. Run `goose configure`. 2. Choose **Custom Providers**, then **Add A Custom Provider**. 3. Fill in the prompts: | Prompt | Value | | ----------------------- | -------------------------------------------- | | API Type | `OpenAI Compatible` | | Name | `Tokens` | | API URL | `https://tokens.bd/v1/chat/completions` | | Authentication Required | Yes, static API key: your `tok_live_...` key | | Available Models | `deepseek/deepseek-v4.1-flash` | | Streaming Support | Yes | Goose stores the key in your system keychain, or in `secrets.yaml` if no keychain is available. In the desktop app the same form is under Settings > Add Custom Provider. Note the API URL: Goose's custom providers take the full chat completions path, not just `/v1`. ## Option B: a provider file Create `~/.config/goose/custom_providers/tokens.json`. On Windows the folder is `%APPDATA%\Block\goose\config\custom_providers\`. ```json title="~/.config/goose/custom_providers/tokens.json" { "name": "tokens", "engine": "openai", "display_name": "Tokens", "description": "Tokens AI gateway", "api_key_env": "TOKENS_API_KEY", "base_url": "https://tokens.bd/v1/chat/completions", "models": [{ "name": "deepseek/deepseek-v4.1-flash", "context_limit": 128000 }], "supports_streaming": true, "requires_auth": true } ``` `context_limit` is a placeholder; copy the real context window from the model's page in [/models](/models). The file has no secret in it, because `api_key_env` tells Goose to read `TOKENS_API_KEY`: :::code-tabs ```bash title="macOS / Linux" export TOKENS_API_KEY="tok_live_your_key" goose session start --provider tokens ``` ```powershell title="Windows PowerShell" $env:TOKENS_API_KEY = "tok_live_your_key" # this window setx TOKENS_API_KEY "tok_live_your_key" # new windows goose session start --provider tokens ``` ::: ## Option C: the built-in OpenAI provider If you'd rather not add a custom provider, Goose's built-in OpenAI provider can point at Tokens. Set these as environment variables (`export` in bash, `$env:` in PowerShell): ```ini GOOSE_PROVIDER=openai OPENAI_HOST=https://tokens.bd OPENAI_BASE_PATH=v1/chat/completions OPENAI_API_KEY=tok_live_your_key GOOSE_MODEL=deepseek/deepseek-v4.1-flash ``` `OPENAI_HOST` is the bare host with no path; the path goes in `OPENAI_BASE_PATH`. This takes over Goose's OpenAI provider, so if you also use OpenAI directly through Goose, Option B is cleaner. ## Switch models For a custom provider, add more entries to the `models` array (or to "Available Models" in the wizard) with exact ids from [/models](/models) or `GET /v1/models`, each with its own `context_limit`. Then pick the model in the desktop app's model picker, or set `GOOSE_MODEL` for the CLI. With the built-in OpenAI provider, `goose configure` won't accept a custom model name. Set `GOOSE_MODEL`, or edit the `providers:` block in `~/.config/goose/config.yaml`, instead. [Choosing a model](/docs/choosing-a-model) helps with the choice. ## Verify it works ```bash goose session start --provider tokens ``` Ask something small, such as "list the files in this folder". The answer should arrive and the request should appear in Usage analytics on the dashboard. In the desktop app, Tokens shows up in the model picker. To rule out Goose, call the endpoint directly: ```bash curl https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "messages": [{"role": "user", "content": "Reply with OK"}], "max_tokens": 10}' ``` ## Troubleshooting **`No api key passed in`.** Goose ignores API keys written into `config.yaml`. Use the keychain (through `goose configure`), `secrets.yaml`, or the environment variable named in `api_key_env`. **404 on every request.** With a custom provider, `base_url` must be the full `https://tokens.bd/v1/chat/completions`. With the built-in provider, `OPENAI_HOST` must be just `https://tokens.bd` and `OPENAI_BASE_PATH` must be `v1/chat/completions`; a wrong base path is the usual cause. **404 `model_not_found` or 403 `model_not_allowed_on_key`.** The model id is misspelled, or the key's allow-list doesn't include it. Allow-lists can't be edited, so create a new key if needed. **The agent can't use tools.** Goose relies on tool calling. If a model doesn't handle tool calls well, switch to another one. **Long sessions fail.** Set `context_limit` to the model's real window so Goose knows how much room it has. For 402 and 429 errors, see [Troubleshooting](/docs/troubleshooting). --- # Qwen Code > Connect Qwen Code to Tokens through its OpenAI protocol in ~/.qwen/settings.json or with three environment variables, then switch between Tokens models with /model. Section: Coding Agents. Page: https://tokens.bd/docs/qwen-code Qwen Code is Alibaba's terminal coding agent, originally forked from Gemini CLI. It speaks several protocols; with Tokens you use its OpenAI protocol, which sends chat completions to `https://tokens.bd/v1`. :::note[Coming from Gemini CLI?] Gemini CLI can only talk to Gemini-format APIs, so it can't use Tokens. Qwen Code keeps a similar workflow and works with any OpenAI-compatible endpoint, which makes it the practical substitute. ::: The [Tokens CLI](/docs/tokens-cli) doesn't configure Qwen Code, so set it up by hand. It's one file or three variables. ## Install Qwen Code ```bash npm install -g @qwen-code/qwen-code@latest # or brew install qwen-code ``` Create a key at [/dashboard/keys](/dashboard/keys). [API keys](/docs/api-keys) covers spend caps and model allow-lists. ## Export TOKENS_API_KEY :::code-tabs ```bash title="macOS / Linux" export TOKENS_API_KEY="tok_live_your_key" ``` ```powershell title="Windows PowerShell" $env:TOKENS_API_KEY = "tok_live_your_key" # this window setx TOKENS_API_KEY "tok_live_your_key" # new windows ``` ::: ## Option A: ~/.qwen/settings.json Add Tokens models under `modelProviders.openai` and tell Qwen Code to use the OpenAI protocol: ```json title="~/.qwen/settings.json" { "modelProviders": { "openai": [ { "id": "deepseek/deepseek-v4.1-flash", "name": "DeepSeek V4.1 Flash (Tokens)", "baseUrl": "https://tokens.bd/v1", "description": "via Tokens", "envKey": "TOKENS_API_KEY" } ] }, "security": { "auth": { "selectedType": "openai" } }, "model": { "name": "deepseek/deepseek-v4.1-flash" } } ``` `envKey` names the environment variable that holds the key, so the file itself has no secret. Qwen Code also accepts an `env` block in the same file (`"env": { "TOKENS_API_KEY": "tok_live_your_key" }`), but that puts the key in plain text. If you use it, keep the file out of any repository. When the same variable is set in several places, Qwen Code takes the shell export first, then a `.env` file, then the settings `env` block. For context window, custom headers and extra request fields, Qwen Code's Model Providers reference documents `generationConfig`, `customHeaders` and `extra_body`. Set the context window to the real value from the model's page in [/models](/models). ## Option B: environment variables only If you don't want a settings file, these three variables are enough: :::code-tabs ```bash title="macOS / Linux" export OPENAI_API_KEY="$TOKENS_API_KEY" export OPENAI_BASE_URL="https://tokens.bd/v1" export OPENAI_MODEL="deepseek/deepseek-v4.1-flash" ``` ```powershell title="Windows PowerShell" $env:OPENAI_API_KEY = $env:TOKENS_API_KEY $env:OPENAI_BASE_URL = "https://tokens.bd/v1" $env:OPENAI_MODEL = "deepseek/deepseek-v4.1-flash" ``` ::: `QWEN_MODEL` works as an alias for `OPENAI_MODEL`. Be aware that other tools read `OPENAI_API_KEY` and `OPENAI_BASE_URL` too, so setting them in your shell profile redirects those tools to Tokens as well. Option A avoids that. You can also run `/auth` inside `qwen` and enter the same values interactively. ### Anthropic protocol Qwen Code also has an `anthropic` protocol that reads `ANTHROPIC_API_KEY`, `ANTHROPIC_BASE_URL` and `ANTHROPIC_MODEL`. With Tokens that would be `ANTHROPIC_BASE_URL=https://tokens.bd` (no `/v1`). The OpenAI protocol is the simpler path and the one shown above. ## Switch models Add one object per model to the `modelProviders.openai` array, each with the exact Tokens id from [/models](/models) or `GET /v1/models` and the same `baseUrl` and `envKey`: ```json { "id": "moonshotai/kimi-k3", "name": "Kimi K3 (Tokens)", "baseUrl": "https://tokens.bd/v1", "envKey": "TOKENS_API_KEY" } ``` The id above is an example; check the exact one in [/models](/models). Then run `/model` inside `qwen` to switch, or change `model.name` in the settings file. With environment variables, change `OPENAI_MODEL`. [Choosing a model](/docs/choosing-a-model) has guidance for agent work. ## Verify it works Start `qwen` in a project folder, run `/model` to confirm the Tokens model is selected, then send a short prompt. The request should appear in Usage analytics on the dashboard. To rule out Qwen Code itself, call the endpoint directly: ```bash curl https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "messages": [{"role": "user", "content": "Reply with OK"}], "max_tokens": 10}' ``` ## Troubleshooting **401 `invalid_api_key` with a key you just changed.** An older value wins somewhere: a shell export beats `.env`, which beats the settings `env` block. Check `echo $TOKENS_API_KEY` (or `$env:TOKENS_API_KEY`). **401 `missing_api_key`.** The variable named in `envKey` isn't set in the shell that started `qwen`. After `setx` on Windows, open a new terminal. **Requests don't reach Tokens.** `security.auth.selectedType` must be `openai`, or Qwen Code uses whichever auth method was selected before. Set it in the file, or pick the OpenAI option in `/auth`. **404 on every request.** `baseUrl` (or `OPENAI_BASE_URL`) must end in `/v1`. **404 `model_not_found` or 403 `model_not_allowed_on_key`.** The id is misspelled, or the key's allow-list excludes the model. Allow-lists can't be edited after creation, so create a new key if needed. For 402, 429 and 5xx errors, see [Troubleshooting](/docs/troubleshooting). --- # Kimi Code CLI > Add Tokens to Moonshot's Kimi Code CLI as an openai provider in ~/.kimi-code/config.toml, read the key from TOKENS_API_KEY, and map local model aliases to Tokens model ids. Section: Coding Agents. Page: https://tokens.bd/docs/kimi-code Kimi Code CLI is Moonshot AI's terminal coding agent. It connects to Tokens through its `openai` provider type, which sends chat completions to `https://tokens.bd/v1`; an `anthropic` type against `https://tokens.bd` also works. :::warning[The old Kimi CLI is archived] The Python Kimi CLI (`MoonshotAI/kimi-cli`) is archived, and Moonshot says existing installations will stop working. This guide is for its successor, Kimi Code CLI (`MoonshotAI/kimi-code`), which uses a different config file. ::: The [Tokens CLI](/docs/tokens-cli) doesn't configure Kimi Code, so add the provider by hand. It's one short TOML file. ## Install Kimi Code CLI :::code-tabs ```bash title="macOS / Linux" curl -fsSL https://code.kimi.com/kimi-code/install.sh | bash # or: npm install -g @moonshot-ai/kimi-code # or: brew install kimi-code ``` ```powershell title="Windows PowerShell" irm https://code.kimi.com/kimi-code/install.ps1 | iex ``` ::: Create a key at [/dashboard/keys](/dashboard/keys). [API keys](/docs/api-keys) explains spend caps and model allow-lists. ## Export TOKENS_API_KEY :::code-tabs ```bash title="macOS / Linux" export TOKENS_API_KEY="tok_live_your_key" ``` ```powershell title="Windows PowerShell" $env:TOKENS_API_KEY = "tok_live_your_key" # this window setx TOKENS_API_KEY "tok_live_your_key" # new windows ``` ::: ## Configure ~/.kimi-code/config.toml Kimi Code reads `~/.kimi-code/config.toml`; set `KIMI_CODE_HOME` to move it. Add a provider and at least one model: ```toml title="~/.kimi-code/config.toml" default_model = "tokens/deepseek-v4.1-flash" [providers.tokens] type = "openai" base_url = "https://tokens.bd/v1" api_key_env = "TOKENS_API_KEY" [models."tokens/deepseek-v4.1-flash"] provider = "tokens" model = "deepseek/deepseek-v4.1-flash" max_context_size = 128000 capabilities = [ "tool_use" ] ``` How the pieces fit: - `[providers.tokens]` defines the connection. `type = "openai"` means Chat Completions. The other types are `kimi`, `anthropic`, `openai_responses`, `google-genai` and `vertexai`. - `[models."..."]` is a local alias. Its name is what you see in Kimi Code; `model` is the id sent to Tokens and must match our catalog exactly. - The alias contains a `/`, so it must be quoted in the table header. Any TOML key with a `.` in it needs quotes too. - `max_context_size = 128000` is a placeholder. Copy the real context window from the model's page in [/models](/models). Set exactly one of `api_key_env` or `api_key`. Kimi Code doesn't fall back to other environment variables, so if neither resolves, it fails at startup. With `type = "openai"`, Kimi Code handles DeepSeek-style `reasoning_content` in responses automatically. If a model returns its reasoning under a different field, `reasoning_key` lets you name it. ### Anthropic-compatible variant To use the Messages API instead, change the provider: ```toml [providers.tokens] type = "anthropic" base_url = "https://tokens.bd" api_key_env = "TOKENS_API_KEY" ``` Leave `/v1` off: this type follows the Anthropic SDK, which appends `/v1/messages` itself. See [Messages API](/docs/messages) for what that endpoint supports. ### Add the provider interactively You can also run `/provider` inside the TUI, or `kimi provider` from the shell, instead of editing the file. ## Switch models Add one `[models."..."]` table per Tokens model, each pointing at the `tokens` provider: ```toml [models."tokens/kimi-k3"] provider = "tokens" model = "moonshotai/kimi-k3" max_context_size = 128000 capabilities = [ "tool_use" ] ``` The id in `model` is an example; copy the exact one from [/models](/models) or `GET /v1/models`, and set its real context size. Then switch with `/model` inside `kimi`, or change `default_model` to make it the default. [Choosing a model](/docs/choosing-a-model) has guidance for agent work. ## Verify it works Run `kimi` in a project folder, then `/model` to confirm the Tokens alias is active, and send a short prompt. The request should appear in Usage analytics on the dashboard. If no credential resolves, Kimi Code says so loudly at startup rather than failing on the first request. To check the key and model without Kimi Code involved: ```bash curl https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "messages": [{"role": "user", "content": "Reply with OK"}], "max_tokens": 10}' ``` ## Troubleshooting **Startup error about missing credentials.** `TOKENS_API_KEY` isn't set in the shell that started `kimi`, or both `api_key` and `api_key_env` are set. Use exactly one. After `setx` on Windows, open a new terminal. **TOML parse error.** A model alias with `/` or `.` isn't quoted. Write `[models."tokens/deepseek-v4.1-flash"]`, not `[models.tokens/deepseek-v4.1-flash]`. **404 `model_not_found`.** The `model` field, not the alias, must be the exact Tokens id. `deepseek/deepseek-v4.1-flash` is the id; `tokens/deepseek-v4.1-flash` is only your local name. **404 on every request.** For `type = "openai"`, `base_url` must end in `/v1`. For `type = "anthropic"`, it must not. **Tools never get called.** Make sure the model entry lists `capabilities = [ "tool_use" ]`, and that the model you picked supports tool calling. **403 `model_not_allowed_on_key`, 402 or 429.** These are key, balance and plan limits. See [Troubleshooting](/docs/troubleshooting). --- # Connect Factory Droid to Tokens > Add Tokens models to Factory's Droid CLI as custom models in ~/.factory/settings.json, using the Chat Completions or Anthropic Messages protocol. Section: Coding Agents. Page: https://tokens.bd/docs/factory-droid Droid is Factory's terminal coding agent. It supports bring-your-own-key **custom models**, declared in `~/.factory/settings.json`. With Tokens, use the `generic-chat-completion-api` provider type, which sends OpenAI Chat Completions requests to `https://tokens.bd/v1`. An Anthropic Messages variant is covered below. ## Install Droid From Factory's quickstart: ```bash # macOS / Linux curl -fsSL https://app.factory.ai/cli | sh # or brew install --cask droid npm install -g droid ``` ```powershell # Windows PowerShell irm https://app.factory.ai/cli/windows | iex ``` ## Add Tokens custom models to ~/.factory/settings.json ### 1. Export your key Create a key in the dashboard ([API keys](/docs/api-keys) covers caps and allow-lists), then export it in your shell profile: ```bash export TOKENS_API_KEY=tok_live_your_key ``` ### 2. Add the custom model Add a `customModels` array to `~/.factory/settings.json`: ```json title="~/.factory/settings.json" { "customModels": [ { "model": "deepseek/deepseek-v4.1-flash", "displayName": "DeepSeek V4.1 Flash (Tokens)", "baseUrl": "https://tokens.bd/v1", "apiKey": "${TOKENS_API_KEY}", "provider": "generic-chat-completion-api", "maxOutputTokens": 8192, "noImageSupport": true } ] } ``` Field by field: - `model` is the exact Tokens ID, sent upstream unchanged. The slash in the ID is fine. - `displayName` is what you see in Droid's model menu. - `baseUrl` is the OpenAI-compatible base, with `/v1`. - `apiKey` uses `${VAR}` expansion, so the key stays in your environment instead of the file. - `provider` sets the protocol. See the table below. - `maxOutputTokens` is an example; use the model's output limit from the [model catalog](/models). - `noImageSupport: true` stops Droid from sending images to a text-only model. Remove it for models the catalog lists with image input. If the file already has other settings, add `customModels` alongside them rather than replacing the file. ### Pick the provider type The `provider` value must be exactly one of these: | `provider` | Protocol | Tokens endpoint | | ----------------------------- | ----------------------- | ---------------------- | | `generic-chat-completion-api` | OpenAI Chat Completions | `/v1/chat/completions` | | `openai` | OpenAI Responses | `/v1/responses` | | `anthropic` | Anthropic Messages | `/v1/messages` | Use `generic-chat-completion-api` for Tokens. The `openai` type looks like the obvious choice but it means the Responses API. Tokens serves `/v1/responses`, but whether it works depends on the upstream provider behind each model. ### Anthropic Messages variant For the Messages API, set `"provider": "anthropic"` and `"baseUrl": "https://tokens.bd"`, without `/v1`, matching Factory's own `https://api.anthropic.com` example. Droid sends the key in `x-api-key` by default, which Tokens accepts. `"authMode": "bearer"` switches it to `Authorization: Bearer` if you prefer. The format is described in [Messages](/docs/messages). ### Legacy config.json Older setups use `~/.factory/config.json` with snake_case `custom_models`. Droid still loads it, but that file does not expand `${VAR}`, so the key would have to be written in plain text. Move to `settings.json`. ## Switch models Add one object to `customModels` per Tokens model. In Droid, run `/model`; custom models appear in their own "Custom models" section. Droid watches the settings file, so new entries show up without a restart. Get exact IDs from the [catalog](/models) or: ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` [Choosing a model](/docs/choosing-a-model) covers which models suit agent work. ## Verify it works Check the key and model directly: ```bash curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` Then run `droid`, choose "DeepSeek V4.1 Flash (Tokens)" from `/model`, and ask it to read a file in the current project. The request appears in your dashboard usage analytics. ## Troubleshooting **The model doesn't appear under Custom models.** The JSON is invalid (a trailing comma is the usual cause), or `customModels` was put in `config.json` with the camelCase name. **401 `missing_api_key`.** `${TOKENS_API_KEY}` expanded to nothing. Export the variable in the shell you start `droid` from. **401 `invalid_api_key`.** Wrong, revoked or rotated key. Rotation stops the old secret immediately. **404 `model_not_found`.** `model` must be the full Tokens ID with its provider prefix. **404 on every request.** Check `baseUrl`: `https://tokens.bd/v1` for `generic-chat-completion-api`, `https://tokens.bd` for `anthropic`. **Odd behavior with tools or caching.** Factory says only official Anthropic and OpenAI models are fully tested with Droid, and prompt caching on generic providers isn't guaranteed. Try a different model before assuming the gateway is at fault. **429 errors.** See `Retry-After`, and [Troubleshooting](/docs/troubleshooting) for each rate limit code. The request format is in [Chat Completions](/docs/chat-completions). Other terminal agents with similar setups: [Claude Code](/docs/claude-code), [OpenCode](/docs/opencode) and [Crush](/docs/crush). Sources: [Factory docs, BYOK overview](https://docs.factory.com/cli/byok/overview), [Droid quickstart](https://docs.factory.com/droid-cli/quickstart), checked October 2026. --- # Connect Warp to Tokens > Point Warp's agents at Tokens with a custom inference endpoint that speaks OpenAI Chat Completions, and know where Warp will not use it. Section: Coding Agents. Page: https://tokens.bd/docs/warp Warp is a terminal with built-in AI agents. It supports a **custom inference endpoint**: any server that implements OpenAI's `POST /v1/chat/completions`. Tokens is one, so you can add `https://tokens.bd/v1` as an endpoint and pick Tokens models in Warp's model picker. ## Limits to know first These come from Warp's documentation (checked October 2026): - **Requests go through Warp's backend.** Warp calls your endpoint from its servers, so the endpoint must be a public URL. `tokens.bd` is public, so this works, but it means your prompts pass through Warp as well as Tokens. - **"Auto" models and custom routers never use your endpoint.** Only a model you pick explicitly goes to Tokens. - **Cloud Agents don't use it.** Warp's docs say the custom endpoint doesn't apply to Cloud Agents. - **Eligibility.** The feature is free for individuals and organizations of 10 or fewer employees. Larger organizations need a Business or Enterprise plan, and those still consume Warp platform credits. If those limits rule Warp out, a terminal agent that runs entirely on your key, such as [Claude Code](/docs/claude-code), [OpenCode](/docs/opencode) or [Aider](/docs/aider), may suit you better. ## Add Tokens as a custom inference endpoint in Warp ### 1. Create a key Create a key in the dashboard. [API keys](/docs/api-keys) covers spend caps and model allow-lists. A separate key for Warp makes its usage easy to spot. ### 2. Enter the endpoint 1. Open Warp **Settings** and search for `inference endpoint`. 2. Add the endpoint URL: `https://tokens.bd/v1`. Warp's docs describe this field as "the base URL that exposes `/v1/chat/completions`", and their example uses a URL ending in `/v1`. 3. Add your API key: `tok_live_your_key`. 4. Add the model ID: `deepseek/deepseek-v4.1-flash`. 5. Save. In summary: ```text title="Warp Settings > inference endpoint" Endpoint URL: https://tokens.bd/v1 API key: tok_live_your_key Model ID: deepseek/deepseek-v4.1-flash ``` Warp stores the key in its settings and doesn't read `TOKENS_API_KEY` from your environment. ### 3. Select the model Open Warp's model picker and select the Tokens model explicitly. If the picker is on Auto, requests go to Warp's own models, not Tokens. ## Switch models Add another model ID in the endpoint settings for each Tokens model you want, then choose between them in the model picker. Use exact IDs, including the provider prefix. Find them in the [model catalog](/models) or with: ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` Warp's agents edit files and run commands, so pick a model with tool calling support. [Choosing a model](/docs/choosing-a-model) explains the trade-offs. ## Verify it works Test the key directly: ```bash export TOKENS_API_KEY=tok_live_your_key curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` Then, with the Tokens model selected, ask Warp's agent a short question. Open your dashboard usage analytics: the request should appear with the model you chose. If Warp answered but your usage shows nothing, Warp used one of its own models; check the picker. ## Troubleshooting **Nothing reaches Tokens.** The picker is on Auto, or you're using a Cloud Agent. Select the Tokens model explicitly, outside Cloud Agents. **The setting is unavailable.** Your organization may be above the 10-employee threshold without a Business or Enterprise plan. **401 `invalid_api_key`.** Wrong, revoked or rotated key. Rotation stops the old secret immediately, so update the key in Warp. **404 `model_not_found`.** The model ID must match the Tokens ID exactly, for example `deepseek/deepseek-v4.1-flash`. **404 on every request.** The endpoint URL should be `https://tokens.bd/v1`. Don't add `/chat/completions`; Warp appends it. **403 `model_not_allowed_on_key`.** The key's allow-list excludes the model. Allow-lists can't be changed after creation, so create a new key. **429 `rate_limited` or `concurrency_limit`.** Wait for `Retry-After`. Details in [Troubleshooting](/docs/troubleshooting). Warp's requests use the [Chat Completions](/docs/chat-completions) format. Source: [Warp docs, custom inference endpoint](https://docs.warp.dev/agents/inference/custom-inference-endpoint), checked October 2026. --- # Connect Amp to Tokens > Route models in Amp's own catalog through Tokens with a Model Routing Custom URL connection. Amp cannot add arbitrary model IDs, so read the limits first. Section: Coding Agents. Page: https://tokens.bd/docs/amp Amp (ampcode.com, from Sourcegraph) is a coding agent with a settings UI and a CLI. Its **Model Routing** feature lets you send requests for a model through your own connection, called a Custom URL. A Custom URL connection can speak OpenAI Chat Completions to `https://tokens.bd/v1`, or Anthropic Messages to `https://tokens.bd`. ## The limitation: Amp routes only its own catalog Read this before you start, because it decides whether Amp works for you. Amp's model mappings use **Amp's canonical model IDs**, written as `provider/model`. A Custom URL connection can only serve models that already exist in Amp's catalog. You can change the ID that gets sent upstream, but you cannot add a new model that Amp doesn't know about. For Tokens that means: - If a model is in Amp's catalog and Tokens carries the same model, you can route it through Tokens. - You can't add `deepseek/deepseek-v4.1-flash`, or any other Tokens model, to Amp unless Amp's catalog has an entry for it. We have not confirmed whether Amp's catalog includes a DeepSeek model that could map to Tokens. If you want to choose any model from the Tokens [catalog](/models) freely, use an agent with open custom provider support, such as [OpenCode](/docs/opencode), [Claude Code](/docs/claude-code) or [Cline](/docs/cline). ## Add a Custom URL connection in Amp Model Routing ### 1. Create a key Create a key in the dashboard. [API keys](/docs/api-keys) covers spend caps and allow-lists. If you want Amp to use only certain models through Tokens, a key with an allow-list enforces that on our side. ### 2. Add the connection 1. Open your personal settings in Amp and go to **Model Routing**. 2. Click **Add** and choose **Custom URL**. From the terminal, `amp config model-providers` does the same. 3. Pick the format and base URL: | Format | Base URL | Amp sends requests to | | ---------------------------- | ---------------------- | --------------------------------------- | | `chat-completions` (default) | `https://tokens.bd/v1` | `https://tokens.bd/v1/chat/completions` | | `anthropic-messages` | `https://tokens.bd` | `https://tokens.bd/v1/messages` | 4. Enter your API key, `tok_live_your_key`. Amp sends it as `Authorization: Bearer`. Get the base URL right for the format. `chat-completions` appends `/chat/completions`, so it needs `/v1` in the base. `anthropic-messages` appends `/v1/messages`, so it must not. Amp stores the key with the connection; it doesn't read `TOKENS_API_KEY`. ### 3. Map Amp models to Tokens IDs For each Amp catalog model you route, the mapping can rename the ID sent upstream. Amp's docs show the format with an exact line like this: ```text moonshotai/kimi-k3 -> kimi-k3-turbo ``` The left side is Amp's canonical ID. The right side is the ID your connection receives. For Tokens, the right side must be the exact Tokens model ID. Tokens IDs also use `provider/model` form, but they don't always match Amp's spelling, so check the exact ID in the [catalog](/models) or with: ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` ## Switch models Switching models in Amp works the way it always does: you pick from Amp's catalog. Which of those go through Tokens is decided by your Model Routing mappings. To route another model, add a mapping from its Amp ID to the matching Tokens ID. [Choosing a model](/docs/choosing-a-model) helps when Tokens offers several candidates. ## Verify it works Test the key directly: ```bash export TOKENS_API_KEY=tok_live_your_key curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` Then use Amp's own checks. In the UI, click **Check Access** on the connection. From the CLI: ```bash amp config model-providers test amp config model-providers check-access --provider-model ``` Replace `` with your connection's ID and `` with the Amp catalog ID you mapped. Finally, run a short task in Amp with that model and confirm the request appears in your dashboard usage analytics. ## Troubleshooting **The model you want isn't selectable.** It isn't in Amp's catalog. Model Routing can't add it. Use another agent for that model. **404 `model_not_found`.** The right-hand side of the mapping doesn't match a Tokens ID. Copy it exactly from `GET /v1/models`. **404 on every request.** The base URL doesn't match the format. Use `https://tokens.bd/v1` for `chat-completions` and `https://tokens.bd` for `anthropic-messages`. **401 `invalid_api_key`.** Wrong, revoked or rotated key. Rotation stops the old secret immediately. **403 `model_not_allowed_on_key` or `tier_permission_denied`.** The key's allow-list or your plan doesn't include the mapped model. **429 errors.** Read `Retry-After`; [Troubleshooting](/docs/troubleshooting) lists each code. Request formats are in [Chat Completions](/docs/chat-completions) and [Messages](/docs/messages). Source: [Amp docs, model routing](https://ampcode.com/docs/customize/model-routing), checked October 2026. --- # Other tools and compatibility > Which coding agents and editors can use a custom OpenAI or Anthropic endpoint like Tokens, which cannot, and a generic recipe for connecting any OpenAI-compatible tool. Section: Coding Agents. Page: https://tokens.bd/docs/other-tools Tokens can serve any tool that lets you set a custom OpenAI-compatible or Anthropic-compatible endpoint. This page lists which coding agents and editors can and can't, and gives a generic recipe for tools without their own guide. ## Compatibility table Based on each tool's official documentation, checked October 2026. "Partial" means some features use Tokens and others don't. | Tool | Custom endpoint? | Protocol to use with Tokens | Guide | | ------------------------ | --------------------- | -------------------------------------------------------------------------- | -------------------------------------- | | Claude Code | Yes | Anthropic Messages, base `https://tokens.bd` | [Claude Code](/docs/claude-code) | | Codex CLI | Yes, with a caveat | Responses API only; works if the model's upstream supports `/v1/responses` | [Codex CLI](/docs/codex-cli) | | OpenCode | Yes | Chat Completions | [OpenCode](/docs/opencode) | | OpenClaw | Yes | Chat Completions or Anthropic Messages | [OpenClaw](/docs/openclaw) | | Hermes Agent | Yes | Chat Completions | [Hermes Agent](/docs/hermes-agent) | | Crush | Yes | Chat Completions (`openai-compat`) | [Crush](/docs/crush) | | Aider | Yes | Chat Completions, `openai/` model prefix | [Aider](/docs/aider) | | Goose | Yes | Chat Completions, full endpoint URL | [Goose](/docs/goose) | | Qwen Code | Yes | Chat Completions | [Qwen Code](/docs/qwen-code) | | Kimi Code CLI | Yes | Chat Completions or Anthropic Messages | [Kimi Code](/docs/kimi-code) | | Factory Droid | Yes | Chat Completions or Anthropic Messages | [Factory Droid](/docs/factory-droid) | | Cline | Yes | Chat Completions | [Cline](/docs/cline) | | Kilo Code | Yes | Chat Completions | [Kilo Code](/docs/kilo-code) | | Roo Code | Yes, but archived | Chat Completions; repo archived 2026-05-15 | [Roo Code](/docs/roo-code) | | Continue | Yes | Chat Completions | [Continue](/docs/continue) | | Zed | Yes | Chat Completions or Anthropic Messages | [Zed](/docs/zed) | | GitHub Copilot (VS Code) | Partial | Chat only; inline suggestions and embeddings stay on GitHub | [GitHub Copilot](/docs/github-copilot) | | Cursor | Partial, undocumented | Base URL override not in Cursor's official docs; Tab never uses it | [Cursor](/docs/cursor) | | Warp | Partial | Chat Completions; not used by Auto models or Cloud Agents | [Warp](/docs/warp) | | Amp | Partial | Only models in Amp's own catalog can be routed | [Amp](/docs/amp) | | Gemini CLI | No | Only accepts Gemini or Vertex wire format | Use Qwen Code | | Kimi CLI (Python) | No | Archived; existing installs will stop working | Use Kimi Code CLI | ### Why Gemini CLI can't use Tokens Gemini CLI's only base URL overrides, `GOOGLE_GEMINI_BASE_URL` and `GOOGLE_VERTEX_BASE_URL`, expect the Gemini or Vertex request format. Tokens doesn't serve that format, so no setting makes it work. [Qwen Code](/docs/qwen-code), originally a Gemini CLI fork, supports OpenAI-compatible endpoints. ## Connect any OpenAI-compatible tool If your tool isn't listed, look for a setting called "OpenAI compatible", "custom provider", "API base" or "base URL". Then you need four things. ### 1. Base URL | Your tool speaks | Base URL | The tool then calls | | ----------------------- | ---------------------- | ---------------------- | | OpenAI Chat Completions | `https://tokens.bd/v1` | `/v1/chat/completions` | | Anthropic Messages | `https://tokens.bd` | `/v1/messages` | Some tools, such as Goose's custom provider file, want the full endpoint `https://tokens.bd/v1/chat/completions`. If the tool's own example ends in `/chat/completions`, copy that shape. A 404 on the first request almost always means `/v1` is missing or doubled. ### 2. API key Tokens reads the key from `Authorization: Bearer ` or `x-api-key: `, so OpenAI-style and Anthropic-style tools both work without extra headers. Where the tool reads environment variables, use `export TOKENS_API_KEY=tok_live_your_key` and keep the key out of config files you might commit. One key per tool makes usage easy to read and revocation painless; see [API keys](/docs/api-keys). ### 3. Model ID Use the exact Tokens ID, including the provider prefix: `deepseek/deepseek-v4.1-flash`, not `deepseek-v4.1-flash`. Get IDs from the [model catalog](/models) or from the API, which returns only models your key can use: ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` Tools that reference models as `/` end up with two slashes, such as `tokens/deepseek/deepseek-v4.1-flash`. Not every tool documents how it splits that, so check there first if the model isn't found. LiteLLM-based tools such as Aider want an `openai/` prefix, which they strip before sending. ### 4. Context window and output limit Many tools can't read a model's limits from the endpoint, so they ask you or assume a default. Use the values from the [catalog](/models). A wrong context window breaks automatic compaction: too high and long conversations get rejected upstream, too low and history is trimmed early. Also check: - **Tool calling.** Agents that edit files need a model with function calling. - **Streaming.** For a usage chunk at the end of an OpenAI-style stream, the client must send `stream_options: {"include_usage": true}`. - **Unsupported endpoints.** Image generation, audio, files, batches, assistants, fine-tuning and moderations return 404 `unsupported_endpoint`. Embeddings work only for embedding models, so keep a tool's embedding features on their current provider. - **Browser-only tools.** Tokens' responses carry no CORS headers, so calls straight from a web page fail. Use a server, CLI or editor extension. ## Test the connection Before you debug a tool's settings, confirm the key and model work. The tester below sends a short request with your key and the model you pick. Or run the curl that follows from your machine. ::connection-tester ```bash curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` If that succeeds and the tool still fails, the problem is in the tool's settings. [Connect Your Agent](/dashboard/connect) also runs a live test and has snippets for common agents. ## Troubleshooting | Status and code | Usual cause in a new tool | | ------------------------ | ------------------------------------------------------------- | | 404 on the first request | Base URL missing `/v1`, or path doubled | | 401 `missing_api_key` | Env var not set where the tool runs | | 404 `model_not_found` | Model ID missing the provider prefix | | 429 `rate_limited` | Agent burst past the per-minute limit; wait for `Retry-After` | Every code is in [Troubleshooting](/docs/troubleshooting); formats are in [Chat Completions](/docs/chat-completions) and [Messages](/docs/messages). The [Tokens CLI](/docs/tokens-cli) can configure OpenCode, Claude Code, Codex CLI and Crush for you. Sources: [Gemini CLI configuration reference](https://github.com/google-gemini/gemini-cli/blob/main/docs/reference/configuration.md), [MoonshotAI/kimi-cli archive notice](https://github.com/MoonshotAI/kimi-cli), plus the official docs linked from each tool's guide, checked October 2026. --- # Connect JetBrains AI Assistant to Tokens > Add Tokens as an OpenAI-compatible provider in JetBrains AI Assistant for chat, and optionally for inline completion. Covers the URL field, tool calling and the known path bug. Section: Coding Agents. Page: https://tokens.bd/docs/jetbrains-ai-assistant JetBrains AI Assistant is the AI plugin for IntelliJ IDEA, PyCharm, WebStorm and the other JetBrains IDEs. Besides JetBrains' own subscription, it can use your own API key with a third-party provider. Tokens fits the **OpenAI-compatible** provider type, which sends OpenAI Chat Completions requests to `https://tokens.bd/v1`. The models then appear in the AI Chat model selector. This guide was checked against JetBrains' AI Assistant 2026.2 documentation (the providers page is dated 15 July 2026 and the overview 5 August 2026; read October 2026). It was checked against the documentation, not run end to end with a live key. :::note[This is not the same as Claude Code or Junie in JetBrains] AI Assistant also hosts coding agents (Junie, Claude Agent, Codex, GitHub Copilot and agents added through the Agent Client Protocol). Those agents are separate from the OpenAI-compatible provider described here, and JetBrains' provider pages do not say that they use it. To run Claude Code through Tokens, follow [Claude Code](/docs/claude-code) and its own settings instead. For an open-source assistant that you configure in a file, see [Continue](/docs/continue), which also runs in JetBrains IDEs. ::: ## What you need - A JetBrains IDE with AI Assistant installed. JetBrains lists CLion, DataGrip, DataSpell, GoLand, IntelliJ IDEA, PhpStorm, PyCharm, Rider, RubyMine, RustRover and WebStorm. Use version 2026.1.2 or later (see the path bug below). - A Tokens key. [API keys](/docs/api-keys) covers spend caps and allow-lists. - A model id from the [model catalog](/models). This guide uses `deepseek/deepseek-v4.1-flash`. If your organization uses JetBrains IDE Services or JetBrains Central, JetBrains says an administrator can restrict providers or block your own API keys. If the settings below are missing, ask your admin. ## Add Tokens as a provider 1. Open **Settings** and go to **Tools | AI Assistant | Providers & API keys**. 2. In **Third-party AI providers**, set **Provider** to **OpenAI-compatible**. 3. In **URL**, enter `https://tokens.bd/v1`. 4. In **API Key**, paste your Tokens key. 5. Set **Tool calling** on or off (see below). 6. Click **Test Connection**, then **Apply**. 7. Open AI Chat and click the model selector. Your Tokens models are listed under the provider's section. ### Which URL to enter JetBrains' pages say only "Specify the URL of the provider's API endpoint". They give no example that includes or leaves out `/v1`, and they do not say whether the field takes a path. Enter `https://tokens.bd/v1`, which ends in `/v1` and has no `/chat/completions`. The reasons: - Two bug reports in JetBrains' tracker show the IDE building the model-list request itself. In LLM-22721 a custom path was ignored and the IDE asked for `/api/v1/models` (closed as a duplicate). In LLM-22911 a base URL ending in `/v4/` was rewritten to `/v1`. JetBrains lists LLM-22911 as fixed in 2026.1.2. - Tokens serves the model list at `https://tokens.bd/v1/models`, so a URL that ends in `/v1` gives the right path whether the IDE keeps your path or replaces it with `/v1`. If Test Connection fails and the Tokens request log shows a 404 `unsupported_endpoint` with a doubled path such as `/v1/v1/models`, the IDE is adding `/v1` itself. Try `https://tokens.bd` (the same host without `/v1`) instead. This fallback is an inference from the bug reports, not something JetBrains documents. JetBrains' pages do not say how the model list is built, but the bug reports above show the IDE requesting `/models`. Tokens serves it filtered by key, plan and wallet, so a key with an allow-list shows only the allowed models. ### Tool calling JetBrains describes the **Tool calling** setting as whether the model supports calling tools configured through MCP, and it appears only for OpenAI-compatible providers. Turn it on for a model that supports tool calling (check [Choosing a model](/docs/choosing-a-model)) if you use MCP servers in AI Chat. Leave it off for a model that does not. JetBrains' page does not document tool use outside MCP. ## Use Tokens for inline completion AI Completion (inline completion and next edit suggestions) has its own provider setting. 1. In the same settings page, find the **AI Completion** section and set **Provider** to **OpenAI Compatible**. 2. Enter your Tokens key as **API key** and `https://tokens.bd/v1` as **Base URL**. 3. Enter the **Model**, **Model context**, **Max output tokens** and **Prompt schema**, then click **Test Connection** and **Apply**. JetBrains says the model must be served by the endpoint in Base URL, and that inline completion needs a model with Fill-in-the-Middle (FIM) support. Next edit suggestions need edit-prediction support. Most chat models in the catalog are not FIM models, so check the [catalog](/models) before pointing completion at Tokens. Completion fires on short pauses while you type, which adds up on a metered key and against the per-minute rate limit. Use a separate key with a spend cap. ## Models Assignment For local models and OpenAI-compatible endpoints, JetBrains asks you to assign models to feature groups yourself under **Models Assignment**: | Group | Used for, per JetBrains | | ----------------- | ----------------------------------------------------------------- | | Core features | In-editor code generation and commit message generation | | Instant helpers | Chat context collection, chat title generation and name suggestions | The page also shows a **Context window** field, which defaults to 64 000 tokens for local models. Set it to the context window listed for your model in the [catalog](/models) so long chats are not cut short. ## Check that it works Test the key and model outside the IDE first: ```bash export TOKENS_API_KEY=tok_live_your_key curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` Then click **Test Connection** in the IDE, pick the Tokens model in AI Chat and send a message. The request appears in your dashboard usage analytics. ## Limits and what does not work - JetBrains' provider pages do not mention agents. Junie, Claude Agent and the other agents are configured separately. - JetBrains states that AI Assistant does not support invoking tools from configured MCP servers when using local models. The page makes no such statement for OpenAI-compatible endpoints, but it also does not confirm that every tool-calling feature works there. - If a feature has no model assigned that supports it, JetBrains can fall back to a JetBrains AI subscription, if you have one. - The model-selector list is only as complete as `GET /v1/models` for your key. ## Troubleshooting **Test Connection fails.** Check the URL first (see above), then the key. Run the `curl` commands above to separate a Tokens problem from an IDE problem. If you are on a version older than 2026.1.2, update, because older builds could replace the path. **401 `invalid_api_key` or `missing_api_key`.** The key is wrong, was left empty, or was rotated. Rotation stops the old secret immediately. Paste the new key and apply again. **The model list is empty.** An empty `data` array usually means no active plan and no wallet balance, or an allow-list that excludes every model. Subscribe or top up in [billing](/dashboard/billing). **404 `model_not_found`.** The model id must be the full Tokens id, including the provider prefix and the slash, such as `deepseek/deepseek-v4.1-flash`. **Errors after turning Tool calling on.** A 400 `invalid_request` can mean the upstream rejected the tool definitions. Turn the setting off or choose a model that supports tools. **402 `insufficient_credits` or 429 `rate_limited`.** Top up in [billing](/dashboard/billing), or wait for `Retry-After`. Inline completion is the usual cause of rate limits. Every code is listed in [Errors](/docs/errors), and the request format is in [Chat Completions](/docs/chat-completions). Sources: JetBrains AI Assistant help, [Providers and API keys](https://www.jetbrains.com/help/ai-assistant/settings-reference-providers-and-api-keys.html), [Use custom models](https://www.jetbrains.com/help/ai-assistant/use-custom-models.html), [Agents](https://www.jetbrains.com/help/ai-assistant/agents.html), tracker issues [LLM-22911](https://youtrack.jetbrains.com/projects/LLM/issues/LLM-22911) and [LLM-22721](https://youtrack.jetbrains.com/projects/LLM/issues/LLM-22721), checked October 2026. --- # Connect Xcode to Tokens > Add Tokens as an internet-hosted chat provider in Xcode's Intelligence settings. Enter the URL without /v1, add your key and pick models in the chat model picker. Section: Coding Agents. Page: https://tokens.bd/docs/xcode Xcode 26 and later can chat with models from a provider you add yourself. Apple requires the provider to support the Chat Completions API, which Tokens does. You add Tokens under **Xcode > Settings > Intelligence** as an **Internet Hosted** chat provider, and Xcode then lists the models your key can call. This guide was checked against Apple's page "Setting up coding intelligence" (Xcode documentation, 2026) and Apple Developer Forums threads on custom providers, October 2026. Apple's page does not name an Xcode version. It was checked against the documentation, not run end to end with a live key. ## What you need - Xcode 26 or later on a Mac. Use 26.3 or later, because a forum thread reports that Xcode 26.2 forgot added providers when it quit, and that 26.3 fixed it. - A Tokens key. [API keys](/docs/api-keys) covers spend caps and allow-lists. - Apple Intelligence turned on in System Settings. Apple's page does not say this, but guides from other providers list it as a requirement for any model provider in Xcode. ## Add Tokens 1. Choose **Xcode > Settings** and select **Intelligence** in the sidebar. 2. Under **Chat**, click **Add a Chat Provider** (some guides call it **Add a Model Provider**). 3. Select **Internet Hosted**. 4. Fill in the dialog, then click **Add**: | Field | What to enter | | ------------------ | -------------------------------------------------------------------------------- | | **URL** | `https://tokens.bd`, the gateway root with no `/v1` | | **API Key** | Your Tokens key, `tok_live_your_key` | | **API Key Header** | Leave empty (see below) | | **Description** | A label for yourself, such as `Tokens` | ### Why the URL has no /v1 Apple's page says Xcode expects the provider to support `{Model provider URL}/v1/models` and `{Model provider URL}/v1/chat/completions`. So the URL you enter is the root, and Xcode adds `/v1` and the endpoint itself. For Tokens that root is `https://tokens.bd`, which makes Xcode call `https://tokens.bd/v1/models` and `https://tokens.bd/v1/chat/completions`. Do not enter `https://tokens.bd/v1`. It already ends in `/v1`, so Xcode would request a doubled `/v1/v1/models` path that Tokens answers with 404 `unsupported_endpoint`. The same applies to a URL ending in `/chat/completions`. ### The API key header Apple's page does not describe the **API Key Header** field. Guides from Vercel and OpenRouter, both written for Xcode 26, say that leaving it empty makes Xcode send the standard `Authorization` header. Tokens reads `Authorization: Bearer `. Tokens also accepts the key in an `x-api-key` header, so if you must name a header, `x-api-key` with the bare key works. Do not type the word `Bearer` into the key field unless you set the header to `Authorization`, as the OpenRouter guide does. ### Where the key is kept Apple does not document how Xcode stores the key, so this guide does not claim it is in the macOS Keychain. Enter the key only in this dialog. Do not put it in a project file, a scheme, an `.xcconfig` or a script that gets committed. Create the key for Xcode alone, with its own monthly spend cap, so you can revoke it without touching anything else. ## Pick a model After you add the provider, click its row in the Intelligence settings. Xcode fills the **Models** table from the provider's model list. Mark the models you plan to use as favorites so they sit at the top of the picker. Open the chat from a project window and choose the model in the picker on the message field. Tokens model ids always have a slash, for example `deepseek/deepseek-v4.1-flash`. Xcode takes ids from the list instead of asking you to type them, and the Vercel guide shows slash ids working, but Apple does not document the id format. The list holds only the models the key can call: a key with an allow-list, or an account with no active plan and no wallet balance, shows a short or empty list. [Choosing a model](/docs/choosing-a-model) helps with the choice. Apple's page has no setting for context window, output limit or tool calling. Xcode sends the model id and the conversation, and the model's limits are whatever Tokens and the upstream enforce. Prefer a model with a large context window for chats that include many files, and see the [catalog](/models) for limits. ## What uses the custom provider | Part of Xcode | Uses Tokens? | | ---------------------------------------------------------- | -------------------------------------------------------------------------------- | | Chat with a provider you added under **Chat** | Yes | | Agents in the **Agents** section (Claude Agent, Codex) | No. Apple says you enable and sign in to each agent separately | | ChatGPT in Xcode, Claude Sonnet and Opus (Apple's built-in options) | No. They use their own accounts | | Agents you add through the Agent Client Protocol | Not documented by Apple for custom providers. Each agent has its own settings | Apple notes that the agent or model you set up in Intelligence settings may access your project files and other information when it handles a request. Requests sent through the Tokens provider include that project context, so use a key and a model you trust with your code. For agentic coding that runs through your Tokens key, use a tool with its own endpoint setting, such as [Claude Code](/docs/claude-code) in a terminal beside Xcode, and let Xcode do the chat. Managed Macs: an MDM profile can turn off the coding assistant by setting `CodingAssistantAllowExternalIntegrations` to `false`. If the Intelligence settings are missing or locked, ask your Mac administrator. ## Check that it works Test the key and the two paths Xcode uses before you add the provider: ```bash export TOKENS_API_KEY=tok_live_your_key curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` The first call must return a list of models. Then add the provider, wait for the model table to fill, pick a model and send a prompt in the chat. The request appears in your dashboard usage analytics. ## Troubleshooting **"Provider is not valid - Models could not be fetched with the provided account details."** Xcode could not get a model list from `https://tokens.bd/v1/models`. Check, in this order: the URL has no `/v1`, no trailing path and no typo; the key is complete; the `curl` call above returns a list. Apple's engineer says Xcode supports only `/v1` path prefixes, which Tokens has. **401 `invalid_api_key` or `missing_api_key`.** The key is wrong, empty, or was rotated, which stops the old secret immediately. If you set a custom header, make sure it matches how the key is written (see above). Edit the provider or add it again. **404 `unsupported_endpoint`.** The path is wrong, usually a doubled `/v1`. Remove `/v1` from the URL. **404 `model_not_found`.** The model id is unknown or inactive. Pick a model from the list again; the catalog can change. **The model table is empty.** The key can call no models. An empty list usually means no active plan and no wallet balance, or an allow-list that excludes everything. Subscribe or top up in [billing](/dashboard/billing). **402 `insufficient_credits`, 403 `monthly_spend_cap_exceeded` or 429 `rate_limited`.** Top up, raise the key's cap, or wait for `Retry-After`. **The provider is gone after restarting Xcode.** Apple's forum reports this for Xcode 26.2. Update to 26.3 or later. All codes are in [Errors](/docs/errors), and the request format is in [Chat Completions](/docs/chat-completions). Sources: Apple, [Setting up coding intelligence](https://developer.apple.com/documentation/xcode/setting-up-coding-intelligence); Apple Developer Forums threads [816031](https://developer.apple.com/forums/thread/816031) and [810665](https://developer.apple.com/forums/thread/810665); Vercel's [Xcode guide](https://vercel.com/docs/ai-gateway/ecosystem/framework-integrations/xcode) and OpenRouter's [Xcode guide](https://openrouter.ai/docs/community/xcode) for the field details Apple does not give. Checked October 2026. --- # GitHub Copilot CLI > Run GitHub Copilot CLI on Tokens models with its bring-your-own-key environment variables: base URL, key, model, token limits, and what to check when it fails. Section: Coding Agents. Page: https://tokens.bd/docs/github-copilot-cli GitHub Copilot CLI is GitHub's terminal coding agent, started with the `copilot` command. Its bring-your-own-key (BYOK) mode replaces GitHub-hosted models with a provider you choose. With Tokens it uses the default `openai` provider type, which sends OpenAI Chat Completions requests to `https://tokens.bd/v1`. This page is about the terminal agent. For Copilot Chat inside VS Code, see [GitHub Copilot in VS Code](/docs/github-copilot). The two are configured differently: | | Copilot CLI (this page) | Copilot in VS Code | | --- | --- | --- | | Where you configure it | Environment variables read by `copilot` | `chatLanguageModels.json` in VS Code | | Model choice | `COPILOT_MODEL` or `--model` | The model picker in Copilot Chat | | Scope | The whole CLI session | Chat and utility tasks only; inline suggestions stay on GitHub | :::note[Checked against the documentation] Based on GitHub's [BYOK reference for Copilot CLI](https://docs.github.com/en/copilot/how-tos/copilot-cli/customize-copilot/use-byok-models), checked October 2026. That page does not name a Copilot CLI version; GitHub announced BYOK in a changelog entry dated 7 April 2026. We have not run Copilot CLI against Tokens end to end. Run `copilot help providers` in your installed version to compare it with this page. ::: ## What you need - A Tokens key. Create one for Copilot CLI alone, with a monthly spend cap ([API keys](/docs/api-keys)). - The model id from [/models](/models). The examples use `deepseek/deepseek-v4.1-flash`. - Copilot CLI installed. GitHub's documented options are `npm install -g @github/copilot` (Node.js 22 or newer), `winget install GitHub.Copilot` on Windows, `brew install --cask copilot-cli`, or `curl -fsSL https://gh.io/copilot-install | bash` on macOS and Linux. On Windows, GitHub requires PowerShell 6 or newer. - A model that supports tool calling and streaming. Copilot CLI returns an error if either is missing. GitHub recommends a context window of 128K tokens or more. ## Set it up Copilot CLI turns BYOK on when `COPILOT_PROVIDER_BASE_URL` is set. Set the base URL, your key and the model, then start `copilot`: :::code-tabs ```bash title="macOS / Linux" export TOKENS_API_KEY="tok_live_your_key" export COPILOT_PROVIDER_BASE_URL="https://tokens.bd/v1" export COPILOT_PROVIDER_API_KEY="$TOKENS_API_KEY" export COPILOT_MODEL="deepseek/deepseek-v4.1-flash" copilot ``` ```powershell title="Windows PowerShell" $env:TOKENS_API_KEY = "tok_live_your_key" $env:COPILOT_PROVIDER_BASE_URL = "https://tokens.bd/v1" $env:COPILOT_PROVIDER_API_KEY = $env:TOKENS_API_KEY $env:COPILOT_MODEL = "deepseek/deepseek-v4.1-flash" copilot ``` ::: The PowerShell lines last for the current window. Keep the key in your shell profile or a secrets manager, not in a file inside a repository. | Variable | Value for Tokens | | --- | --- | | `COPILOT_PROVIDER_BASE_URL` | `https://tokens.bd/v1`. Required. It ends in `/v1`, like the OpenAI example on GitHub's page. | | `COPILOT_PROVIDER_TYPE` | Leave unset. The default is `openai`. Do not use `azure` or `anthropic` for Tokens. | | `COPILOT_PROVIDER_API_KEY` | Your Tokens key. | | `COPILOT_MODEL` | The Tokens model id, or pass `--model ` to `copilot`. Required. | The key is the only credential Tokens needs. If you would rather use `COPILOT_PROVIDER_BEARER_TOKEN`, Tokens also accepts `Authorization: Bearer`, but the API key variable is the simpler choice. ### Token limits Copilot CLI looks up context limits from a well-known model name. Tokens ids are not well-known names to it, so set the limits yourself: ```bash export COPILOT_PROVIDER_MAX_PROMPT_TOKENS=120000 export COPILOT_PROVIDER_MAX_OUTPUT_TOKENS=8192 ``` The numbers are placeholders. Copy the real context window and maximum output from the model's page in [/models](/models), and leave room for the output inside the context window. `COPILOT_PROVIDER_MODEL_ID` and `COPILOT_PROVIDER_WIRE_MODEL` exist for cases where the name the CLI should look up differs from the name sent to the provider. You do not need either with Tokens, because the id you set in `COPILOT_MODEL` is the id Tokens expects. ### The wire API setting GitHub's page also lists `COPILOT_PROVIDER_WIRE_API`, but it does not say which values are accepted. Leave it unset. Tokens supports both [Chat Completions](/docs/chat-completions) and the [Responses API](/docs/responses), so the default works. Other guides set it to `responses` for specific models; if you try that, test with the check below first. ## Check that it works Test the key and model outside Copilot CLI first: ```bash curl https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` Then run Copilot CLI with one prompt: ```bash copilot -p "Reply with OK" ``` You should get a short answer, and the request should appear in usage analytics on your dashboard. To check tool calling, start `copilot` in a project and ask it to read a file. If the model answers but never touches your files, it probably does not support tool calling well; see [Tool calling](/docs/tool-calling). ## Choosing a model Pick a model with reliable tool calling and a long context. [Choosing a model](/docs/choosing-a-model) compares them, and the `/models` catalog shows context and output limits for each. Switch for a session with `copilot --model ` or by changing `COPILOT_MODEL`. Only one provider configuration is active at a time. ## Limits and what to know - **Spend.** An agent loops: one prompt can become many requests, each carrying the growing conversation. Use a key with a monthly spend cap ([API keys](/docs/api-keys)) and watch usage on the dashboard. - **GitHub sign-in.** GitHub's announcement says signing in to GitHub is optional in BYOK mode, and that signing in adds GitHub features such as `/delegate`, GitHub code search and the GitHub MCP server. GitHub's install page lists an active Copilot subscription as a prerequisite. The BYOK page does not say which plans support BYOK. Ask whoever manages your Copilot plan if you are unsure. - **Organization policy.** If your organization or enterprise has disabled Copilot CLI, you cannot use it, with or without BYOK. - **What still goes to GitHub.** GitHub's page does not list it. It does say `COPILOT_OFFLINE=true` stops Copilot CLI contacting GitHub's servers. That setting does not make Tokens local: prompts and code context still go to `https://tokens.bd/v1`. - **Model ids with a slash.** Tokens ids look like `provider/model`. GitHub's page does not discuss ids with a slash. If `COPILOT_MODEL` is rejected, test the id with the `curl` above first. - **Sessions.** Start `copilot` in a shell where the variables are set. A session started elsewhere does not see them. ## Troubleshooting **Copilot CLI still uses GitHub models.** `COPILOT_PROVIDER_BASE_URL` is not set in the shell that started `copilot`. Run `echo $COPILOT_PROVIDER_BASE_URL` (or `$env:COPILOT_PROVIDER_BASE_URL` in PowerShell). **401 `missing_api_key` or `invalid_api_key`.** `COPILOT_PROVIDER_API_KEY` is empty, or it holds another provider's key. A Tokens key starts with `tok_live_`. See [Errors](/docs/errors). **404 `model_not_found`.** `COPILOT_MODEL` is not an exact Tokens id. Copy it from [/models](/models) or `GET https://tokens.bd/v1/models`. **404 `unsupported_endpoint`.** The base URL is missing `/v1`, or has an extra path. Use `https://tokens.bd/v1` exactly. **An error about tool calling or streaming.** The model does not support one of them. Pick a model that does ([Streaming](/docs/streaming), [Tool calling](/docs/tool-calling)). **403 `monthly_spend_cap_exceeded` or 402 `insufficient_credits`.** The key's cap or your balance is used up. Raise the cap or top up in [billing](/dashboard/billing). **429 `rate_limited` or `concurrency_limit`.** The agent sent requests faster than your limits allow. Wait for `Retry-After`; see [Rate limits](/docs/rate-limits). For the full list of codes, see [Errors](/docs/errors). Other problems: [Troubleshooting](/docs/troubleshooting). --- # OpenHands > Connect OpenHands to Tokens as an OpenAI-compatible endpoint, in the web UI or the CLI, with the right model string for ids that already contain a slash. Section: Coding Agents. Page: https://tokens.bd/docs/openhands OpenHands is an open-source AI software engineer that edits code, runs commands and browses inside a sandbox. It calls models through LiteLLM, so Tokens is added as an OpenAI-compatible endpoint: a model string starting with `openai/`, a Base URL of `https://tokens.bd/v1`, and your Tokens key. Requests are OpenAI Chat Completions. :::note[Checked against the documentation] Based on the OpenHands documentation on docs.openhands.dev (the OpenAI, local LLM, LLM overview and CLI pages), checked October 2026. The pages do not state a CLI version; the install page we read pins agent server image `1.26.0-python`. We have not run OpenHands against Tokens end to end. ::: ## What you need - A Tokens key. Create one only for OpenHands, with a monthly spend cap ([API keys](/docs/api-keys)). - The model id from [/models](/models). The examples use `deepseek/deepseek-v4.1-flash`. - OpenHands installed. The CLI and the local web UI both need Python 3.12 and [uv](https://docs.astral.sh/uv/), or you can use the binary or Docker. The web UI also needs Docker running. On Windows, OpenHands says to run everything inside WSL (Ubuntu). - A model that supports tool calling. OpenHands says it needs a powerful model, and that open-weight models vary in how reliably they call tools. ## Which model string to type LiteLLM chooses the provider from the text before the first slash. For a custom OpenAI-compatible endpoint that text must be `openai/`, and OpenHands says the prefix is required in Custom Model. Tokens ids already contain a slash (`provider/model`), so the value you type has two: ```text openai/deepseek/deepseek-v4.1-flash ``` The pattern is `openai/`. OpenHands documents the same shape for proxies that have their own routing prefix (`openai//`), and its local LLM page uses `openai/qwen/qwen3.6-35b-a3b` for an LM Studio model whose id contains a slash. Neither OpenHands nor LiteLLM documents which part is sent to the endpoint. These examples rely on LiteLLM using the first `openai/` only to pick the provider and sending the rest, here `deepseek/deepseek-v4.1-flash`, which is the id Tokens expects. That is our reading of the examples, not a documented guarantee. If requests come back as `model_not_found`, see the troubleshooting section below. ## Set it up in the web UI Start the UI: ```bash uv tool install openhands --python 3.12 openhands serve ``` Open `http://localhost:3000`. Then: 1. Select the Settings button (gear icon), then the **LLM** tab. 2. Turn on the **Advanced** toggle. 3. Set **Custom Model** to `openai/deepseek/deepseek-v4.1-flash`. 4. Set **Base URL** to `https://tokens.bd/v1`. 5. Set **API Key** to your Tokens key. 6. Save the settings. OpenHands keeps its state in `~/.openhands` on your machine. Treat that folder as secret and never commit it, because it holds your settings and may hold the key. ## Set it up in the CLI On first start the CLI asks for an LLM provider and API key and saves them under `~/.openhands/`. The pages we read do not list a Base URL or Custom Model prompt in that first-run dialog, so use environment variables for Tokens: ```bash export LLM_API_KEY="tok_live_your_key" export LLM_MODEL="openai/deepseek/deepseek-v4.1-flash" export LLM_BASE_URL="https://tokens.bd/v1" openhands --override-with-envs ``` :::warning[Environment variables are ignored without the flag] OpenHands ignores `LLM_*` variables unless you start the CLI with `--override-with-envs`. The override is not saved: a plain `openhands` the next day goes back to whatever is stored in `~/.openhands/`. Put the four lines in a small shell script or alias if you use Tokens every time. ::: To change the stored model later, press `Ctrl+P` in the CLI and choose Settings, or edit the `model` field in `~/.openhands/agent_settings.json`. The CLI docs name both `settings.json` and `agent_settings.json` for the stored LLM settings, so look in the folder to see which your version uses. We could not confirm from the documentation whether the CLI's Settings screen has a Base URL field, which is why the steps above use the environment variables. On Windows, run these commands in the WSL shell, not in PowerShell. ## Check that it works First rule out OpenHands, with the id as Tokens expects it (no `openai/` prefix): ```bash curl https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $LLM_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` Then, in OpenHands, ask for something small ("List the files in the workspace and tell me what the project does"). The agent should reply and run a command, and the request should show up in usage analytics on your dashboard. A reply with no actions at all usually means the model is not calling tools. ## Choosing a model OpenHands works in long loops of tool calls, so pick a model with strong tool calling and a long context. [Choosing a model](/docs/choosing-a-model) compares them, and [/models](/models) shows each model's context and output limits. OpenHands' own advice is to use the strongest model you can afford for long or high-stakes tasks. The pages we read do not mention a field for the context window or maximum output. If the model has a native tool calling switch under model customization, OpenHands says it can be toggled there. If you see malformed JSON errors or poor output, OpenHands suggests a stronger model or a larger context window before anything else. ## Limits and what to know - **Spend.** OpenHands retries failed calls (`LLM_NUM_RETRIES`, default 4) and loops until a task is done. A single task can send hundreds of requests, each carrying a growing conversation. Use a spend-capped key ([API keys](/docs/api-keys)) and watch the dashboard. OpenHands itself warns to set spending limits. - **LiteLLM costs.** Any cost OpenHands shows is an estimate from LiteLLM's own price table, which may not list Tokens ids. Your real spend is in the dashboard. - **Other LLM settings.** A few options (`LLM_API_VERSION`, `LLM_DROP_PARAMS`, `LLM_DISABLE_VISION`, `LLM_CACHING_PROMPT`) are environment variables or `config.toml` entries only, not UI fields. - **Version changes.** OpenHands changed its settings format at version 1.0.0. If you upgrade from an older install, redo the setup. ## Troubleshooting **404 `model_not_found`.** The model string reached Tokens with the wrong shape. The Custom Model value must be `openai/` followed by the exact Tokens id (`openai/deepseek/deepseek-v4.1-flash`). Without the `openai/` prefix, LiteLLM can treat the first part of the id as a provider name. If the prefix is there and it still fails, run the `curl` above with the id from [/models](/models) to confirm the id itself is right. **LiteLLM says the provider is not provided or not recognized.** The Custom Model value has no `openai/` prefix. Add it. **401 `missing_api_key` or `invalid_api_key`.** The API Key field or `LLM_API_KEY` is empty or holds a different provider's key. A Tokens key starts with `tok_live_`. **404 on every request.** The Base URL must be `https://tokens.bd/v1` and nothing more. LiteLLM adds the path itself, so do not append `/chat/completions`. **The CLI ignores my settings.** `LLM_*` variables need `--override-with-envs`. Without it the stored settings win. **403 `monthly_spend_cap_exceeded`, 402 `insufficient_credits`, or 429 `rate_limited`.** The key's cap, your balance or your rate limit stopped the agent. Raise the cap or top up in [billing](/dashboard/billing), or wait for `Retry-After`. The OpenHands retry loop makes 429 more likely during busy sessions. **From the web UI container, the endpoint is unreachable.** Tokens is a public address, so this is usually a Docker network or proxy problem on your side. Run the `curl` from the same machine first. For all codes, see [Errors](/docs/errors); for other problems, [Troubleshooting](/docs/troubleshooting). --- # Open WebUI > Add Tokens as an OpenAI connection in Open WebUI, list your models, move background tasks such as chat titles to a cheap model, and fix connection errors. Section: Chat Apps & Automation. Page: https://tokens.bd/docs/open-webui Open WebUI is a self-hosted chat interface. It talks to any OpenAI-compatible server through an OpenAI connection, which sends Chat Completions requests, so Tokens works with a base URL of `https://tokens.bd/v1` and a Tokens key. You add the connection in the admin settings, or with environment variables. This guide is based on Open WebUI's documentation, checked on 11 October 2026 (the latest release then was v0.12.0). It was checked against the documentation, not run end to end against a live Open WebUI install. ## What you need - A Tokens key. Create one in the dashboard; [API keys](/docs/api-keys) covers spend caps and allowed models. - Admin access to your Open WebUI instance. - At least one model id from the [model catalog](/models). This guide uses `deepseek/deepseek-v4.1-flash`. Ids have the form `provider/model`. ## Add Tokens as a connection 1. In Open WebUI, go to **Settings > Admin > Connections**. 2. In the **Manage OpenAI API Connections** list, click the add button (**Add Connection**). 3. Leave **Connection Type** as **External**. 4. Set **URL** to `https://tokens.bd/v1`. 5. Paste your Tokens key into **API Key**. 6. Click **Save**. Enter the base URL only. The examples in Open WebUI's documentation all end at `/v1` and none include a path such as `/chat/completions`, and Tokens' own base URL ends at `/v1` too. Saving does not test the connection. To test it, use **Verify Connection**, which calls `GET /models` with a Bearer token. Tokens answers that call with the models your key can use, so a successful check also proves the key works. Leave the **Provider** setting under **Advanced** at **Default**, and leave **Forward cookies** off. Open WebUI's docs say to enable that only for a server you control that authenticates by cookie, and never for third-party endpoints. ### Or set it with environment variables Open WebUI reads the connection from environment variables as well. For one connection: ```bash title="Environment of the Open WebUI container" OPENAI_API_BASE_URL=https://tokens.bd/v1 OPENAI_API_KEY=tok_live_your_key ``` `OPENAI_API_BASE_URLS` and `OPENAI_API_KEYS` take several values separated by semicolons. If you use them to add Tokens next to another provider, keep the two lists in the same order. In a Docker Compose file, pass the key from your shell or an `.env` file next to the compose file instead of typing it into the file you commit: ```yaml title="docker-compose.yml (excerpt)" services: open-webui: environment: - OPENAI_API_BASE_URL=https://tokens.bd/v1 - OPENAI_API_KEY=${TOKENS_API_KEY} ``` These variables are marked as `ConfigVar` in Open WebUI's reference: the value is stored on first launch, and after that Open WebUI uses the stored value rather than the environment, so later changes are made in **Settings > Admin > Connections**. The reference describes `ENABLE_PERSISTENT_CONFIG=False` as the switch that makes environment variables win again, with UI changes lost on restart. :::note[Docker and host.docker.internal] Open WebUI's documentation uses `host.docker.internal` for model servers that run on the Docker host. Tokens is a remote address, so the URL stays `https://tokens.bd/v1`. ::: ## Choose which models appear By default Open WebUI shows every model the connection returns from `GET /models`. For Tokens that is the list your key is allowed to use, so a key with an allowed-models list shows just those models. The **Model IDs** field under **Advanced** changes this: - Empty (the default): all models from the provider are detected. - Filled: the list you type replaces the fetched one. Open WebUI stops calling the provider's `/models` for that connection and shows only the ids you entered. Use it to hide models from your users. Type the exact Tokens id, for example `deepseek/deepseek-v4.1-flash`, and click the plus button to add it. Duplicates are refused and surrounding spaces are removed. ### Ids with a slash Tokens ids always contain a slash. Open WebUI's documentation does not describe special handling of slashes in model ids, so the ids are shown as the provider returns them. If you want a label per connection, the **Prefix ID** field joins your prefix and the model id with a dot (a prefix `tokens` gives `tokens.provider/model`) and removes it again before the request goes upstream. Type the prefix without a trailing slash: the docs show that `groq/` produces `groq/.model`. ## Check that it works 1. Open a new chat and pick a Tokens model from the model selector. 2. Send "Reply with OK". 3. The reply streams in, and the request appears in your Tokens [usage](/dashboard/usage). If you want to rule out Open WebUI first, run this from any machine: ```bash curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` ## Streaming and tool calling Open WebUI's docs list `POST /v1/chat/completions` as the one required endpoint, with streaming and the usual parameters such as temperature, top_p and max_tokens. Tokens streams that endpoint as Server-Sent Events; see [Streaming](/docs/streaming). Tool use needs a model and a provider that accept `tools` and `tool_choice`, which Open WebUI's docs call out as a requirement. Tokens passes these through, but not every model handles them well. Check [Choosing a model](/docs/choosing-a-model) and the model's page in the [catalog](/models) before enabling tools for your users, and read [Tool calling](/docs/tool-calling) for the request format. ## Background tasks cost credits: set a cheap task model Open WebUI sends extra requests behind each conversation. Its task-models page lists them: chat titles, tags, follow-up suggestions, autocomplete, retrieval and web search query rewriting, image prompts and context compaction summaries. Each one is a billed Tokens request. By default the task model is **Current Model**, so these requests run on whatever model the user chatted with, an expensive one included. To change that: 1. Go to **Settings > Admin > Interface** and find the **Tasks** section. 2. Set **External Task Model** to a small, fast, non-reasoning model from your Tokens list. Open WebUI's docs recommend a small, non-reasoning model because reasoning models add latency and cost for simple outputs. This is the picker that applies to Tokens, because it is not a Local connection such as Ollama. The environment variable is `TASK_MODEL_EXTERNAL`. 3. Under **Task Model Parameters**, set `max_tokens` if you want a hard cap on background requests. The docs note that setting any parameter removes the built-in 1000 token limit for titles and summaries, so include `max_tokens` yourself. The variable is `TASK_MODEL_PARAMS`, a JSON object. If the named task model is unavailable, Open WebUI falls back to the chat's own model instead of failing, which also means a task model that your key may not call silently costs you the expensive model again. Keep the task model in the key's allowed-models list. To stop tasks altogether, turn off their toggles under **Generation** in the same admin page: Title Generation (`ENABLE_TITLE_GENERATION`), Tags Generation (`ENABLE_TAGS_GENERATION`), Follow Up Generation (`ENABLE_FOLLOW_UP_GENERATION`) and Autocomplete Generation (`ENABLE_AUTOCOMPLETE_GENERATION`). Open WebUI's reference lists autocomplete as off by default and the others as on. Autocomplete fires while users type, so leave it off on a metered key. Context compaction has its own picker, **Context Compaction Model**, under **Chat**. ## Run it on a shared server Everyone who uses the instance spends the one key you put in the connection, and shares that account's rate and concurrency limits ([Rate limits](/docs/rate-limits)). Before you give the instance to other people: - Create a key just for this instance, with a monthly spend cap and an allowed-models list ([API keys](/docs/api-keys)). - Put the task model in that list, as described above. - Keep the key out of files you commit. Use the environment variable or the admin page. - Watch usage in the dashboard; Open WebUI does not show you what the Tokens account has left. `GET /v1/tokens/usage` does ([Models and usage](/docs/models-and-usage)). ## What does not go through this connection - **Images, speech and transcription.** Open WebUI has separate settings and variables for image generation, text-to-speech and speech-to-text. Tokens does not serve those endpoints; they return 404 `unsupported_endpoint`. Point those features at another provider. - **Embeddings for RAG.** Open WebUI's docs mark `/v1/embeddings` as optional, used for RAG. Tokens serves embeddings only for models that are embedding models ([Embeddings](/docs/embeddings)). Keep the default RAG embedding setup unless you pick one of those. ## Troubleshooting **Verify Connection fails, or the model list is empty.** Check the URL ends in `/v1` and has no trailing path. An empty list from Tokens usually means no active plan and no wallet balance; see [Models and usage](/docs/models-and-usage). Open WebUI's docs also say a failed check does not always mean chat is broken: add the model id under **Model IDs** and try a chat. **The settings page hangs or the list loads slowly.** Open WebUI times out the model list fetch after 10 seconds by default; raise `AIOHTTP_CLIENT_TIMEOUT_MODEL_LIST` on slow networks. Open WebUI's connection-error guide has a section on model list loading problems for a saved URL that cannot be reached. **401 `invalid_api_key` or `missing_api_key`.** The key is wrong, rotated or empty. Paste a fresh one into the connection. Rotation stops the old secret immediately. **404 `model_not_found`.** The model id is not the full Tokens id. Copy it from the model list, `provider/model` included. **403 `model_not_allowed_on_key`.** The key's allowed-models list does not include the model. This also happens when a task model is outside the list; chat titles fail while normal chat works. **402 `insufficient_credits`, 403 `monthly_spend_cap_exceeded`.** Out of credits, or the key hit its cap. Top up in [billing](/dashboard/billing) or use a key with a higher cap. **429 `rate_limited` or `concurrency_limit`.** Many users, or many background tasks, on one key. Wait for `Retry-After`, turn off the task toggles you don't need, or use a second key for a second instance. Every code is listed in [Errors](/docs/errors); more fixes are in [Troubleshooting](/docs/troubleshooting). If you ask [support](/docs/support) for help, include the `x-tokens-request-id` header from a curl run with `-i`. Sources: Open WebUI documentation, [OpenAI-compatible providers](https://docs.openwebui.com/getting-started/quick-start/connect-a-provider/starting-with-openai-compatible/), [Starting with OpenAI](https://docs.openwebui.com/getting-started/quick-start/connect-a-provider/starting-with-openai), [task models](https://docs.openwebui.com/features/administration/task-models/) and the [environment variable reference](https://docs.openwebui.com/reference/env-configuration), checked 11 October 2026. --- # LibreChat > Add Tokens to LibreChat as a custom endpoint in librechat.yaml: base URL, key from .env, model list, a cheap title model, Docker mounting and error fixes. Section: Chat Apps & Automation. Page: https://tokens.bd/docs/librechat LibreChat is a self-hosted chat application. It adds OpenAI-compatible providers as custom endpoints in a file called `librechat.yaml`. A custom endpoint with the base URL `https://tokens.bd/v1` sends Chat Completions requests to Tokens. LibreChat can also talk to the Anthropic Messages API through a custom endpoint, covered near the end. This guide is based on LibreChat's documentation, checked on 11 October 2026. The latest release then was v0.8.8, and the `librechat.example.yaml` in its repository uses config `version: 1.3.17`. It was checked against the documentation, not run end to end against a live LibreChat install. ## What you need - A Tokens key. Create one in the dashboard; [API keys](/docs/api-keys) covers spend caps and allowed models. - A LibreChat install where you can edit `librechat.yaml` and `.env`, and restart it. - A model id from the [model catalog](/models). This guide uses `deepseek/deepseek-v4.1-flash`. Ids have the form `provider/model`. ## Add the key to .env LibreChat's documentation recommends referencing keys from `.env` with `${VAR_NAME}` instead of writing them into `librechat.yaml`. Add the key to the `.env` file in the LibreChat project root: ```bash title=".env" TOKENS_API_KEY=tok_live_your_key ``` Keep `.env` out of git. Every `${...}` reference in the config needs a matching entry. If one is missing, the endpoint still appears, and the error `Missing API Key for ` shows up when a message is sent. ## Add the endpoint to librechat.yaml Create `librechat.yaml` next to `.env` (or set `CONFIG_PATH` in `.env` to its full path) and add Tokens under `endpoints.custom`: ```yaml title="librechat.yaml" version: 1.3.17 cache: true endpoints: custom: - name: 'Tokens' apiKey: '${TOKENS_API_KEY}' baseURL: 'https://tokens.bd/v1' models: default: ['deepseek/deepseek-v4.1-flash'] fetch: true titleConvo: true titleModel: 'deepseek/deepseek-v4.1-flash' modelDisplayLabel: 'Tokens' ``` If you already have a `librechat.yaml`, add only the `- name: 'Tokens'` block under the existing `endpoints.custom` list and keep your own `version`. What each field does, from LibreChat's reference: - `name`: required, unique (compared case-insensitively). It is the title in the endpoint selector. - `apiKey` and `baseURL`: required. `${TOKENS_API_KEY}` reads the value from `.env`. LibreChat also accepts `user_provided`, where each user types their own key in the interface. Writing the key into the file as plain text is possible but the docs advise against it. - `models`: required. `default` is a non-empty list and is the fallback if fetching fails. `fetch: true` makes LibreChat ask the endpoint for its model list. - `titleConvo` and `titleModel`: turn on conversation titles and choose the model that writes them. See below, because the default is not a Tokens model. - `modelDisplayLabel`: the name shown next to responses. LibreChat silently drops an entry if `name`, `baseURL`, `apiKey` or `models` is missing or misspelled, or if `models` has neither `fetch: true` nor a `default` list. A schema error anywhere in the file stops the server and disables all custom endpoints, so check `docker compose logs api` if LibreChat does not come up. ### Docker: mount the file A Docker container cannot see `librechat.yaml` unless you mount it. Copy `docker-compose.override.yml.example` to `docker-compose.override.yml`, then uncomment the bind mount under the `api` service that maps `./librechat.yaml` to `/app/librechat.yaml`. Compose merges the override file with `docker-compose.yml` automatically. After any change to `librechat.yaml` or `.env`, restart: ```bash docker compose down && docker compose up -d ``` For a local install without Docker, stop the process and run `npm run backend`. Then open LibreChat and check that **Tokens** is in the endpoint selector. ## Models and the slash in ids With `fetch: true`, LibreChat calls the endpoint's model list when it needs it. Tokens answers `GET https://tokens.bd/v1/models` with only the models your key can use, so a key with an allowed-models list shows just those. LibreChat's docs warn that fetching can slow first use if the endpoint responds late. `default` stays as the fallback. To pin a fixed list instead, set `fetch: false` and list the ids you want: ```yaml models: default: ['deepseek/deepseek-v4.1-flash', 'provider/another-model'] fetch: false ``` Replace `provider/another-model` with an id from the catalog. LibreChat's documentation does not describe special handling of slashes in model ids. Its own OpenRouter example, whose ids look like `meta-llama/llama-3-70b-instruct`, uses them as plain strings in `default` and `titleModel`, and the same shape applies here: write the full Tokens id, `provider/model`. ## Check that it works 1. Select **Tokens** in the endpoint selector and pick a model. 2. Send "Reply with OK". The response streams in, and the request shows in your Tokens [usage](/dashboard/usage). To test the key outside LibreChat: ```bash curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` ## Streaming and tool calling Chat responses stream as Server-Sent Events, which Tokens passes through ([Streaming](/docs/streaming)). LibreChat's custom-endpoint page does not cover agents or tool calling, so this guide does not describe how LibreChat uses tools with a custom endpoint. If you use a feature that sends tool definitions, the model needs to support tool calling; see [Tool calling](/docs/tool-calling) and [Choosing a model](/docs/choosing-a-model). If Tokens returns 400 `invalid_request` because the upstream rejected a request parameter, LibreChat has `dropParams` to remove a default parameter from requests (for example `dropParams: ['stop']`) and `addParams` to add or override one. Use them only after you have seen which parameter is rejected. ## Chat titles cost credits: set a cheap titleModel When `titleConvo` is `true`, LibreChat sends an extra request to name each conversation. Two settings matter, both from LibreChat's reference: - `titleConvo` defaults to `false`. Leave it that way to send no title requests. - `titleModel` defaults to `gpt-3.5-turbo`. That is not a Tokens model, so with `titleConvo: true` and no `titleModel`, the title request fails with 404 `model_not_found`. Set `titleModel` to a small, fast Tokens model, not the model the user is chatting with: ```yaml titleConvo: true titleModel: 'provider/cheap-model-id' ``` Replace it with a cheap model from the [catalog](/models). The value `current_model` is also accepted and uses the conversation's own model, which is the expensive choice when users chat with a large model. Keep that model in the key's allowed-models list. Otherwise titles fail with 403 `model_not_allowed_on_key` while normal chat works. ## Use the Anthropic Messages API instead For models you want to reach through `/v1/messages`, LibreChat supports a custom endpoint with `provider: 'anthropic'`. Its docs say to use that only for endpoints that implement the native Anthropic Messages API, and to omit `provider` for OpenAI-compatible gateways such as the entry above. Tokens serves both. The `baseURL` for this provider is the API root, not the `/v1/messages` path, so it is `https://tokens.bd`: ```yaml title="librechat.yaml (second entry under endpoints.custom)" - name: 'Tokens Messages' provider: 'anthropic' apiKey: '${TOKENS_API_KEY}' baseURL: 'https://tokens.bd' headers: anthropic-version: '2023-06-01' models: default: ['deepseek/deepseek-v4.1-flash'] fetch: false titleConvo: true titleModel: 'provider/cheap-model-id' ``` For this provider `models.fetch` is not used, so every model must be listed under `models.default`. The `anthropic-version` header comes from LibreChat's own Anthropic example. Tokens accepts the key in `x-api-key` or as a Bearer token ([Authentication](/docs/authentication)); the format is in [Messages](/docs/messages). Most people only need the first entry. ## Run it on a shared server All users of one LibreChat instance spend one Tokens account if you put a single key in `.env`, and share its rate and concurrency limits ([Rate limits](/docs/rate-limits)). Create a key just for the instance with a monthly spend cap and an allowed-models list ([API keys](/docs/api-keys)), including your `titleModel` in the list. The other option, `apiKey: 'user_provided'`, lets each user enter their own Tokens key in the interface, so each person's usage and cap stay separate. LibreChat stores those keys encrypted. ## What does not go through Tokens Tokens serves chat. Image, audio and file endpoints return 404 `unsupported_endpoint`, and embeddings work only for embedding models ([Models and usage](/docs/models-and-usage)). Keep LibreChat's features that need those on another provider. ## Troubleshooting **Tokens is missing from the endpoint selector.** LibreChat drops entries without a clear error. Check `name`, `baseURL`, `apiKey` and `models` for typos, that the block is indented under `endpoints.custom`, that no other endpoint has the same name, and that Docker mounts `librechat.yaml`. Restart after every change. **LibreChat does not start.** A schema error in the file disables custom endpoints and stops the server. Read `docker compose logs api`. **`Missing API Key for Tokens`, or 401 `missing_api_key` / `invalid_api_key`.** `TOKENS_API_KEY` is not in `.env` or the container was not restarted, or the key is wrong or rotated. Rotation stops the old secret immediately. **404 `model_not_found`.** The id is not a full Tokens id, or the title model is still the `gpt-3.5-turbo` default. Check ids with `GET https://tokens.bd/v1/models`. **403 `model_not_allowed_on_key`.** The model, often the `titleModel`, is not in the key's allowed list. **402 `insufficient_credits`, 403 `monthly_spend_cap_exceeded`.** Out of credits or the key hit its cap. See [billing](/dashboard/billing). **429 `rate_limited` or `concurrency_limit`.** Too many users or requests on one key; wait for `Retry-After`. Titles add one request per new chat. **404 on every request.** `baseURL` must be `https://tokens.bd/v1` with `/v1` and nothing after it, unless you set `directEndpoint: true`, which LibreChat's reference describes as using `baseURL` as the full completions URL. Every code is listed in [Errors](/docs/errors), with more fixes in [Troubleshooting](/docs/troubleshooting). Sources: LibreChat documentation, [Custom Endpoints quick start](https://www.librechat.ai/docs/quick_start/custom_endpoints), [librechat.yaml](https://www.librechat.ai/docs/configuration/librechat_yaml) and the [custom endpoint object reference](https://www.librechat.ai/docs/configuration/librechat_yaml/object_structure/custom_endpoint), plus the LibreChat repository's `librechat.example.yaml`, checked 11 October 2026. --- # Connect n8n to Tokens > Point n8n's OpenAI credential at Tokens by setting its Base URL, then use the OpenAI Chat Model with the AI Agent node: model IDs, the Responses API toggle, tool calling and spend limits. Section: Chat Apps & Automation. Page: https://tokens.bd/docs/n8n n8n is a workflow automation tool with AI nodes. Its OpenAI credential has a Base URL field, so the OpenAI Chat Model node, and through it the AI Agent node, can send OpenAI Chat Completions requests to Tokens at `https://tokens.bd/v1`. You set the address once on the credential and every node that uses that credential follows it. This guide was checked against n8n 2.42.6 (the stable release of 9 October 2026), using n8n's documentation and the source of the OpenAI Chat Model node (node version 1.3), in October 2026. It was checked against the documentation and source, not run end to end against a live Tokens key. ## What you need - A Tokens key. [API keys](/docs/api-keys) covers creating one with a spend cap. - A model ID from the [model catalog](/models). This guide uses `deepseek/deepseek-v4.1-flash`. - n8n, either n8n Cloud or self-hosted, with a workflow where you can add an AI Agent node. ## Where the Base URL lives The sources disagree, so here is what was checked. - **n8n's documentation does not mention it.** The [OpenAI credentials page](https://docs.n8n.io/integrations/builtin/credentials/openai/) lists only an API key and an Organization ID. The [OpenAI Chat Model page](https://docs.n8n.io/integrations/builtin/cluster-nodes/sub-nodes/n8n-nodes-langchain.lmchatopenai/) does not mention a base URL either. Not documented by n8n. - **n8n's source does have it.** The OpenAI credential defines a field named Base URL, described as "Override the default base URL for the API", with the default `https://api.openai.com/v1`. The credential's own connection test calls `/models` on that URL. - **The node used to have its own.** The OpenAI Chat Model node has a Base URL option only in node version 1. From node version 1.1 it is hidden, and the node reads the credential's Base URL instead. A workflow created today gets version 1.3, so use the credential. - **GitHub issue 14431** ("Allow setting custom base URL in OpenAI node") was closed as not planned on the day it was opened, with a maintainer reply that this is already possible. Later comments in that thread suggest using the provider's base URL directly. That matches the credential field. If your n8n is older than the version checked and the credential has no Base URL field, update n8n first. ## Set up the credential 1. In n8n, open Credentials and create a new credential of type **OpenAI**. 2. **API Key**: your Tokens key. 3. **Organization ID (optional)**: leave it empty. 4. **Base URL**: `https://tokens.bd/v1`. End it at `/v1`, with no trailing slash and no `/chat/completions`. 5. Leave **Add Custom Header** off. Tokens needs no extra header. 6. Save. n8n tests the credential by calling `GET /models` on the Base URL. A green result means the key and address are right. ```text API Key: tok_live_your_key Base URL: https://tokens.bd/v1 ``` :::note[Not the gateway credits] On n8n Cloud, supported nodes offer "Use Gateway credits" in the credential field. That is n8n's own billing and does not go through Tokens. Pick the credential you just made instead. ::: On self-hosted n8n, keep the key in n8n's credential store rather than in workflow JSON or a node field you might export. ## Use it in the AI Agent node 1. Add an **AI Agent** node, then click the Chat Model connector and add an **OpenAI Chat Model** sub-node. 2. **Credential to connect with**: the credential you made. 3. **Model**: choose **By ID** and type `deepseek/deepseek-v4.1-flash`. Details are in the next section. 4. If the node shows a **Use Responses API** toggle, turn it **off** (see below). 5. Connect tools or memory to the agent as usual. ### Model IDs with a slash Tokens IDs look like `provider/model`. In node version 1.2 and later the Model field has two modes, **From List** and **By ID**. - **From List** calls `GET /models` on your Base URL and lists what comes back. For a Base URL that is not api.openai.com, the node lists every model returned without filtering. The list shows only the models your key can use, so a key with an allowed-models list shows a shorter list. - **By ID** sends the text you type as the model name, unchanged. Use it if the list is empty or the model you want is missing. The slash needs no escaping. A third-party report says you can switch the field to expression mode and type the ID as a string. The n8n documentation does not describe that, and By ID makes it unnecessary, so this guide does not rely on it. On older workflows whose node version is 1.1 or lower, the Model field is a plain dropdown without a By ID mode. n8n's documentation does not say how to enter an ID there. Recreate the node, which gives you the current version. ### Turn off Use Responses API In node version 1.3 the OpenAI Chat Model has a **Use Responses API** toggle. n8n's documentation says Chat Completions is the default. The node source defaults the toggle to on for newly added nodes. Open the node and check the toggle yourself. Tokens serves `/v1/responses` only when the upstream behind the model supports it ([Responses API](/docs/responses)). Chat Completions works across the catalog, so switch the toggle off. The Built-in Tools (Web Search, File Search, Code Interpreter) belong to the Responses API and to OpenAI itself, so they do not apply to Tokens. ## Check that it works Test the key and model with curl first. If this works and n8n fails, the problem is in n8n's settings. ```bash export TOKENS_API_KEY=tok_live_your_key curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` In n8n, add a Chat Trigger and the AI Agent node, open the chat and send a short message. The agent answers, and the request appears in your Tokens dashboard usage analytics. To compare against the node, an HTTP Request node pointed at the same URL gives you the raw error body, which n8n's own documentation also suggests when the OpenAI node reports "Bad request". ## Tool calling The AI Agent node always works as a tools agent: it decides which connected tools to call, and n8n documents the OpenAI Chat Model as one of the models it supports. Tool calling needs a model that supports it, and no n8n setting turns it on. When a Base URL override is set, the OpenAI Chat Model node itself warns that models other than OpenAI's may not support tool calling or JSON response format. Check the model's page in the [model catalog](/models) before building an agent on it, and see [Tool calling](/docs/tool-calling) and [Choosing a model](/docs/choosing-a-model). The node sets strict tool calling off, which is what OpenAI-compatible backends expect. ## What costs credits in the background Every model call is billed to the key. An n8n agent can make many. - The agent calls the model again after each tool result. **Max Iterations** (default 10 in the Tools Agent options) is the limit on those rounds for one run. - Each round sends the conversation again, plus memory and tool output, so input tokens grow with every round. - The Chat Model has a **Max Retries** option. n8n's source uses 2 when it is unset, so a failing call can be sent more than once. - A Schedule or Webhook trigger starts the loop with nobody watching, and a workflow that processes many items runs the agent once per item. Use a dedicated key for n8n with a monthly spend cap and an allowed-models list ([API keys](/docs/api-keys)). When the cap is reached, requests fail with 403 `monthly_spend_cap_exceeded` instead of running up a bill. Lower Max Iterations for simple agents and set **Maximum Number of Tokens** in the node's options. ## Limits and what is not covered - Only the **OpenAI Chat Model** sub-node was checked. The standalone **OpenAI** node and the Embeddings OpenAI node use the same credential, but n8n does not document them against a custom Base URL, so they are not covered here. Tokens' embeddings endpoint works only for embedding models ([Models and usage](/docs/models-and-usage)). - Tokens does not support image, audio, file or assistants endpoints, so those operations in the OpenAI node return 404 `unsupported_endpoint`. - The Responses-only options and built-in tools are not usable (see above). ## Troubleshooting **The credential test fails with 401.** `missing_api_key` or `invalid_api_key` means the key was not pasted correctly, or it was revoked or rotated. Copy it again from the dashboard. **404 on the credential test or on every request.** Check that the Base URL is `https://tokens.bd/v1`: it needs `/v1`, no trailing slash, and no `/chat/completions` on the end. **404 `model_not_found`.** The model ID must be the full Tokens ID, for example `deepseek/deepseek-v4.1-flash`, not the part after the slash. Choose By ID and check it against `GET /v1/models`. **403 `model_not_allowed_on_key`.** The key has an allowed-models list that does not include this model. Use an allowed model or a key without the restriction. **400 `invalid_request`, or an error about the Responses API.** Turn off Use Responses API. For other 400s, use the HTTP Request node to see the whole error body, then check that the model supports the parameters the node sends. **The agent never calls its tools, or tool arguments come back broken.** The model probably does not handle tool calling well. Try another model from the catalog. **429 `rate_limited` or `concurrency_limit`.** A workflow processing many items sends requests in parallel. n8n's documentation suggests the Loop Over Items node with a Wait node at the end to send smaller batches. Wait for `Retry-After` seconds ([Rate limits](/docs/rate-limits)). **402 `insufficient_credits`.** Top up or renew in [billing](/dashboard/billing). Note that n8n's own help text for "Insufficient quota" is written for OpenAI accounts and does not apply to Tokens. Every code is listed in [Errors](/docs/errors). When you contact [support](/docs/support), include the `x-tokens-request-id` header from a curl run of the same call. Sources: [n8n OpenAI credentials](https://docs.n8n.io/integrations/builtin/credentials/openai/), [OpenAI Chat Model node](https://docs.n8n.io/integrations/builtin/cluster-nodes/sub-nodes/n8n-nodes-langchain.lmchatopenai/) and its [common issues](https://docs.n8n.io/integrations/builtin/cluster-nodes/sub-nodes/n8n-nodes-langchain.lmchatopenai/common-issues/), [AI Agent node](https://docs.n8n.io/integrations/builtin/cluster-nodes/root-nodes/n8n-nodes-langchain.agent/) and its [Tools Agent page](https://docs.n8n.io/integrations/builtin/cluster-nodes/root-nodes/n8n-nodes-langchain.agent/tools-agent/), [n8n issue 14431](https://github.com/n8n-io/n8n/issues/14431), and the n8n source tagged n8n@2.42.6 (the OpenAI credential and the OpenAI Chat Model node), checked October 2026. --- # Connect Dify to Tokens > Add Tokens models to Dify with the OpenAI-API-compatible provider plugin: the form fields, model IDs with a slash, the Function Call Type setting for agents, and Dify Cloud versus self-hosted. Section: Chat Apps & Automation. Page: https://tokens.bd/docs/dify Dify is an open-source platform for building LLM apps, chatbots, agents and workflows. Model providers in Dify are plugins. Its OpenAI-API-compatible provider sends OpenAI Chat Completions requests to any base URL, which for Tokens is `https://tokens.bd/v1`. You add each Tokens model to Dify by hand, as a custom model. This guide was checked against Dify 1.17.1 (released 10 September 2026) and version 0.0.68 of the OpenAI-API-compatible plugin, using Dify's documentation and the plugin's source in the langgenius/dify-official-plugins repository, in October 2026. It was checked against the documentation and source, not run end to end against a live Tokens key. ## What you need - A Tokens key. [API keys](/docs/api-keys) covers creating one with a spend cap. - A model ID from the [model catalog](/models). This guide uses `deepseek/deepseek-v4.1-flash`. - The model's context window and maximum output from its catalog page. `GET /v1/models` does not return them ([Models and usage](/docs/models-and-usage)), and Dify asks for both. - A Dify workspace where you are the owner or an admin. Dify allows only those roles to manage model providers. ## Install the provider plugin In Dify, open **Integrations** and then **Model Provider**, and install **OpenAI-API-compatible** from the Marketplace. Dify's documentation lists three plugin sources: the Marketplace, a public GitHub repository, and a local `.zip` upload. ## Add a Tokens model Open the OpenAI-API-compatible provider card and click **Add Model**. Dify's documentation describes Add Model as the way to add a model that is not in a provider's list. If a model with the same name and type already exists, Dify attaches the new key to it instead of creating a duplicate. Fill in the form. The labels below come from the plugin's source at version 0.0.68, and your Dify version may word them differently. | Field | Value | | -------------------------------------- | ---------------------------------------------------------------------- | | Model Type | `LLM` | | Model Name | `deepseek/deepseek-v4.1-flash` | | Model display name (optional) | Any label you want to see in Dify, such as `DeepSeek V4.1 Flash` | | API Key | Your Tokens key | | API Base URL | `https://tokens.bd/v1` | | model name for API endpoint (optional) | Leave empty (see below) | | Completion mode | `Chat` | | Model context size | The context window from the model's catalog page | | Upper bound for max tokens | The maximum output from the model's catalog page | | Function Call Type | `Tool Call` for a model that supports tools (see Tool calling below) | | Stream function calling | `Support` if you want streamed tool calls, when the model supports it | | Vision Support | `Support` only for a model that accepts images | Save the form. Dify checks the credentials when you save, and the plugin does this by sending a short chat request to the endpoint. If the form saves, the key, address and model name are accepted. A few notes on the form: - **API Base URL.** The plugin's placeholder is `https://api.openai.com/v1`, so the address ends in `/v1`. The plugin adds `chat/completions` itself. Do not add it. Some third-party guides call this field "API endpoint URL". - **The key is optional in the form**, but Tokens always needs one. Without it you get 401 `missing_api_key`. - **Model context size and Upper bound for max tokens both default to 4096.** If you leave the defaults, Dify limits long prompts and outputs on its side, whatever the model can do. Set both from the catalog page. - **Include Usage in Stream** is on by default in the plugin. Keep it on: it makes Dify ask for token counts in the stream. Tokens supports this ([Streaming](/docs/streaming)). - **Add each model separately.** The plugin does not fetch a model list from the endpoint. Repeat Add Model for every Tokens model you want to pick in Dify. ### Model IDs with a slash Tokens IDs look like `provider/model`. Dify's documentation does not say whether a slash is allowed in a model name, and the plugin's Model Name placeholder says "Enter full model name". The plugin sends the model name to the endpoint as is, unless **model name for API endpoint** is filled in, in which case it sends that value instead. 1. First try the full ID in **Model Name**. 2. If Dify rejects the name, put a short name in Model Name, for example `tokens-flash`, and the full Tokens ID in **model name for API endpoint**. This second field exists for that case: the name Dify shows can differ from the name the endpoint expects. Whichever you choose, the value that reaches Tokens must be the full ID, or you get 404 `model_not_found`. ## Check that it works Test the key and model with curl first. ```bash export TOKENS_API_KEY=tok_live_your_key curl -s https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}' ``` Then in Dify create an app, pick the model you added, and send a message in the preview. The request appears in your Tokens dashboard usage analytics. ## Tool calling Dify's Agent node and agent apps work with tools. Dify offers two agent strategies: - **Function Calling** uses the model's native tool calling and passes tool definitions through the `tools` parameter. Dify says to make sure the model supports function calling when you use it. - **ReAct** guides the model with structured prompts instead. Dify recommends it for models without native tool calling. For a custom model, the setting that tells Dify a model can call tools is **Function Call Type** in the Add Model form. Its default is `Not Support`. The plugin's source shows what the choices do. - **Tool Call** sends the OpenAI `tools` format. This is the one that matches [Tool calling](/docs/tool-calling) on Tokens. - **Function Call** sends the older `functions` format. Tokens' documentation does not cover it, so do not use it. - **Not Support** sends no tool definitions, so the Function Calling strategy cannot work with that model. Dify's documentation does not say what happens if you pick the Function Calling strategy for a model set to Not Support, so set it correctly rather than testing it. Tool calling still needs a model that supports it. Check the model's page in the [model catalog](/models) and see [Choosing a model](/docs/choosing-a-model). ## Dify Cloud and self-hosted The setup is the same in both. The difference is where the requests to Tokens come from: Dify's servers on Dify Cloud, your own deployment when self-hosted. Tokens is a public HTTPS address, so a self-hosted Dify needs outbound access to it, and nothing on the Tokens side needs to know which one you use. - **Dify's AI credits** are Dify's own billing. They do not pay for a Tokens model. Tokens usage is billed to your Tokens account. - On a self-hosted Dify, if you open the Marketplace outside Dify to install the plugin, set your deployment's URL under **Install Preference** first. Not confirmed from Dify's documentation: the network requirements of a self-hosted plugin daemon, and whether any Dify Cloud plan limits custom model providers. Check Dify's self-hosting documentation and your plan. ## What costs credits in the background Every request Dify sends to a Tokens model is billed to the key. - Each LLM node in a workflow is one request per run. A workflow with several LLM nodes, or one inside an iteration or loop, makes one request per pass. - An agent calls the model again after each tool result. Dify's documentation describes **Max Iterations** as a safety limit that prevents infinite loops, and suggests 3 to 5 for simple tasks and 10 to 15 for complex research. Keep it as low as the task allows. - Each round sends the conversation again, so input tokens grow with every round. - Dify's documentation does not list other features that call a model automatically. Check which of your app's features use a model, and which model they use, in the app and workspace settings. Use a dedicated key for Dify with a monthly spend cap and an allowed-models list ([API keys](/docs/api-keys)). When the cap is reached, requests fail with 403 `monthly_spend_cap_exceeded` instead of running up a bill. ## Limits and what is not covered - This page covers LLM models only. The plugin also offers text embedding, rerank, speech and text-to-speech types. Tokens' embeddings endpoint works only for embedding models, and Tokens has no audio endpoints ([Models and usage](/docs/models-and-usage)), so the others are not covered here. - The plugin has an **API Type** setting with a Responses API choice. Leave it on Chat Completions. Responses works on Tokens only when the upstream supports it ([Responses API](/docs/responses)). - Dify's built-in model providers, such as its OpenAI provider, are not used here. ## Troubleshooting **Saving the model fails with a credentials error.** The message contains the status code and response body from Tokens, so read it. 401 `invalid_api_key` means a wrong or revoked key. 404 means the API Base URL or the model name is wrong. Compare with the curl test above. **404 on every request.** Check that API Base URL is `https://tokens.bd/v1`. It needs `/v1`, and `/chat/completions` must not be added. **404 `model_not_found`.** The model name that reaches Tokens is not the full ID. Check Model Name and "model name for API endpoint" against `GET /v1/models`. **403 `model_not_allowed_on_key`.** The key's allowed-models list does not include this model. Use an allowed model or another key. **The Agent node refuses the model, or never calls tools.** Set Function Call Type to `Tool Call` on the model, or switch the agent to the ReAct strategy. If it still fails to call tools, the model is probably a poor fit for tool use. **400 `invalid_request` mentioning the `user` field.** The plugin has a **User Identity Support** setting, described as whether the endpoint accepts the optional top-level `user` parameter. Set it to Not Support to leave the field out. **Long prompts are cut short or outputs stop early.** Check Model context size and Upper bound for max tokens. Both default to 4096. **429 `rate_limited`, `concurrency_limit` or `window_exhausted`.** Wait for `Retry-After` seconds, or see [Rate limits](/docs/rate-limits). **402 `insufficient_credits`**: top up in [billing](/dashboard/billing). Every code is listed in [Errors](/docs/errors). When you contact [support](/docs/support), include the `x-tokens-request-id` header from a curl run of the same call. Sources: [Dify Model Providers](https://docs.dify.ai/en/use-dify/workspace/model-providers), [Integrations](https://docs.dify.ai/en/use-dify/workspace/plugins) and the [self-hosted Integrations page](https://docs.dify.ai/en/self-host/use-dify/workspace/plugins), [Agent node](https://docs.dify.ai/en/use-dify/nodes/agent), the [OpenAI-API-compatible plugin source and README](https://github.com/langgenius/dify-official-plugins/tree/main/models/openai_api_compatible) (version 0.0.68) and the [Dify plugin SDK's OpenAI-compatible model class](https://github.com/langgenius/dify-plugin-sdks/blob/main/src/dify_plugin/interfaces/model/openai_compatible/llm.py), checked October 2026. --- # cURL > Call the Tokens API with cURL: chat completions, streaming, Anthropic Messages, listing models, checking usage and reading request IDs from error responses. Section: SDKs & Libraries. Page: https://tokens.bd/docs/curl cURL is the fastest way to check that a key works and see exactly what the Tokens API returns, with no SDK in the way. Every example below reads your key from `TOKENS_API_KEY`, so set that first. ```bash export TOKENS_API_KEY="tok_live_your_key" ``` On Windows PowerShell, use `$env:TOKENS_API_KEY = "tok_live_your_key"` and call `curl.exe` rather than `curl`. In Windows PowerShell 5.1, `curl` is an alias for `Invoke-WebRequest`, which takes different flags. The `\` line continuations below are for bash and zsh; in PowerShell put the command on one line or use a backtick. ## Send a chat completion with cURL The OpenAI-compatible base URL is `https://tokens.bd/v1`. Authenticate with `Authorization: Bearer`. ```bash curl https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "messages": [ {"role": "system", "content": "You are a terse code reviewer."}, {"role": "user", "content": "What does `set -euo pipefail` do?"} ], "max_tokens": 300 }' ``` The response is a standard chat completion object: the reply is in `choices[0].message.content` and token counts are in `usage`. Field-by-field details are in [Chat Completions](/docs/chat-completions). `model` is required. Request bodies can be up to 10 MB; anything larger returns `413`. ## Stream tokens with -N Add `"stream": true` and pass `-N` (`--no-buffer`) so cURL prints each server-sent event as it arrives instead of waiting for the whole response. ```bash curl -N https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "messages": [{"role": "user", "content": "Count from 1 to 10, one number per line."}], "stream": true, "stream_options": {"include_usage": true} }' ``` You get `data: {...}` lines with `choices[0].delta.content`, then `data: [DONE]`. The usage chunk (with an empty `choices` array) only appears because the request sets `stream_options.include_usage`. Leave it out and the stream carries no token counts. More on the event format in [Streaming](/docs/streaming). ## Call the Anthropic Messages API The same key works on the Anthropic-compatible endpoint. Send it as `x-api-key` (or `Authorization: Bearer`, both are accepted). ```bash curl https://tokens.bd/v1/messages \ -H "x-api-key: $TOKENS_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "max_tokens": 300, "messages": [{"role": "user", "content": "Explain a git rebase in two sentences."}] }' ``` The model doesn't have to be an Anthropic model. The gateway translates the request for whichever provider serves the model you name. Add `"stream": true` and `-N` to stream Anthropic-style events. See [Messages](/docs/messages). :::note SDKs and tools built for Anthropic usually want the base URL **without** `/v1` (`https://tokens.bd`), because they append `/v1/messages` themselves. With raw cURL you type the full path. ::: ## List the models your key can use ```bash curl https://tokens.bd/v1/models \ -H "Authorization: Bearer $TOKENS_API_KEY" ``` This returns the models available to this key right now, taking into account your plan, your wallet balance and the key's allowed-models list. Copy model IDs from here rather than typing them from memory. The full catalog with prices is at [/models](/models). To pull out just the IDs with `jq`: ```bash curl -s https://tokens.bd/v1/models \ -H "Authorization: Bearer $TOKENS_API_KEY" | jq -r '.data[].id' ``` ## Check plan usage and wallet balance `GET /v1/tokens/usage` is specific to Tokens. It reports your plan, active usage windows, wallet balance and this key's limits. It is read-only and is never billed. ```bash curl -s https://tokens.bd/v1/tokens/usage \ -H "Authorization: Bearer $TOKENS_API_KEY" | jq ``` The response has these top-level fields: | Field | What it holds | | --------- | ----------------------------------------------------------------------------------------------------------------------------- | | `plan` | Plan name, tier and `periodEnd`, or `null` if you have no active plan | | `windows` | Active usage windows (`session_5h`, `weekly`, `monthly`) with `unit`, `limit`, `used`, `remaining`, `percentUsed`, `resetsAt` | | `wallet` | `balanceUsd`, or `null` | | `key` | `monthlySpendCapUsd` and `allowedModels` for the key you used | This is the call to make when a request fails with `429 window_exhausted` and you want to know when the window resets. ## Debug errors with -i and x-tokens-request-id Pass `-i` to print response headers along with the body. Every response carries `x-tokens-request-id`, which is the first thing support will ask for. ```bash curl -i https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "not-a-real/model", "messages": [{"role": "user", "content": "hi"}]}' ``` ```http HTTP/2 404 content-type: application/json x-tokens-request-id: 7f3c... x-request-id: ... {"error":{"message":"...","type":"...","code":"model_not_found","param":null,"request_id":"7f3c..."}} ``` Read `error.code`, not just the HTTP status. A `403` can mean an inactive key, a model that isn't on the key's allow-list, or a spend cap you've hit, and each has a different fix. On `429`, check the `Retry-After` header. Tokens doesn't send `X-RateLimit-*` headers. For one-line status checks in scripts, `-w` prints just the code: ```bash curl -s -o /dev/null -w "%{http_code}\n" https://tokens.bd/v1/models \ -H "Authorization: Bearer $TOKENS_API_KEY" ``` `200` means the key is valid. `401` means it's missing or wrong. The full list of codes and what to do about each is in [Errors](/docs/errors) and [Troubleshooting](/docs/troubleshooting). ## What isn't supported The gateway serves `/v1/chat/completions`, `/v1/completions` (legacy), `/v1/messages`, `/v1/responses`, `/v1/embeddings` (embedding models only), `/v1/models` and `/v1/tokens/usage`. Image generation, audio, files, batches, assistants, fine-tuning and moderations endpoints return `404 unsupported_endpoint`. Sending an image *to* a model as part of a chat request works: see [Image input](/docs/vision). --- # Python (OpenAI SDK) > Use the official openai Python package with Tokens: client setup, sync and async calls, streaming, tool calls, timeouts, retries and error handling. Section: SDKs & Libraries. Page: https://tokens.bd/docs/python The official `openai` Python package works with Tokens unchanged. You point it at `https://tokens.bd/v1`, pass your Tokens key, and use model IDs from the Tokens catalog. Everything else (streaming, tool calls, async) is the SDK you already know. ## Install and configure the OpenAI Python SDK ```bash pip install --upgrade openai export TOKENS_API_KEY="tok_live_your_key" ``` ```python title="hello.py" import os from openai import OpenAI client = OpenAI( base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], ) completion = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", messages=[ {"role": "system", "content": "Answer in one short paragraph."}, {"role": "user", "content": "When should I use a dataclass instead of a dict?"}, ], max_tokens=400, ) print(completion.choices[0].message.content) print(completion.usage) ``` Use the exact model ID from [/models](/models) or `client.models.list()`. IDs follow a `provider/model` pattern, and a typo returns `404 model_not_found`. ### Alternative: OPENAI_BASE_URL and OPENAI_API_KEY The SDK reads `OPENAI_BASE_URL` and `OPENAI_API_KEY` from the environment when you don't pass them, so existing code can switch to Tokens with no edits: ```bash export OPENAI_BASE_URL="https://tokens.bd/v1" export OPENAI_API_KEY="$TOKENS_API_KEY" ``` ```python from openai import OpenAI client = OpenAI() # picks up OPENAI_BASE_URL and OPENAI_API_KEY ``` This is convenient, but those variables are global. Every tool in that shell that reads them (Aider, some agents, other scripts) will also send traffic to Tokens. If you want that scoped, pass `base_url` and `api_key` explicitly as in the first example. ## Stream responses ```python stream = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": "Write a haiku about merge conflicts."}], stream=True, stream_options={"include_usage": True}, ) for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="", flush=True) if chunk.usage: print(f"\n\n{chunk.usage.prompt_tokens} in, {chunk.usage.completion_tokens} out") ``` The `if chunk.choices` guard matters. With `include_usage` on, the final chunk has an empty `choices` list and carries only `usage`. Without `include_usage`, streams carry no token counts at all. See [Streaming](/docs/streaming). ## Use the async client `AsyncOpenAI` takes the same arguments. Use it in FastAPI, aiohttp or anything else running an event loop. ```python import asyncio import os from openai import AsyncOpenAI client = AsyncOpenAI( base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], ) async def main() -> None: stream = await client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": "Name three uses for asyncio.Semaphore."}], stream=True, ) async for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="", flush=True) asyncio.run(main()) ``` If you fan out many requests with `asyncio.gather`, cap them with a semaphore. Tokens limits concurrent requests per account (the limit comes from your plan), and going over it returns `429 concurrency_limit`. See [Rate limits](/docs/rate-limits). ## Tool calls Tool calling works on models that support it. Check the model's page in [/models](/models) before you rely on it. ```python import json tools = [{ "type": "function", "function": { "name": "get_weather", "description": "Current weather for a city", "parameters": { "type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"], }, }, }] messages = [{"role": "user", "content": "Is it raining in Dhaka?"}] first = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", messages=messages, tools=tools ) msg = first.choices[0].message if msg.tool_calls: messages.append(msg) for call in msg.tool_calls: args = json.loads(call.function.arguments) result = {"city": args["city"], "condition": "light rain", "temp_c": 29} # your real lookup here messages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)}) final = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", messages=messages, tools=tools ) print(final.choices[0].message.content) ``` The full request and response shapes are in [Tool calling](/docs/tool-calling). ## Configure timeouts and retries The SDK retries some failures (connection errors, 408, 409, 429 and 5xx) twice by default with backoff. Its default timeout is 10 minutes. Both can be set per client or per call: ```python import os import httpx from openai import OpenAI client = OpenAI( base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], timeout=httpx.Timeout(120.0, connect=10.0), max_retries=3, ) # Override for one call client.with_options(timeout=30.0, max_retries=0).chat.completions.create(...) ``` Some notes specific to Tokens: - The gateway already fails over to another upstream source on server errors, 429, timeouts and connection errors before it replies, so a 5xx you see means that failover didn't help. A couple of SDK retries is plenty. - Long generations are fine. The gateway waits up to 600 seconds for response headers. For long outputs, stream so you aren't holding an idle connection. - Retrying `429 window_exhausted` won't help until the window resets. `Retry-After` gives the seconds until it does, and that can be hours. ## Handle Tokens error codes Errors use the OpenAI shape, so `openai` raises its usual exception classes. The useful part for Tokens is `error.code`. ```python import openai try: client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": "hi"}], ) except openai.APIStatusError as e: request_id = e.response.headers.get("x-tokens-request-id") print(e.status_code, e.code, e.message, request_id) if e.code == "insufficient_credits": print("Top up at https://tokens.bd/dashboard/billing") except openai.APIConnectionError as e: print("Network problem:", e) ``` Log `x-tokens-request-id` with every failure. Support can trace a request from that ID. The codes you're likely to meet are `invalid_api_key` (401), `model_not_allowed_on_key` and `tier_permission_denied` (403), `insufficient_credits` (402), and `rate_limited`, `concurrency_limit` and `window_exhausted` (429). [Errors](/docs/errors) lists them all, and [Troubleshooting](/docs/troubleshooting) gives a fix for each. :::warning Keep the key in an environment variable or a secrets manager. If it ends up in a committed file or a notebook you've shared, rotate it in [/dashboard/keys](/dashboard/keys). The old secret stops working immediately. ::: --- # Node.js and TypeScript (OpenAI SDK) > Use the official openai npm package with Tokens from Node.js and TypeScript: client setup, streaming, error handling with APIError, and a Next.js route handler that keeps your key on the server. Section: SDKs & Libraries. Page: https://tokens.bd/docs/nodejs The official `openai` npm package talks to Tokens with two settings changed: `baseURL` and `apiKey`. This page covers a TypeScript setup, streaming, reading Tokens error codes from `APIError`, and a Next.js route handler so your key never reaches the browser. ## Install the OpenAI Node.js SDK ```bash npm install openai export TOKENS_API_KEY="tok_live_your_key" ``` The examples use ES modules and top-level `await`, so run them as `.mjs` or `.ts` files (for example with `npx tsx hello.ts`). They assume a current major version of `openai` (v5 or later). ```ts title="lib/tokens.ts" import OpenAI from "openai"; export const tokens = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY, timeout: 120_000, // ms; SDK default is 10 minutes maxRetries: 2, // SDK default; retries connection errors, 408, 409, 429, 5xx }); ``` ```ts title="hello.ts" import { tokens } from "./lib/tokens"; const completion = await tokens.chat.completions.create({ model: "deepseek/deepseek-v4.1-flash", messages: [ { role: "system", content: "Reply in one short paragraph." }, { role: "user", content: "What's the difference between interface and type in TypeScript?" }, ], max_tokens: 400, }); console.log(completion.choices[0]?.message.content); console.log(completion.usage); ``` If you'd rather not touch code, the SDK also reads `OPENAI_BASE_URL` and `OPENAI_API_KEY` from the environment. Keep in mind those variables affect every OpenAI-based tool in that environment. Copy model IDs from [/models](/models) or `await tokens.models.list()`. Don't guess them. ## Stream chat completions ```ts const stream = await tokens.chat.completions.create({ model: "deepseek/deepseek-v4.1-flash", messages: [{ role: "user", content: "List five git commands I should know, one per line." }], stream: true, stream_options: { include_usage: true }, }); for await (const chunk of stream) { const text = chunk.choices[0]?.delta?.content; if (text) process.stdout.write(text); if (chunk.usage) { console.log(`\n${chunk.usage.prompt_tokens} in, ${chunk.usage.completion_tokens} out`); } } ``` The last chunk has an empty `choices` array when `include_usage` is on, which is why the optional chaining is there. To stop early, `break` out of the loop or call `stream.controller.abort()`. More in [Streaming](/docs/streaming). ## Handle errors with APIError Non-2xx responses throw a subclass of `OpenAI.APIError`. Tokens puts a specific `code` on every error, and that tells you far more than the HTTP status. ```ts import OpenAI from "openai"; import { tokens } from "./lib/tokens"; try { await tokens.chat.completions.create({ model: "deepseek/deepseek-v4.1-flash", messages: [{ role: "user", content: "hi" }], }); } catch (err) { if (err instanceof OpenAI.APIConnectionError) { console.error("Network problem:", err.message); } else if (err instanceof OpenAI.APIError) { const requestId = err.headers?.get("x-tokens-request-id"); console.error(err.status, err.code, err.message, requestId); switch (err.code) { case "insufficient_credits": // 402: top up at /dashboard/billing break; case "window_exhausted": // 429: plan window used up; Retry-After header = seconds until reset console.error("Resets in", err.headers?.get("retry-after"), "s"); break; case "model_not_allowed_on_key": case "tier_permission_denied": // 403: key allow-list or plan doesn't cover this model break; } } else { throw err; } } ``` Two things to know: - `err.headers` is a standard `Headers` object in current SDK versions, so use `.get()`. Always log `x-tokens-request-id`. Support needs it to find your request. - The SDK automatically retries `429` and `5xx`. The gateway has already tried failing over to another upstream source before it returns `502`, `503` or `504`. Retrying a `window_exhausted` won't help until `Retry-After` has passed, and that can be hours, so set `maxRetries` low if you'd rather fail fast. Every code is listed in [Errors](/docs/errors), and [Troubleshooting](/docs/troubleshooting) gives the fix for each. ## Keep the key on the server: a Next.js route handler Tokens doesn't send CORS headers, so calling it straight from browser JavaScript fails, and it would expose your key anyway. Put the call in a route handler and have your frontend call that. ```ts title="app/api/chat/route.ts" import OpenAI from "openai"; export const runtime = "nodejs"; const tokens = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY, // server-only: no NEXT_PUBLIC_ prefix }); type ChatMessage = { role: "user" | "assistant"; content: string }; export async function POST(req: Request) { const body = (await req.json()) as { messages?: ChatMessage[] }; if (!Array.isArray(body.messages) || body.messages.length === 0) { return Response.json({ error: "messages required" }, { status: 400 }); } try { const stream = await tokens.chat.completions.create( { model: "deepseek/deepseek-v4.1-flash", // pick on the server, not from the client messages: body.messages.slice(-20), max_tokens: 1024, stream: true, }, { signal: req.signal } // stop paying for tokens if the user navigates away ); const encoder = new TextEncoder(); const text = new ReadableStream({ async start(controller) { try { for await (const chunk of stream) { const delta = chunk.choices[0]?.delta?.content; if (delta) controller.enqueue(encoder.encode(delta)); } controller.close(); } catch (e) { controller.error(e); } }, }); return new Response(text, { headers: { "Content-Type": "text/plain; charset=utf-8", "Cache-Control": "no-store" }, }); } catch (err) { if (err instanceof OpenAI.APIError) { return Response.json( { error: err.code ?? "upstream_error", requestId: err.headers?.get("x-tokens-request-id") }, { status: err.status ?? 502 } ); } throw err; } } ``` The client side is a plain `fetch` that reads the body as a stream: ```ts const res = await fetch("/api/chat", { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({ messages: [{ role: "user", content: "Hello" }] }), }); const reader = res.body!.pipeThrough(new TextDecoderStream()).getReader(); for (;;) { const { value, done } = await reader.read(); if (done) break; console.log(value); } ``` A few choices in that handler are deliberate. The model and `max_tokens` are set on the server so a visitor can't switch your app to an expensive model. Message history is trimmed. You should also add your own authentication and per-user rate limiting in front of this route, because anyone who can reach it is spending your balance. A key with a monthly spend cap ([API keys](/docs/api-keys)) puts a hard ceiling on the damage. If you're on the Vercel AI SDK, [Vercel AI SDK](/docs/vercel-ai-sdk) shows the same pattern with `streamText`. --- # Anthropic SDK (Python and TypeScript) > Point the official Anthropic Python and TypeScript SDKs at Tokens with base URL https://tokens.bd, then call messages.create and stream with any model in the Tokens catalog. Section: SDKs & Libraries. Page: https://tokens.bd/docs/anthropic-sdk Tokens exposes an Anthropic-compatible Messages API, so the official Anthropic SDKs work with one change: the base URL. If your code is already written against `messages.create`, you can keep it and switch providers by changing the `model` string. ## Use the Anthropic SDK base URL: https://tokens.bd Set the base URL to `https://tokens.bd`, **without** `/v1`. The SDKs append `/v1/messages` themselves. If you set `https://tokens.bd/v1`, requests go to `/v1/v1/messages` and fail with a `404`. | Setting | Value | | -------- | ------------------------------------------------------------------- | | Base URL | `https://tokens.bd` | | API key | your Tokens key (`tok_live_...`) | | Model | any ID from [/models](/models), e.g. `deepseek/deepseek-v4.1-flash` | The SDK sends the key as `x-api-key`. Tokens accepts that or `Authorization: Bearer`. ## Which models work Any model in the Tokens catalog, not just Claude models. When you send an Anthropic-format request for a model from another provider, the gateway translates the request and the response, so your code receives normal Anthropic `Message` objects and stream events either way. Check which IDs your key can use with `GET https://tokens.bd/v1/models` (see [cURL](/docs/curl)). Features beyond plain text depend on the model you pick. Tool use, extended thinking and prompt caching only work where the underlying model and provider support them. Check the model's page in [/models](/models) before you rely on one of them. ## Python ```bash pip install --upgrade anthropic export TOKENS_API_KEY="tok_live_your_key" ``` ```python title="hello_anthropic.py" import os import anthropic client = anthropic.Anthropic( base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"], ) message = client.messages.create( model="deepseek/deepseek-v4.1-flash", max_tokens=512, system="You are a concise senior engineer.", messages=[{"role": "user", "content": "When is a B-tree index the wrong choice?"}], ) for block in message.content: if block.type == "text": print(block.text) print(message.usage) ``` `max_tokens` is required by the Messages API, as it is with Anthropic directly. ### Stream with the Python SDK ```python with client.messages.stream( model="deepseek/deepseek-v4.1-flash", max_tokens=1024, messages=[{"role": "user", "content": "Write a bash one-liner that finds the 10 largest files."}], ) as stream: for text in stream.text_stream: print(text, end="", flush=True) final = stream.get_final_message() print("\n", final.usage) ``` The async client is `anthropic.AsyncAnthropic` with the same arguments. Use `async with client.messages.stream(...)` and `async for text in stream.text_stream`. ## TypeScript The Anthropic TypeScript SDK needs Node.js 20 or later. ```bash npm install @anthropic-ai/sdk ``` ```ts title="hello-anthropic.ts" import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic({ baseURL: "https://tokens.bd", apiKey: process.env.TOKENS_API_KEY, }); const message = await client.messages.create({ model: "deepseek/deepseek-v4.1-flash", max_tokens: 512, messages: [{ role: "user", content: "Explain optimistic locking in three sentences." }], }); for (const block of message.content) { if (block.type === "text") console.log(block.text); } ``` ### Stream with the TypeScript SDK ```ts const stream = client.messages .stream({ model: "deepseek/deepseek-v4.1-flash", max_tokens: 1024, messages: [{ role: "user", content: "Draft a short commit message for a null-check fix." }], }) .on("text", (text) => process.stdout.write(text)); const final = await stream.finalMessage(); console.log("\n", final.usage); ``` If you only need the raw events and want to use less memory, `client.messages.create({ ..., stream: true })` returns an async iterable instead. ## Environment variables instead of code Both SDKs read `ANTHROPIC_API_KEY` and `ANTHROPIC_BASE_URL` when you construct the client with no arguments: ```bash export ANTHROPIC_BASE_URL="https://tokens.bd" export ANTHROPIC_API_KEY="$TOKENS_API_KEY" ``` Be careful with this in a shell where you also run Claude Code or other Anthropic-based tools, because they read the same variables. Claude Code has its own setup with `ANTHROPIC_AUTH_TOKEN` in `~/.claude/settings.json`, covered in [Claude Code](/docs/claude-code). ## Errors and retries The SDKs raise their usual `APIError` subclasses (`AuthenticationError` for 401, `PermissionDeniedError` for 403, `RateLimitError` for 429, and so on). Every error on `/v1/messages` uses the Anthropic error shape, `{"type": "error", "error": {"type", "message", "code"}}`. Errors from Tokens itself (bad key, credits, plan limits) also carry `error.code`, so read it when you need to tell `insufficient_credits` from `window_exhausted`. ```python try: client.messages.create(model="deepseek/deepseek-v4.1-flash", max_tokens=64, messages=[{"role": "user", "content": "hi"}]) except anthropic.APIStatusError as e: print(e.status_code, e.response.headers.get("x-tokens-request-id"), e.body) ``` Include the `x-tokens-request-id` header value in any support ticket. Both SDKs retry 429 and 5xx twice by default. That's reasonable here, but retrying `429 window_exhausted` won't succeed until the time in `Retry-After` has passed. Codes and fixes are in [Errors](/docs/errors), and the request format is in [Messages](/docs/messages). --- # Vercel AI SDK > Use Tokens as a provider in the Vercel AI SDK with createOpenAICompatible from @ai-sdk/openai-compatible, then call generateText and streamText, including a Next.js route. Section: SDKs & Libraries. Page: https://tokens.bd/docs/vercel-ai-sdk The Vercel AI SDK talks to Tokens through `@ai-sdk/openai-compatible`, the provider package for any OpenAI-compatible API. Create the provider once with `createOpenAICompatible`, and then `generateText`, `streamText` and tool calling work the same as they do with the first-party providers. OpenCode uses this same package under the hood, so if you've set Tokens up there ([Tokens CLI](/docs/tokens-cli) does it for you), this will look familiar. ## Install the packages ```bash npm install ai @ai-sdk/openai-compatible zod export TOKENS_API_KEY="tok_live_your_key" ``` The examples target AI SDK 7 (`ai@7`, `@ai-sdk/openai-compatible@3`), which needs Node.js 22 or later. On AI SDK 5 or 6, the code is the same except that you pass the system prompt as `system` instead of `instructions`. ## Create the Tokens provider with createOpenAICompatible ```ts title="lib/tokens.ts" import { createOpenAICompatible } from "@ai-sdk/openai-compatible"; export const tokens = createOpenAICompatible({ name: "tokens", baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY, includeUsage: true, // ask for token counts on streamed responses }); ``` `apiKey` is sent as `Authorization: Bearer `. `includeUsage: true` makes the provider set `stream_options.include_usage` on streaming calls. Without it, Tokens streams carry no usage data, and `usage` ends up empty after a `streamText` call. Models are addressed by their Tokens ID: `tokens("deepseek/deepseek-v4.1-flash")`. Copy IDs from [/models](/models) or `GET /v1/models` rather than typing them. ## generateText ```ts title="summarize.ts" import { generateText } from "ai"; import { tokens } from "./lib/tokens"; const { text, usage, finishReason } = await generateText({ model: tokens("deepseek/deepseek-v4.1-flash"), instructions: "You summarize pull requests for busy reviewers.", prompt: "Summarize: refactored auth middleware to use a single session lookup; removed two duplicate DB calls.", maxOutputTokens: 300, }); console.log(text); console.log(finishReason, usage.inputTokens, usage.outputTokens); ``` ## streamText `streamText` isn't awaited. It returns immediately, and you consume `textStream`. ```ts title="stream.ts" import { streamText } from "ai"; import { tokens } from "./lib/tokens"; const result = streamText({ model: tokens("deepseek/deepseek-v4.1-flash"), prompt: "Explain the difference between TCP and UDP for a junior developer.", maxOutputTokens: 600, onError: ({ error }) => console.error(error), }); for await (const delta of result.textStream) { process.stdout.write(delta); } console.log("\n", await result.usage); ``` Errors that happen during a stream are delivered through `onError` rather than thrown, so wire it up. Otherwise a `402 insufficient_credits` just looks like an empty response. ## Stream from a Next.js route handler Tokens doesn't allow browser calls (no CORS headers), and you don't want your key in client code anyway. Run `streamText` in a route handler: ```ts title="app/api/chat/route.ts" import { streamText } from "ai"; import { tokens } from "@/lib/tokens"; export async function POST(req: Request) { const { prompt } = (await req.json()) as { prompt?: string }; if (!prompt) return Response.json({ error: "prompt required" }, { status: 400 }); const result = streamText({ model: tokens("deepseek/deepseek-v4.1-flash"), prompt, maxOutputTokens: 1024, abortSignal: req.signal, }); return result.toTextStreamResponse(); } ``` `toTextStreamResponse()` sends plain text chunks that you can read with `fetch` and a stream reader. If your frontend uses `useChat` from `@ai-sdk/react`, take `messages` from the request body instead of `prompt`, convert them with `convertToModelMessages`, and return `result.toUIMessageStreamResponse()`. The Tokens part (the provider) doesn't change. Choose the model and output limit on the server so visitors can't pick an expensive model for you. Put authentication in front of the route too. ## Tool calling Tools work on models that support them. Check the model page in [/models](/models) first. ```ts import { generateText, tool, stepCountIs } from "ai"; import { z } from "zod"; import { tokens } from "./lib/tokens"; const { text, steps } = await generateText({ model: tokens("deepseek/deepseek-v4.1-flash"), tools: { getWeather: tool({ description: "Current weather for a city", inputSchema: z.object({ city: z.string() }), execute: async ({ city }) => ({ city, condition: "light rain", tempC: 29 }), }), }, stopWhen: stepCountIs(3), prompt: "Do I need an umbrella in Dhaka right now?", }); console.log(text, `(${steps.length} steps)`); ``` If a model doesn't support tools, the upstream provider rejects the request or the model ignores the tools and answers in text. Switch to a model that lists tool support. Wire-level details are in [Tool calling](/docs/tool-calling). ## Errors HTTP errors arrive as `APICallError` (exported from `ai`), which carries `statusCode`, `responseHeaders` and `responseBody`. The Tokens error code is inside `responseBody` as JSON, under `error.code`. ```ts import { APICallError, generateText } from "ai"; import { tokens } from "./lib/tokens"; try { await generateText({ model: tokens("deepseek/deepseek-v4.1-flash"), prompt: "hi" }); } catch (err) { if (APICallError.isInstance(err)) { console.error(err.statusCode, err.responseHeaders?.["x-tokens-request-id"], err.responseBody); } } ``` The AI SDK retries failed calls twice by default (`maxRetries`). Retrying won't fix a `429 window_exhausted`, which only clears when the plan window resets. Codes and fixes are in [Errors](/docs/errors) and [Troubleshooting](/docs/troubleshooting). --- # LangChain and LiteLLM > Use Tokens from LangChain (Python and JavaScript) with ChatOpenAI and a custom base URL, and from LiteLLM as an SDK or proxy with the openai/ model prefix. Section: SDKs & Libraries. Page: https://tokens.bd/docs/langchain LangChain and LiteLLM both have an OpenAI client that takes a custom base URL, and that's all Tokens needs. This page shows the LangChain `ChatOpenAI` setup in Python and JavaScript, then LiteLLM as a Python SDK and as a proxy. All examples assume your key is exported: ```bash export TOKENS_API_KEY="tok_live_your_key" ``` ## LangChain Python: ChatOpenAI with base_url ```bash pip install --upgrade langchain-openai ``` ```python title="chain.py" import os from langchain_openai import ChatOpenAI llm = ChatOpenAI( model="deepseek/deepseek-v4.1-flash", base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], temperature=0.2, max_tokens=500, timeout=120, max_retries=2, ) reply = llm.invoke([ ("system", "You are a precise SQL reviewer."), ("human", "Is SELECT * in a view a problem? Two sentences."), ]) print(reply.content) print(reply.usage_metadata) ``` `model` takes the Tokens model ID exactly as listed in [/models](/models) or `GET /v1/models`. LangChain passes it through unchanged. ### Streaming ```python llm = ChatOpenAI( model="deepseek/deepseek-v4.1-flash", base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], stream_usage=True, # sets stream_options.include_usage so you get token counts ) for chunk in llm.stream("Give me three naming tips for Python modules."): print(chunk.content, end="", flush=True) ``` Without `stream_usage=True`, streamed responses from Tokens carry no usage data. See [Streaming](/docs/streaming). ### Tools and chains `bind_tools`, `with_structured_output` and LCEL chains all work, since they only depend on the OpenAI chat format. Tool calling and structured output still need a model that supports them. Check the model page in [/models](/models) before building on one, and see [Tool calling](/docs/tool-calling). ```python from langchain_core.tools import tool @tool def get_weather(city: str) -> str: """Current weather for a city.""" return f"{city}: light rain, 29C" llm_with_tools = llm.bind_tools([get_weather]) msg = llm_with_tools.invoke("Do I need an umbrella in Dhaka?") print(msg.tool_calls) ``` ## LangChain JS: ChatOpenAI with configuration.baseURL ```bash npm install @langchain/openai @langchain/core ``` In JavaScript the base URL goes inside `configuration`, which is passed straight to the underlying OpenAI client. ```ts title="chain.ts" import { ChatOpenAI } from "@langchain/openai"; const llm = new ChatOpenAI({ model: "deepseek/deepseek-v4.1-flash", apiKey: process.env.TOKENS_API_KEY, configuration: { baseURL: "https://tokens.bd/v1" }, temperature: 0.2, maxTokens: 500, streamUsage: true, }); const reply = await llm.invoke("Explain debouncing vs throttling in two sentences."); console.log(reply.content); for await (const chunk of await llm.stream("List three uses for a WeakMap.")) { process.stdout.write(String(chunk.content)); } ``` Run LangChain code on a server or in a CLI, not in browser bundles. Tokens doesn't send CORS headers, and the key would be visible to anyone who opens dev tools. ## LiteLLM: use the openai/ prefix with api_base LiteLLM routes a request by the prefix of the model name. `openai/` tells it to use its generic OpenAI-compatible client against whatever `api_base` you give it. LiteLLM strips that first `openai/`, so the Tokens model ID (which has its own `provider/` part) is sent as-is. ### LiteLLM Python SDK ```bash pip install --upgrade litellm ``` ```python title="litellm_example.py" import os import litellm response = litellm.completion( model="openai/deepseek/deepseek-v4.1-flash", # "openai/" + Tokens model ID api_base="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], messages=[{"role": "user", "content": "One tip for faster pytest runs?"}], max_tokens=200, ) print(response.choices[0].message.content) ``` Keep `/v1` on `api_base`. Leaving it off is the most common cause of 404s with LiteLLM. ### LiteLLM proxy config If your team runs a LiteLLM proxy, add Tokens models to `model_list` and keep the key in the environment: ```yaml title="config.yaml" model_list: - model_name: deepseek-v4.1-flash # the name your apps will request litellm_params: model: openai/deepseek/deepseek-v4.1-flash api_base: https://tokens.bd/v1 api_key: os.environ/TOKENS_API_KEY ``` ```bash litellm --config config.yaml ``` Clients of the proxy then ask for `deepseek-v4.1-flash`, and LiteLLM forwards those requests to Tokens. A few things to know: - Spend, plan windows and rate limits are enforced per Tokens account, not per proxy user. Everyone behind one Tokens key shares that account's per-minute and concurrency limits ([Rate limits](/docs/rate-limits)). - If you want separate ceilings for different apps, create separate Tokens keys, each with a monthly spend cap and an allowed-models list ([API keys](/docs/api-keys)), and use one key per `model_list` entry. - LiteLLM's own cost tracking won't know Tokens prices. Use the usage page in the dashboard or `GET /v1/tokens/usage` as the source of truth. ## When something fails LangChain and LiteLLM both wrap the OpenAI SDK's errors, so the HTTP status and the Tokens `error.code` (`invalid_api_key`, `model_not_found`, `insufficient_credits`, `window_exhausted`, and so on) end up in the exception message. The [Troubleshooting](/docs/troubleshooting) guide maps each code to a fix. When you open a ticket, include the `x-tokens-request-id` response header. To get at it, reproduce the call once with [cURL](/docs/curl) and `-i`. --- # OpenAI Agents SDK > Run agents built with the OpenAI Agents SDK (Python and TypeScript) on Tokens: Chat Completions model class, the slash-in-model-id problem, tracing, tools and streaming. Section: SDKs & Libraries. Page: https://tokens.bd/docs/openai-agents-sdk The OpenAI Agents SDK is OpenAI's library for building agents: an `Agent` with instructions and tools, run by a `Runner`. It talks to models through the OpenAI API, so it can reach Tokens at `https://tokens.bd/v1`. Two defaults get in the way, and both are easy to fix: the SDK calls the Responses API unless you tell it otherwise, and it reads a model id that contains a slash as `provider/model`. Every Tokens model id contains a slash, so this page shows how to set both up. :::note[Checked against the documentation] Based on the OpenAI Agents SDK documentation, checked October 2026 against the Python package `openai-agents` 0.23.1 (released 2 October 2026) and the TypeScript package `@openai/agents` 0.20.0. The code was checked against the documentation and the SDK's own examples, not run end to end against Tokens. The SDK changes quickly, so compare with the [official docs](https://openai.github.io/openai-agents-python/models/) if something here no longer matches your version. ::: ## What you need - A Tokens key from [API keys](/docs/api-keys), exported as `TOKENS_API_KEY`. - A model id from [/models](/models), for example `deepseek/deepseek-v4.1-flash`. Agents call tools, so pick a model that supports tool calling ([Choosing a model](/docs/choosing-a-model)). - Python 3.10 or newer for the Python package. The TypeScript package needs the `openai` package version 7.2 or newer when you pass your own client. ```bash export TOKENS_API_KEY="tok_live_your_key" ``` ## Why the defaults need changing | Default in the SDK | What happens with Tokens | Fix | | ----------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- | | Uses the Responses API | Tokens passes `/v1/responses` through, but it only works if the provider behind the model implements it ([Responses](/docs/responses)). Chat Completions works for every model. | Use `OpenAIChatCompletionsModel`, or `set_default_openai_api("chat_completions")`. | | Reads `prefix/name` as a provider prefix | The SDK documents that unknown prefixes raise `UserError`, and that `openai/...` is shortened to the part after the slash. A Tokens id such as `deepseek/deepseek-v4.1-flash` hits this rule. | Pass a model object, use a custom `ModelProvider`, or switch the prefix modes. | | Uploads traces to OpenAI | Traces go to OpenAI's servers and need an OpenAI key. With a Tokens key the upload fails with a 401 in your logs. | Turn tracing off. | The SDK's documentation lists these under its "non-OpenAI models" section. It recommends Chat Completions when a provider lacks Responses support, and says unknown prefixes raise `UserError` instead of being passed through. ## Python: use a Chat Completions model object This is the setup to start with. You create an `AsyncOpenAI` client that points at Tokens, wrap it in `OpenAIChatCompletionsModel`, and give that object to the agent. Because the agent receives a model object and not a model name string, the SDK does not parse the model id, so the slash is not an issue. ```bash pip install --upgrade openai-agents ``` ```python title="agent.py" import asyncio import os from openai import AsyncOpenAI from agents import Agent, OpenAIChatCompletionsModel, Runner, set_tracing_disabled set_tracing_disabled(True) # traces would go to OpenAI; see the Tracing section client = AsyncOpenAI( base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], ) agent = Agent( name="Reviewer", instructions="You are a concise senior engineer. Answer in at most three sentences.", model=OpenAIChatCompletionsModel(model="deepseek/deepseek-v4.1-flash", openai_client=client), ) async def main() -> None: result = await Runner.run(agent, "When should I use a dataclass instead of a dict?") print(result.final_output) asyncio.run(main()) ``` `base_url` keeps the `/v1` at the end. The model string goes to Tokens exactly as written, so it must match an id from [/models](/models) or `GET https://tokens.bd/v1/models`. ## Python: set a model for every agent in one place If you have many agents, a custom `ModelProvider` passed through `RunConfig` applies one model setup to a whole run, and your `Agent` objects can keep plain model names. This follows the SDK's own `custom_example_provider.py` example. ```python title="provider.py" import asyncio import os from openai import AsyncOpenAI from agents import ( Agent, Model, ModelProvider, OpenAIChatCompletionsModel, RunConfig, Runner, set_tracing_disabled, ) set_tracing_disabled(True) client = AsyncOpenAI( base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], ) class TokensModelProvider(ModelProvider): def get_model(self, model_name: str | None) -> Model: # The name arrives untouched, slash included. return OpenAIChatCompletionsModel( model=model_name or "deepseek/deepseek-v4.1-flash", openai_client=client, ) agent = Agent( name="Assistant", instructions="You are a helpful assistant.", model="deepseek/deepseek-v4.1-flash", ) async def main() -> None: result = await Runner.run( agent, "Name two uses for asyncio.Semaphore.", run_config=RunConfig(model_provider=TokensModelProvider()), ) print(result.final_output) asyncio.run(main()) ``` `RunConfig` applies to that one `Runner.run` call. A run without it goes to OpenAI's own endpoint with whatever `OPENAI_API_KEY` is set. ### Alternative: the global client The SDK also has a global default, which is the shortest setup when every agent should use Tokens: ```python import os from openai import AsyncOpenAI from agents import set_default_openai_api, set_default_openai_client, set_tracing_disabled set_default_openai_client( AsyncOpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]), use_for_tracing=False, ) set_default_openai_api("chat_completions") set_tracing_disabled(True) ``` With this setup, a plain string such as `model="deepseek/deepseek-v4.1-flash"` still goes through the SDK's default `MultiProvider`, which splits on the first slash. The part before the slash is not a provider the SDK knows, and the default behavior is to raise `UserError: Unknown prefix`. If the prefix is `openai`, the SDK silently sends only the part after the slash. To use model strings with the global setup, build the provider yourself and tell it to keep the full id: ```python import os from agents import MultiProvider, RunConfig provider = MultiProvider( openai_base_url="https://tokens.bd/v1", openai_api_key=os.environ["TOKENS_API_KEY"], openai_use_responses=False, # Chat Completions openai_prefix_mode="model_id", # keep a leading "openai/" as part of the id unknown_prefix_mode="model_id", # keep any other "provider/" as part of the id ) run_config = RunConfig(model_provider=provider) ``` `openai_prefix_mode` and `unknown_prefix_mode` are documented by the SDK for sending namespaced ids such as `openrouter/openai/gpt-4.1-mini` to an OpenAI-compatible backend. The simpler routes (a model object, or the `ModelProvider` above) avoid the question, so prefer them. :::note `use_for_tracing=False` stops the SDK from using this client's key (your Tokens key) to upload traces to OpenAI. The default is `True`. The argument is described in the SDK's source docstring, not in its guide pages, so check it against the version you install. ::: ## Tracing: turn it off The SDK uploads traces to OpenAI's servers by default. Without an OpenAI platform key you get 401 errors in your logs, even though the agent itself works. The documentation gives three ways to turn tracing off: | Scope | How | | ------------ | ------------------------------------------------------------------------ | | Whole process | `set_tracing_disabled(True)`, or the environment variable `OPENAI_AGENTS_DISABLE_TRACING=1` | | One run | `RunConfig(tracing_disabled=True)` | You can also keep tracing and set a separate OpenAI key only for uploads with `set_tracing_export_api_key(...)`. That key has to come from platform.openai.com. It is not your Tokens key, and your prompts then go to OpenAI as trace data. If your prompts are private, leave tracing off. ## Tools Give the agent Python functions with type hints and a docstring. The SDK builds the tool schema from them. ```python title="tools.py" import asyncio import os from openai import AsyncOpenAI from agents import ( Agent, OpenAIChatCompletionsModel, Runner, function_tool, set_tracing_disabled, ) set_tracing_disabled(True) client = AsyncOpenAI( base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], ) @function_tool def get_weather(city: str) -> str: """Current weather for a city. Args: city: The city to look up. """ return f"{city}: light rain, 29C" # replace with a real lookup agent = Agent( name="Weather helper", instructions="Use the tool when the user asks about the weather.", model=OpenAIChatCompletionsModel(model="deepseek/deepseek-v4.1-flash", openai_client=client), tools=[get_weather], ) async def main() -> None: result = await Runner.run(agent, "Do I need an umbrella in Dhaka?") print(result.final_output) asyncio.run(main()) ``` `function_tool` is the decorator in the released package. Newer SDK documentation imports the same decorator as `from agents.decorators import tool`, which is an alias for it. Either works on the version checked. Each tool call is another request to Tokens, so a run with several tool rounds costs several requests and counts against your [rate limits](/docs/rate-limits). Tool calling needs a model that supports it ([Tool calling](/docs/tool-calling)). The SDK's documentation lists tools that only work on the Responses API (for example `ToolSearchTool`), and says they are rejected on Chat Completions backends. Hosted tools such as web search or file search run on OpenAI's side and are not available through Tokens. ## Stream the output `Runner.run_streamed` returns a result you iterate for events. This is the SDK's documented text-streaming loop: ```python title="stream.py" import asyncio import os from openai import AsyncOpenAI from openai.types.responses import ResponseTextDeltaEvent from agents import Agent, OpenAIChatCompletionsModel, Runner, set_tracing_disabled set_tracing_disabled(True) client = AsyncOpenAI( base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], ) agent = Agent( name="Writer", instructions="You write short, plain answers.", model=OpenAIChatCompletionsModel(model="deepseek/deepseek-v4.1-flash", openai_client=client), ) async def main() -> None: result = Runner.run_streamed(agent, input="Write a haiku about merge conflicts.") async for event in result.stream_events(): if event.type == "raw_response_event" and isinstance(event.data, ResponseTextDeltaEvent): print(event.data.delta, end="", flush=True) print() asyncio.run(main()) ``` Keep reading the stream until it ends. The run is not finished until the iterator is. The SDK's streaming page shows this loop for the default setup. It does not say whether it behaves the same with `OpenAIChatCompletionsModel`, so test it once with your model. See also [Streaming](/docs/streaming). Tokens streams carry token counts only when the request asks for them. On Chat Completions the SDK has a `ModelSettings(include_usage=True)` field for that (the field is documented as available for Chat Completions only). Import `ModelSettings` from `agents` and pass it as `Agent(..., model_settings=ModelSettings(include_usage=True))`. If a provider sends broken tool-call fragments while streaming, the SDK has an option, `openai_buffer_streamed_tool_calls=True` on `MultiProvider`, that buffers them. ## TypeScript The TypeScript package is `@openai/agents`. It needs `zod` 4 as a peer dependency, and the `openai` package (version 7.2 or newer) when you pass your own client. ```bash npm install @openai/agents zod openai ``` ```ts title="agent.ts" import OpenAI from "openai"; import { Agent, OpenAIChatCompletionsModel, run, setTracingDisabled, tool, } from "@openai/agents"; import { z } from "zod"; setTracingDisabled(true); const client = new OpenAI({ apiKey: process.env.TOKENS_API_KEY, baseURL: "https://tokens.bd/v1", }); const getWeather = tool({ name: "get_weather", description: "Current weather for a city", parameters: z.object({ city: z.string() }), async execute({ city }) { return `${city}: light rain, 29C`; // replace with a real lookup }, }); const agent = new Agent({ name: "Weather helper", instructions: "Use the tool when the user asks about the weather.", model: new OpenAIChatCompletionsModel(client, "deepseek/deepseek-v4.1-flash"), tools: [getWeather], }); const result = await run(agent, "Do I need an umbrella in Dhaka?"); console.log(result.finalOutput); ``` Passing an `OpenAIChatCompletionsModel` instance keeps that agent on Chat Completions, and the model id is not parsed. The SDK's own example also offers a global route, `setDefaultOpenAIClient(client)` with `setOpenAIAPI("chat_completions")`, and a `new Runner({ modelProvider })` route with an `OpenAIProvider({ openAIClient: client })`. The TypeScript documentation does not say how those routes treat a model name with a slash, so use the model object above. To stream text in Node: ```ts const stream = await run(agent, "Write a haiku about merge conflicts.", { stream: true }); stream.toTextStream({ compatibleWithNodeStreams: true }).pipe(process.stdout); await stream.completed; ``` You can turn tracing off with `setTracingDisabled(true)` as above, or with `OPENAI_AGENTS_DISABLE_TRACING=1`. Run this code on a server or in a CLI, not in a browser bundle: Tokens does not send CORS headers and the key would be visible to anyone who opens dev tools. ## Check that it works Run `python agent.py`. A short answer printed to the terminal means the key, base URL and model id are right. Then open Usage analytics in the [dashboard](/dashboard) and look for the request. If the call fails, test the endpoint with [cURL](/docs/curl) first: a 200 there with a failure in your agent points at the SDK setup, not the key. ## Choosing a model Agents loop: they call tools, read results and call again, so the model must handle tool calls well and keep a long context. Some models answer plain chat but fail on tools. [Choosing a model](/docs/choosing-a-model) covers which suit agent work, and each model's page in [/models](/models) shows its context window and whether tools are supported. Set `max_tokens` through `ModelSettings` if you want to cap each reply. A run with several tool rounds can use a lot of tokens, so set a spend cap on the key you give an agent ([API keys](/docs/api-keys)). ## Limits and what does not work - **Responses-only features.** Hosted tools (web search, file search, code interpreter), `previous_response_id` and other Responses-only fields are not available on the Chat Completions path. The SDK drops Responses-only fields silently unless you turn on strict validation (`strict_feature_validation=True` on `OpenAIProvider`, `openai_strict_feature_validation=True` on `MultiProvider`). - **Structured output.** The SDK sends `json_schema` response formats. If the model behind a Tokens id does not support them, the upstream rejects the request with a 400 (`invalid_request`). Use a model that supports structured output ([Structured output](/docs/structured-output)). - **Audio and Realtime.** The Chat Completions adapter raises `AgentsException("Audio is not currently supported")` for audio output. Realtime and voice agents need OpenAI's own endpoints and do not go through Tokens. - **Tracing.** Traces go to OpenAI, not Tokens, and the Tokens key cannot upload them. - **Empty replies on `finish_reason="length"`.** The adapter raises `ModelBehaviorError` when the model stops at the length limit with no output. That is a token or reasoning budget problem. Raise `max_tokens` or pick a model with less hidden reasoning. ## Troubleshooting **`UserError: Unknown prefix: `.** The model was given to the agent as a string and the SDK's default provider read the part before the slash as a provider prefix. Pass an `OpenAIChatCompletionsModel` object, use the `ModelProvider` above, or set `unknown_prefix_mode="model_id"` on a `MultiProvider`. **404 `model_not_found`.** Either the id has a typo, or the SDK removed a leading `openai/` from it. Compare the id in your Tokens usage log with the one you meant to send, and check it against `GET https://tokens.bd/v1/models`. **404 or 400 from `/v1/responses`.** The SDK is still on the Responses API. Use `OpenAIChatCompletionsModel` or `set_default_openai_api("chat_completions")`. **401 errors from `api.openai.com` or "incorrect API key" in the log, and the agent still answers.** These are trace uploads. Turn tracing off. **401 `invalid_api_key` from Tokens.** The key in `TOKENS_API_KEY` is wrong or revoked. Create a new one in [/dashboard/keys](/dashboard/keys). **402 `insufficient_credits`, 403 `model_not_allowed_on_key`, 429 `window_exhausted`.** These are account limits, not SDK problems. See [Errors](/docs/errors) and [Troubleshooting](/docs/troubleshooting). An agent that fires many requests in parallel can hit `429 concurrency_limit`: lower the parallelism ([Rate limits](/docs/rate-limits)). **Errors surface as SDK exceptions.** The SDK wraps the OpenAI client, so a Tokens error arrives as an `openai` exception. Read `e.code` and `e.response.headers.get("x-tokens-request-id")` and include the request id when you contact [support](/docs/support). --- # Claude Agent SDK > Run agents built with the Claude Agent SDK (Python and TypeScript) on Tokens: set ANTHROPIC_BASE_URL and a Tokens key, choose a model id, handle the opus, sonnet and haiku aliases, stream and add tools. Section: SDKs & Libraries. Page: https://tokens.bd/docs/claude-agent-sdk The Claude Agent SDK is Anthropic's library for running the Claude Code agent loop inside your own Python or TypeScript program: built-in tools to read and edit files and run commands, permissions, sessions, subagents and MCP. It is not a thin API client. The SDK starts a Claude Code process and talks to it, and that process sends Anthropic Messages requests to whatever `ANTHROPIC_BASE_URL` says. Pointed at `https://tokens.bd`, those requests go to Tokens' [/v1/messages](/docs/messages) endpoint. If you only need to call a model, the [Anthropic SDK](/docs/anthropic-sdk) is the smaller choice. :::note[Checked against the documentation] Based on the Claude Agent SDK documentation at code.claude.com, checked October 2026 against `claude-agent-sdk` 0.2.165 for Python and `@anthropic-ai/claude-agent-sdk` 0.3.296 for TypeScript. The code was checked against the documentation, not run end to end against Tokens. ::: :::warning[Non-Claude models are best effort] Anthropic's gateway documentation says it "doesn't support routing Claude Code to non-Claude models through any gateway". The Agent SDK runs the same agent as Claude Code, so the same applies. It works when the model handles tool calls well, and that is per model. If an agent loses its way in long runs, switch models before debugging the SDK. [Claude Code](/docs/claude-code) has the same caveat. ::: ## What you need - A Tokens key from [API keys](/docs/api-keys), exported as `TOKENS_API_KEY`. - A model id from [/models](/models), for example `deepseek/deepseek-v4.1-flash`, that supports tool calling ([Choosing a model](/docs/choosing-a-model)). - Python 3.10 or newer, or Node.js 18 or newer. Both packages bundle a native Claude Code binary, so most installs need no separate Claude Code install. Some do not: a pip source install (for example on Windows on ARM), or an npm install that skips optional dependencies (`npm ci --omit=optional`). In those cases install Claude Code natively ([Claude Code](/docs/claude-code)), and in TypeScript set `pathToClaudeCodeExecutable`. ```bash export TOKENS_API_KEY="tok_live_your_key" ``` ## How the SDK reaches Tokens The SDK has no gateway options of its own. Anthropic's documentation says it passes environment variables to the Claude Code process it starts, and each SDK has an `env` option for that. You set three things: | Setting | Value | | ------------------------------------------------ | ----------------------------------------------------------------------------------------------------- | | `ANTHROPIC_BASE_URL` | `https://tokens.bd`, without `/v1`. The process adds `/v1/messages` itself. | | `ANTHROPIC_AUTH_TOKEN` | Your Tokens key. Sent as `Authorization: Bearer`, which Tokens accepts. | | The model (`model` option and the alias variables) | A Tokens model id such as `deepseek/deepseek-v4.1-flash`, not `sonnet` or `opus`. See the next section. | `ANTHROPIC_API_KEY` also works (it is sent as `x-api-key`). Set only one of the two. This page uses `ANTHROPIC_AUTH_TOKEN`, the same as [Claude Code](/docs/claude-code). The two SDKs treat `env` differently, and the difference matters: - **TypeScript.** The process inherits your environment by default, but setting `options.env` replaces it entirely. Spread `process.env` into it, or the process loses `PATH` and everything else. - **Python.** `ClaudeAgentOptions(env=...)` is merged on top of the inherited environment. The SDK does not read `.env` files. Load them yourself before you start the SDK, for example with `dotenv`. ## Choose the model id and fix the aliases The `model` option takes "a Claude model alias or full model name" (`ClaudeAgentOptions.model` in Python, `options.model` in TypeScript). Pass a Tokens model id and the process sends it to Tokens as written. Aliases are the catch. Anthropic's model documentation says `sonnet`, `opus` and `haiku` resolve to the latest Claude model for your provider, and that `ANTHROPIC_BASE_URL` "changes where requests are sent, not which model answers them". So with Tokens, an alias still resolves to a Claude model name unless you redirect it, and that is not necessarily a model Tokens serves. Parts of the agent use aliases without you asking: | What | Which model it uses | | --------------------------------------------------------------------------------- | -------------------------------------------------------------------- | | Your main loop | The `model` option | | Background work such as session titles | The `haiku` alias, or `ANTHROPIC_DEFAULT_HAIKU_MODEL` when it is set | | `sonnet`, `opus` and `haiku` when you or a subagent definition name them | `ANTHROPIC_DEFAULT_SONNET_MODEL`, `ANTHROPIC_DEFAULT_OPUS_MODEL`, `ANTHROPIC_DEFAULT_HAIKU_MODEL` | | Subagents with no model of their own | `CLAUDE_CODE_SUBAGENT_MODEL` | Set all of them to the same Tokens id, as the [Claude Code page](/docs/claude-code) does, so no part of the agent asks for a model Tokens does not have. Each variable has to be a full model id. Copy ids from [/models](/models) or `GET https://tokens.bd/v1/models`. Two more variables from the Claude Code page also apply, because the SDK runs the same process: - `CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1` removes most Claude-only beta fields from requests, which prevents `400` errors from non-Claude models. - `CLAUDE_CODE_MAX_CONTEXT_TOKENS` tells Claude Code the real context window. For ids it does not recognize it assumes 200K tokens, so a model with a smaller window fails with "prompt too long" instead of compacting. Take the number from the model's page in [/models](/models). ## Python: a minimal agent ```bash pip install --upgrade claude-agent-sdk ``` ```python title="agent.py" import asyncio import os from claude_agent_sdk import AssistantMessage, ClaudeAgentOptions, ResultMessage, query MODEL = "deepseek/deepseek-v4.1-flash" ENV = { "ANTHROPIC_BASE_URL": "https://tokens.bd", "ANTHROPIC_AUTH_TOKEN": os.environ["TOKENS_API_KEY"], "ANTHROPIC_DEFAULT_OPUS_MODEL": MODEL, "ANTHROPIC_DEFAULT_SONNET_MODEL": MODEL, "ANTHROPIC_DEFAULT_HAIKU_MODEL": MODEL, "CLAUDE_CODE_SUBAGENT_MODEL": MODEL, "CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS": "1", } options = ClaudeAgentOptions( model=MODEL, allowed_tools=["Read", "Glob", "Grep"], # read-only tools, approved automatically max_turns=8, setting_sources=[], # ignore ~/.claude/settings.json, see the note below env=ENV, ) async def main() -> None: async for message in query( prompt="List the Python files in this folder and say in one sentence what each one does.", options=options, ): if isinstance(message, AssistantMessage): for block in message.content: if hasattr(block, "text"): print(block.text) elif hasattr(block, "name"): print(f"[tool: {block.name}]") elif isinstance(message, ResultMessage): print(f"Done: {message.subtype}, {message.num_turns} turns") asyncio.run(main()) ``` `query()` returns an async iterator. Each item is a message: the model's text, a tool call, a tool result, then a final `ResultMessage`. `allowed_tools` only approves tools without a prompt. It does not remove the others. To remove a tool, use `disallowed_tools`. :::note[Settings files can override your variables] Claude Code reads `~/.claude/settings.json` and project settings. Anthropic's documentation says that when a shell export and a settings-file `env` block set the same variable, the settings-file value wins. If you already set up [Claude Code](/docs/claude-code) on the same machine, that file's base URL, key or model would apply to your agent. `setting_sources=[]` (Python) skips those files. It also means the agent does not load `CLAUDE.md`, skills or project settings. Leave the line out if you want them. ::: ## TypeScript: a minimal agent ```bash npm install @anthropic-ai/claude-agent-sdk npm install --save-dev tsx ``` Set `"type": "module"` in `package.json` so top-level `await` works, or name the file `agent.mts`. ```ts title="agent.ts" import { query } from "@anthropic-ai/claude-agent-sdk"; const model = "deepseek/deepseek-v4.1-flash"; const env = { ...process.env, // required: setting env replaces the whole environment ANTHROPIC_BASE_URL: "https://tokens.bd", ANTHROPIC_AUTH_TOKEN: process.env.TOKENS_API_KEY, ANTHROPIC_DEFAULT_OPUS_MODEL: model, ANTHROPIC_DEFAULT_SONNET_MODEL: model, ANTHROPIC_DEFAULT_HAIKU_MODEL: model, CLAUDE_CODE_SUBAGENT_MODEL: model, CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS: "1", }; for await (const message of query({ prompt: "List the TypeScript files in this folder and say in one sentence what each one does.", options: { model, allowedTools: ["Read", "Glob", "Grep"], // read-only tools, approved automatically maxTurns: 8, settingSources: [], // ignore ~/.claude/settings.json env, }, })) { if (message.type === "assistant" && message.message?.content) { for (const block of message.message.content) { if ("text" in block) console.log(block.text); else if ("name" in block) console.log(`[tool: ${block.name}]`); } } else if (message.type === "result") { console.log(`Done: ${message.subtype}`); } } ``` ```bash npx tsx agent.ts ``` If you would rather not pass `env` in code, export the same variables in the shell that runs your program. In TypeScript the process inherits them. In Python they pass through as well, since `env` is merged on top of the environment. ## Stream the output By default the SDK yields a whole text block or tool call after the model finishes it. For token-by-token text, turn on partial messages. The SDK then also yields raw Messages API stream events, and you read the text deltas. :::code-tabs ```python title="Python" from claude_agent_sdk import ClaudeAgentOptions, query from claude_agent_sdk.types import StreamEvent # MODEL and ENV are defined in the minimal program above options = ClaudeAgentOptions( model=MODEL, include_partial_messages=True, setting_sources=[], env=ENV, ) async def stream() -> None: async for message in query( prompt="Explain optimistic locking in two sentences.", options=options ): if isinstance(message, StreamEvent): event = message.event if event.get("type") == "content_block_delta": delta = event.get("delta", {}) if delta.get("type") == "text_delta": print(delta.get("text", ""), end="", flush=True) ``` ```ts title="TypeScript" // model and env are defined in the minimal program above for await (const message of query({ prompt: "Explain optimistic locking in two sentences.", options: { model, includePartialMessages: true, settingSources: [], env }, })) { if (message.type === "stream_event") { const event = message.event; if (event.type === "content_block_delta" && event.delta.type === "text_delta") { process.stdout.write(event.delta.text); } } } ``` ::: The events have the Anthropic streaming shape, which Tokens produces for any model ([Messages](/docs/messages), [Streaming](/docs/streaming)). Stream events cover the main agent only. Token deltas from subagents are not forwarded. ## Add your own tools Besides the built-in tools, you can give the agent functions from your own code as an in-process MCP server. In Python, `@tool` and `create_sdk_mcp_server` build it. In TypeScript, `tool` and `createSdkMcpServer` do, with a Zod schema. :::code-tabs ```python title="Python" from typing import Any from claude_agent_sdk import ClaudeAgentOptions, create_sdk_mcp_server, tool @tool("get_weather", "Current weather for a city", {"city": str}) async def get_weather(args: dict[str, Any]) -> dict[str, Any]: return {"content": [{"type": "text", "text": f"{args['city']}: light rain, 29C"}]} weather = create_sdk_mcp_server(name="weather", version="1.0.0", tools=[get_weather]) # MODEL and ENV are defined in the minimal program above options = ClaudeAgentOptions( model=MODEL, mcp_servers={"weather": weather}, allowed_tools=["mcp__weather__get_weather"], # MCP tools are named mcp____ max_turns=5, setting_sources=[], env=ENV, ) ``` ```ts title="TypeScript" import { createSdkMcpServer, query, tool } from "@anthropic-ai/claude-agent-sdk"; import { z } from "zod"; const getWeather = tool( "get_weather", "Current weather for a city", { city: z.string() }, async ({ city }) => ({ content: [{ type: "text", text: `${city}: light rain, 29C` }], }), ); const weather = createSdkMcpServer({ name: "weather", version: "1.0.0", tools: [getWeather] }); // model and env are defined in the minimal program above const options = { model, mcpServers: { weather }, allowedTools: ["mcp__weather__get_weather"], maxTurns: 5, settingSources: [], env, }; ``` ::: Pass `options` to `query()` as in the minimal examples. The TypeScript package needs `zod` 4 as a peer dependency. Each tool round is another `/v1/messages` request, so an agent run costs several requests and counts against your [rate limits](/docs/rate-limits). See [Tool calling](/docs/tool-calling) for how tool calls are handled for non-Claude models. ## Check that it works First test the endpoint and the model id without the SDK (macOS, Linux, WSL or Git Bash): ```bash curl -sS -w '\n%{http_code}\n' -X POST "https://tokens.bd/v1/messages" \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "anthropic-version: 2023-06-01" -H "content-type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 1, "messages": [{"role": "user", "content": "."}]}' ``` A `200` means the key and model are fine. Then run the minimal program. You should see the tool names it calls, the model's answer and `Done: success`. The requests also show up in Usage analytics in the [dashboard](/dashboard), under the model id you set. If the log lists a different model, one of the alias variables was not set. `ResultMessage` has a `total_cost_usd` field. That is the SDK's own estimate, not what Tokens bills. Tokens' [usage page and `GET /v1/tokens/usage`](/docs/models-and-usage) are the source of truth. ## Choosing a model The agent loop relies on tool calls, long contexts and following a system prompt. Models differ a lot here. [Choosing a model](/docs/choosing-a-model) covers which suit agent work. Use `fallback_model` (Python) or `fallbackModel` (TypeScript) only with another Tokens id, never a Claude name. Lowering `max_turns` and setting a monthly spend cap on the key ([API keys](/docs/api-keys)) limits the damage from a runaway loop. ## Limits and what does not work - **Claude-only features.** Fast mode checks Anthropic's API directly and does not work through Tokens. Remote Control and voice dictation are disabled while a gateway credential is set. Extended thinking and prompt caching work only where the model behind the Tokens id supports them ([Reasoning](/docs/reasoning), [Prompt caching](/docs/prompt-caching)). - **Claude Code's own model picker.** The `/model` picker belongs to the interactive CLI. In the SDK you set the model in options. - **Server-side Anthropic tools.** Tools that run on Anthropic's side, such as hosted web search, are not documented for gateways, and Anthropic does not document them for non-Claude models. Test one before you rely on it, and prefer your own tools or MCP servers. - **Login.** Anthropic does not allow third-party developers to offer claude.ai login or its rate limits in products built on the SDK, so use an API key as shown here. - **Cloud provider modes.** Variables such as `CLAUDE_CODE_USE_BEDROCK` or `CLAUDE_CODE_USE_VERTEX` select other providers. Do not set them with Tokens. - **Background traffic.** Claude Code sends version checks and telemetry to Anthropic outside the gateway. `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1` turns that off, at the cost of auto-updates. ## Troubleshooting **404 on every request.** The base URL ends in `/v1`. The process appends `/v1/messages` itself, so use `https://tokens.bd`. **401 `invalid_api_key`, or `Not logged in`.** The key is wrong, or the variables never reached the process. In TypeScript, check that `env` has `...process.env`. In Python, check the `env` dict. Also check that a `~/.claude/settings.json` is not overriding your values (use `settingSources: []`). The SDK does not load `.env` files. **404 `model_not_found`.** The id is not one Tokens knows, or an alias (`sonnet`, `opus`, `haiku`) resolved to a Claude name. Set the alias variables to your Tokens id, and check the id against `GET https://tokens.bd/v1/models`. **400 errors about `thinking`, `effort`, `context_management` or unknown fields.** Claude Code sends adaptive thinking and beta fields that non-Claude models reject. `CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1` removes most of them. If one model still fails, try another. **"Prompt is too long" or the session never compacts.** Claude Code assumes a 200K context for unknown ids. Set `CLAUDE_CODE_MAX_CONTEXT_TOKENS` to the model's real window. **The process fails to start.** No bundled binary was installed. Install Claude Code natively, and in TypeScript set `pathToClaudeCodeExecutable`. **402 `insufficient_credits`, 403 `model_not_allowed_on_key` or `tier_permission_denied`, 429 `window_exhausted` or `concurrency_limit`.** These are account limits, not SDK problems. A run with several subagents sends requests in parallel, which can hit `concurrency_limit`. See [Errors](/docs/errors) and [Troubleshooting](/docs/troubleshooting), and include the `x-tokens-request-id` header value when you contact [support](/docs/support). --- # LlamaIndex > Use Tokens from LlamaIndex (Python) with the OpenAILike class: api_base, is_chat_model, context window, streaming, function-calling agents, and embeddings for RAG. Section: SDKs & Libraries. Page: https://tokens.bd/docs/llamaindex LlamaIndex is a Python framework for building RAG pipelines and agents over your own data. It talks to models through LLM classes, and `OpenAILike` is the one meant for any OpenAI-compatible server. Pointed at `https://tokens.bd/v1`, it sends Chat Completions requests to Tokens. A RAG pipeline also needs an embedding model, and LlamaIndex picks an OpenAI one unless you say otherwise, so that part gets its own section below. :::note[Checked against the documentation] Based on the LlamaIndex documentation and the integration packages' source and README files, checked October 2026 against `llama-index-llms-openai-like` 0.8.1 (released 1 October 2026, needs `llama-index-core` 0.14.3 or newer and Python 3.10 or newer). The code was checked against the documentation, not run end to end against Tokens. ::: ## What you need - A Tokens key from [API keys](/docs/api-keys), exported as `TOKENS_API_KEY`. - A model id from [/models](/models), for example `deepseek/deepseek-v4.1-flash`, and its context window from the model's page. - For RAG, an embedding model id from the catalog (see [Embeddings](#embeddings-for-rag)). ```bash pip install --upgrade llama-index-llms-openai-like export TOKENS_API_KEY="tok_live_your_key" ``` ## Set up OpenAILike ```python title="chat.py" import os from llama_index.core.llms import ChatMessage from llama_index.llms.openai_like import OpenAILike llm = OpenAILike( model="deepseek/deepseek-v4.1-flash", api_base="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], is_chat_model=True, context_window=128000, # set this to the window shown on the model's page in /models ) response = llm.chat( [ ChatMessage(role="system", content="You are a precise SQL reviewer."), ChatMessage(role="user", content="Is SELECT * in a view a problem? Two sentences."), ] ) print(response) ``` The arguments, from the integration's README: | Argument | What to set | | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | | `model` | The Tokens model id, exactly as listed in [/models](/models) or by `GET https://tokens.bd/v1/models`. LlamaIndex sends it unchanged. | | `api_base` | `https://tokens.bd/v1`. Keep the `/v1`. | | `api_key` | Your Tokens key. | | `is_chat_model` | `True`. See below. | | `context_window` | The model's real context window in tokens. LlamaIndex uses it to size prompts and chunks. | | `is_function_calling_model` | `True` if the model supports tool calling and you use agents or tools. The default is `False`. | Two defaults catch people out: - **`is_chat_model` defaults to `False`.** In that mode `OpenAILike` sends your calls to the completions endpoint, `/v1/completions`. Many chat models do not serve it and the call fails ([Legacy completions](/docs/legacy-completions)). Set `is_chat_model=True` so requests go to `/v1/chat/completions`. - **`context_window` has a small default** (about 3,900 tokens, according to the class docstring). Leave it and LlamaIndex splits and truncates context far more than needed. Take the real value from the model's page. `OpenAILike` inherits its timeout (60 seconds) and retries (3) from LlamaIndex's `OpenAI` class. Both can be set in the constructor as `timeout` and `max_retries`. For long generations raise `timeout`, or stream. Tokens already fails over between upstream sources before it answers ([Rate limits](/docs/rate-limits)), so a few retries are enough. ### Make it the default for a pipeline ```python from llama_index.core import Settings Settings.llm = llm ``` Everything that uses the default LLM (query engines, chat engines) now goes to Tokens. You can also pass the model to one engine: `index.as_query_engine(llm=llm)`. ## Stream responses ```python messages = [ChatMessage(role="user", content="Give me three naming tips for Python modules.")] for chunk in llm.stream_chat(messages): print(chunk.delta, end="", flush=True) print() ``` `delta` is the new text in each chunk. Tokens' streamed responses carry token counts only when the request asks for them ([Streaming](/docs/streaming)), and the source does not show `OpenAILike` asking for them, so streamed calls may show no usage in LlamaIndex. The dashboard still records the request. ## Tool calling and agents Set `is_function_calling_model=True` and give LlamaIndex's `FunctionAgent` plain Python functions. Their names, type hints and docstrings become the tool schema. ```python title="agent.py" import asyncio import os from llama_index.core.agent.workflow import AgentStream, FunctionAgent from llama_index.llms.openai_like import OpenAILike llm = OpenAILike( model="deepseek/deepseek-v4.1-flash", api_base="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], is_chat_model=True, is_function_calling_model=True, context_window=128000, # set this to the window shown on the model's page in /models ) def get_weather(city: str) -> str: """Current weather for a city.""" return f"{city}: light rain, 29C" # replace with a real lookup agent = FunctionAgent( tools=[get_weather], llm=llm, system_prompt="You answer questions about the weather. Use the tool when you need data.", ) async def main() -> None: response = await agent.run(user_msg="Do I need an umbrella in Dhaka?") print(response) # The same agent, with streamed text: handler = agent.run(user_msg="And in Chattogram?") async for event in handler.stream_events(): if isinstance(event, AgentStream): print(event.delta, end="", flush=True) await handler print() asyncio.run(main()) ``` `agent.run(...)` returns a handler. `await` it for the final answer, or iterate `handler.stream_events()` first and then `await handler`, as in the LlamaIndex streaming guide. `AgentStream.delta` is the newest piece of text. If your model cannot stream a reply that includes tool calls, create the agent with `streaming=False`, which the LlamaIndex documentation gives for that case. The agent calls the model once per step, so one question that needs two tool calls is at least three requests to Tokens, and they count against your [rate limits](/docs/rate-limits). Tool calling needs a model that supports it ([Tool calling](/docs/tool-calling)). `should_use_structured_outputs=True` on `OpenAILike` switches structured output on through `response_format`. Only set it for a model that supports JSON schema output ([Structured output](/docs/structured-output)). ## Embeddings for RAG If you set `Settings.llm` and nothing else, `VectorStoreIndex` still embeds with LlamaIndex's default, OpenAI's `text-embedding-ada-002` through `OpenAIEmbedding`. That call needs an OpenAI key and goes to OpenAI, not Tokens. You have two choices. **Embed through Tokens.** `/v1/embeddings` works only with embedding models. A chat model such as `deepseek/deepseek-v4.1-flash` is not one, and the call fails. Find an embedding model in the catalog (search for `embed` in [/models](/models)) and read its id from an environment variable, as the [Embeddings](/docs/embeddings) page does, because the catalog changes. If the catalog has no embedding model, none is available to your account yet, so use the second choice. ```bash pip install --upgrade llama-index-embeddings-openai-like export EMBEDDING_MODEL="the-id-from-the-catalog" ``` ```python import os from llama_index.core import Settings from llama_index.embeddings.openai_like import OpenAILikeEmbedding Settings.embed_model = OpenAILikeEmbedding( model_name=os.environ["EMBEDDING_MODEL"], api_base="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], embed_batch_size=10, ) ``` The class takes `model_name`, not `model`, and passing `model` raises a `ValueError`. `embed_batch_size` is how many texts go in one request. Larger batches mean fewer requests, within the provider's limit for that model. Not documented by LlamaIndex: which `encoding_format` it requests. If calls fail or return odd vectors, read the base64 note on the [Embeddings](/docs/embeddings) page, since the OpenAI client library defaults to base64 and not every provider supports it. Do not mix vectors from different models in one index, and re-embed if you change models. **Embed somewhere else.** Set another embedding model, for example a local one, in `Settings.embed_model` (the LlamaIndex documentation lists the options) and use Tokens only for the LLM. The query step then still sends your retrieved text to Tokens as part of the prompt, like any other chat request. ## Check that it works Run `python chat.py`. A short answer printed to the terminal means the key, base URL and model id are right. Then open Usage analytics in the [dashboard](/dashboard): the request should be listed under the model id you set. To list the ids your key can use, run `curl -H "Authorization: Bearer $TOKENS_API_KEY" https://tokens.bd/v1/models` ([cURL](/docs/curl)). ## Choosing a model For plain question answering over retrieved text, most chat models work. Agents and structured extraction need reliable tool calls or JSON output, which differs by model. [Choosing a model](/docs/choosing-a-model) covers which suit which job, and each model's page in [/models](/models) lists its context window. A bigger `context_window` lets LlamaIndex put more retrieved chunks in a prompt, which also costs more input tokens per question. ## Limits and what does not work - **Model names.** `OpenAILike` sends whatever `model` you give it, so the id must be one Tokens serves. Nothing in LlamaIndex checks it against the catalog. - **Embeddings from chat models.** Not possible, see above. - **OpenAI-only integrations.** Features of LlamaIndex's own `OpenAI` class, such as the Responses API classes, assume OpenAI. Use `OpenAILike` for Tokens. - **Token counting.** Tokens bills on the usage the provider reports, not on LlamaIndex's own counts ([Token counting](/docs/token-counting)). - **Hosted tools.** Tools run by OpenAI itself (hosted web search, file search) do not exist on Tokens. Use your own Python tools. - **Other LlamaIndex LLM classes.** This page covers only `OpenAILike`. LlamaIndex's Anthropic class is not documented for custom gateways, so it is not covered. ## Troubleshooting **A 400 or 404 on `/v1/completions`.** `is_chat_model` is still `False`, so the call goes to the completions endpoint, which many chat models do not serve. Set `is_chat_model=True`. **404 on every request.** `api_base` is missing the `/v1`, or has it twice. Use `https://tokens.bd/v1` exactly. **404 `model_not_found`.** The id has a typo, or it is an embedding model used for chat (or the reverse). Check it against `GET https://tokens.bd/v1/models`. **An OpenAI error about a missing or wrong key during indexing.** Embeddings are going to OpenAI, since `Settings.embed_model` is still the default. Set it as shown above. **The index fails with 400 `invalid_request` or 404 `model_not_found` on `/v1/embeddings`.** The id you gave `OpenAILikeEmbedding` is not an embedding model. See [Embeddings](/docs/embeddings). **The agent answers without calling the tool, or errors about tool calling.** Check `is_function_calling_model=True` and that the model supports tools ([Tool calling](/docs/tool-calling)). Try another model. **Context errors or very short answers.** `context_window` is the default. Set it to the model's real window. **401 `invalid_api_key`, 402 `insufficient_credits`, 403 `model_not_allowed_on_key`, 429 `window_exhausted` or `concurrency_limit`.** These are account limits, not LlamaIndex problems. Indexing many documents in parallel can reach `concurrency_limit`. Lower the parallelism or the number of workers. See [Errors](/docs/errors) and [Troubleshooting](/docs/troubleshooting), and include the `x-tokens-request-id` response header when you contact [support](/docs/support). --- # CrewAI > Run CrewAI agents and crews on Tokens: the LLM class with a custom base URL and custom_openai, how to write the model id, streaming, tool calls, rate limits and troubleshooting. Section: SDKs & Libraries. Page: https://tokens.bd/docs/crewai CrewAI is a Python framework for agents that work together as a crew. Each agent gets an `LLM`, and CrewAI can send that LLM's requests to any OpenAI-compatible endpoint. Tokens is one: CrewAI sends OpenAI Chat Completions requests to `https://tokens.bd/v1` with your Tokens key. The one detail that trips people up is the model string. Tokens model ids already contain a slash (`provider/model`), and CrewAI also reads the part before the first slash. The next section shows what to type. :::note[What was checked] Based on CrewAI's LLMs documentation (docs.crewai.com, checked October 2026) and the `crewai` 1.15.27 source on PyPI and GitHub (released 9 October 2026). The routing rules below come from `llm.py` in that source, because the docs page does not describe them. The code was checked against the documentation and source, not run end to end against Tokens. ::: ## What you need - A Tokens key from [API keys](/docs/api-keys), exported as `TOKENS_API_KEY`. - A model id from [/models](/models). Pick one that supports tool calling if your agents use tools ([Choosing a model](/docs/choosing-a-model)). - Python 3.10 to 3.13 (CrewAI 1.15.27 declares `>=3.10,<3.14`). ```bash pip install crewai export TOKENS_API_KEY="tok_live_your_key" ``` CrewAI's docs install it with `uv` (`uv tool install crewai` for the command line tool, `uv add crewai` inside a project). `pip` works the same way for a plain script. `crewai` already depends on the `openai` package, so you do not need the `litellm` extra for this setup. ## Set up the LLM ```python title="llm.py" import os from crewai import LLM llm = LLM( model="deepseek/deepseek-v4.1-flash", base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], custom_openai=True, timeout=120, max_retries=2, ) ``` What each part does: | Setting | Why | | --- | --- | | `model` | The Tokens model id, exactly as listed in [/models](/models). | | `base_url` | `https://tokens.bd/v1`, including `/v1`. CrewAI hands it to the OpenAI Python SDK, which appends `/chat/completions`. | | `custom_openai=True` | Forces CrewAI's native OpenAI client for this LLM, whatever the model string looks like. CrewAI's own gateway example uses it. It also needs a custom endpoint, so it raises an error if you forget `base_url`. | | `timeout`, `max_retries` | Seconds to wait for a response and retry count. CrewAI's docs list both. If you leave them out, the OpenAI SDK defaults apply (10 minutes, 2 retries). | ### How to write the model id With `custom_openai=True`, type the Tokens id exactly as it appears in the catalog, with no extra prefix. CrewAI removes a leading `openai/` if there is one and sends the rest unchanged, so `deepseek/deepseek-v4.1-flash` reaches Tokens as `deepseek/deepseek-v4.1-flash`. Two cases to watch: - **Do not drop `custom_openai=True`.** Without it, CrewAI reads the part before the first slash as a provider name. A Tokens id can start with a name that CrewAI treats as one of its own providers (`deepseek/` is one), and the request then goes to that provider's client instead of Tokens. The LiteLLM-style form, `model="openai/"` plus `base_url`, also reaches the OpenAI client, but `custom_openai=True` removes the guesswork. - **Ids that start with `openai/`.** The leading `openai/` is stripped. If a Tokens id itself begins with `openai/`, write it twice: `model="openai/openai/"`. CrewAI's LLMs page says to always include a provider prefix. The Tokens id already is `provider/model`, so that rule is met, and its own custom-endpoint example passes the gateway's id this way. ### Use environment variables instead CrewAI also reads `OPENAI_BASE_URL` and `OPENAI_API_KEY` (its docs list both): ```bash export OPENAI_BASE_URL="https://tokens.bd/v1" export OPENAI_API_KEY="$TOKENS_API_KEY" ``` ```python from crewai import LLM llm = LLM(model="deepseek/deepseek-v4.1-flash", custom_openai=True) ``` Those variables are global. Any other OpenAI-based tool in the same shell will also send its traffic to Tokens. Passing `base_url` and `api_key` in code keeps the change local to CrewAI. ## Check that it works First call the LLM on its own, with no agents: ```python title="check.py" from llm import llm print(llm.call("Reply with the single word: ready")) ``` A normal sentence back means the key, the base URL and the model id are all right. To check the id without spending tokens, list the catalog: ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` ## A minimal crew ```python title="crew.py" from crewai import Agent, Crew, Task from llm import llm reviewer = Agent( role="Python code reviewer", goal="Point out bugs and unclear names in short snippets", backstory="You review pull requests for a small team and keep feedback short.", llm=llm, max_iter=5, ) task = Task( description="Review this function and list the problems:\n\ndef avg(xs): return sum(xs)/len(xs)", expected_output="A bullet list with at most three items.", agent=reviewer, ) crew = Crew(agents=[reviewer], tasks=[task]) result = crew.kickoff() print(result.raw) print(crew.usage_metrics) ``` `result.raw` is the final text. `crew.usage_metrics` reports token counts for the run. Your Tokens bill comes from the dashboard's [usage page](/docs/usage-and-alerts), not from CrewAI's numbers. `max_iter` caps how many reasoning steps an agent may take before it must answer (CrewAI's default is 20). Each step is one request that Tokens bills, so a lower cap is cheaper while you are still testing. ## Stream responses Set `stream=True` on the LLM: ```python streaming_llm = LLM( model="deepseek/deepseek-v4.1-flash", base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], custom_openai=True, stream=True, ) ``` CrewAI emits an `LLMStreamChunkEvent` for each chunk. This listener is the example from CrewAI's LLMs page: ```python from crewai.events import BaseEventListener, LLMStreamChunkEvent class MyCustomListener(BaseEventListener): def setup_listeners(self, crewai_event_bus): @crewai_event_bus.on(LLMStreamChunkEvent) def on_llm_stream_chunk(self, event: LLMStreamChunkEvent): print(f"Received chunk: {event.chunk}") my_listener = MyCustomListener() ``` Create the listener before you run the crew. When streaming is on, CrewAI asks Tokens for `stream_options.include_usage`, so the usage numbers still arrive. See [Streaming](/docs/streaming) for the wire format. ## Tool calls Give an agent tools with `@tool` from `crewai.tools` and the `tools` argument: ```python from crewai import Agent from crewai.tools import tool from llm import llm @tool("Get weather") def get_weather(city: str) -> str: """Current weather for a city. Use it when the user asks about the weather.""" return f"{city}: light rain, 29C" agent = Agent( role="Travel helper", goal="Answer weather questions for travellers", backstory="You know Bangladesh well.", llm=llm, tools=[get_weather], ) ``` There is no function-calling switch to set for a custom endpoint. In CrewAI's source, the OpenAI provider's `supports_function_calling()` returns true unless the model is an o1-style model, so CrewAI sends tools in the OpenAI format. Whether the model then calls them well depends on the model: check its page in [/models](/models) and read [Tool calling](/docs/tool-calling). If a model ignores tools, try another one before changing code. ## Limits and what to watch - **Requests per minute.** Crews with several agents can fire many requests in a short time. Set `max_rpm` on the `Agent` (CrewAI's docs describe it as the cap on requests per minute) to stay under your plan's limit. Going over returns `429 rate_limited`. See [Rate limits](/docs/rate-limits). - **Parallel tasks.** Tasks that run at the same time count against your account's concurrent request limit and can return `429 concurrency_limit`. Run them in sequence or lower the parallelism. - **Long generations.** The gateway waits up to 600 seconds for a response. Keep CrewAI's `timeout` at or below that, and use `stream=True` for very long outputs. - **Not covered here.** CrewAI features that call an embeddings model on their own (memory, knowledge sources) are not set up by this page. Tokens serves embeddings ([Embeddings](/docs/embeddings)), but you have to point those features at the same endpoint yourself. ## Troubleshooting | Symptom | Cause and fix | | --- | --- | | `ImportError: Unable to initialize LLM ... LiteLLM fallback package is not installed` | CrewAI did not choose its OpenAI client and fell back to LiteLLM. Add `custom_openai=True` and `base_url`. Installing `crewai[litellm]` also silences it, but then LiteLLM does the routing. | | The error mentions another provider's API key | The model string matched one of CrewAI's own providers. Add `custom_openai=True`. | | `401 missing_api_key` or `invalid_api_key` | The key did not reach the request. Pass `api_key=` explicitly and check `TOKENS_API_KEY` is set in the process that runs the crew. See [API keys](/docs/api-keys). | | `404 model_not_found` | The id is wrong, or CrewAI stripped an `openai/` that belonged to the id. Copy the id from [/models](/models); for ids that start with `openai/`, write the prefix twice. | | `402 insufficient_credits` | The plan credits and wallet cannot cover the request. Top up in [billing](/dashboard/billing). Loops that run up to `max_iter` steps use credits fast. | | `429 rate_limited` or `concurrency_limit` | Too many requests. Set `max_rpm`, run tasks in sequence, wait the `Retry-After` seconds. `window_exhausted` means a plan window is used up, so retrying will not help until it resets. | | Timeouts | Raise `timeout` on the `LLM`, or stream. A `504 upstream_timeout` comes from the upstream after 600 seconds. Try a smaller request. | Every code is in [Errors](/docs/errors). When you ask [support](/docs/support) for help, include the `x-tokens-request-id` response header. The examples on this page do not expose response headers, so reproduce the call once with [cURL](/docs/curl) and `-i` to read it. --- # Laravel AI SDK > Use Tokens from the official Laravel AI SDK (laravel/ai) with the openai-compatible provider: config/ai.php, .env, a first call, streaming, tools, timeouts and error handling for Laravel 12 and 13. Section: SDKs & Libraries. Page: https://tokens.bd/docs/laravel-ai-sdk The Laravel AI SDK (`laravel/ai`) is Laravel's official package for agents, text generation and other AI features. It has an `openai-compatible` driver that talks to any server with an OpenAI-style `/chat/completions` endpoint, which is what Tokens is. You declare a provider in `config/ai.php` with the Tokens URL and key, then pick it by name when you call an agent. If you only want the plain OpenAI client for PHP, without Laravel's agent layer, see [PHP](/docs/php). :::note[What was checked] Based on the Laravel AI SDK documentation for Laravel 13.x and 12.x (laravel.com/docs, checked October 2026) and on the `laravel/ai` 1.2.0 source (released 7 October 2026). The 12.x page does not describe the `openai-compatible` driver, so that part comes from the 13.x docs and the package source. The code was checked against the documentation and source, not run end to end against Tokens. ::: ## What you need - PHP 8.3 or newer and Laravel 12 or 13. `laravel/ai` 1.2.0 requires `php ^8.3` and `illuminate/* ^12.0|^13.0`. - A Tokens key from [API keys](/docs/api-keys). - A model id from [/models](/models). Use one that supports tool calling if your agents use tools ([Choosing a model](/docs/choosing-a-model)). ## Install ```bash composer require laravel/ai php artisan vendor:publish --provider="Laravel\Ai\AiServiceProvider" php artisan migrate ``` The publish step creates `config/ai.php` and a migration. The migration creates the `agent_conversations` and `agent_conversation_messages` tables that the SDK uses to store conversations. Run it even if you do not plan to use conversation memory yet. If you already had `config/ai.php` from an older release, run `composer update laravel/ai` and compare your file with the current published one. The `openai-compatible` entry may be missing from an old copy, and you can add it by hand as shown next. ## Configure the Tokens provider Add a provider to the `providers` array in `config/ai.php`. The name is yours to choose; this page uses `tokens`. ```php title="config/ai.php" return [ 'default' => 'tokens', // ... 'providers' => [ // ... the providers that ship in the file stay as they are 'tokens' => [ 'driver' => 'openai-compatible', 'url' => env('TOKENS_BASE_URL', 'https://tokens.bd/v1'), 'key' => env('TOKENS_API_KEY'), 'models' => [ 'text' => [ 'default' => env('TOKENS_MODEL', 'deepseek/deepseek-v4.1-flash'), ], ], ], ], ]; ``` ```ini title=".env" TOKENS_API_KEY=tok_live_your_key ``` How each option works: - **`url`** is required. Use `https://tokens.bd/v1` including `/v1`. The driver removes a trailing slash and posts to `chat/completions` under it, so the request goes to `https://tokens.bd/v1/chat/completions`. - **`key`** is optional in the SDK and sent as a bearer token when present. For Tokens it is required. - **`models.text.default`** is the model used when a call does not name one. Without it, a call that does not pass `model:` throws an `InvalidArgumentException` that says the provider needs a default text model. - **`'default' => 'tokens'`** makes this provider the default for text. If you would rather leave the default alone, skip it and pass `provider: 'tokens'` on each call, as the examples below do. :::warning[Keep the key out of config files] Read the key with `env()` as above and keep `.env` out of version control. After you change `.env` or `config/ai.php` on a server that caches config, run `php artisan config:clear` (or `config:cache` again). ::: ### Laravel 12 and 13: the variable names differ, the config does not The variable names in the Laravel docs differ between versions, and they only matter for the built-in `openai` provider: | Docs | Variable for a custom OpenAI URL | `openai-compatible` provider | | --- | --- | --- | | Laravel 13.x | `OPENAI_URL` | Documented. Env names in the shipped `config/ai.php`: `OPENAI_COMPATIBLE_URL`, `OPENAI_COMPATIBLE_API_KEY`. | | Laravel 12.x | `OPENAI_BASE_URL` | Not described on the page. The driver is in the `laravel/ai` package, which supports Laravel 12 and 13. | A variable name is only whatever your `config/ai.php` passes to `env()`. The `tokens` entry above uses its own names, so it works on both versions. If you prefer the entry that already ships in the file, set `OPENAI_COMPATIBLE_URL=https://tokens.bd/v1` and `OPENAI_COMPATIBLE_API_KEY` in `.env` and add `models.text.default` to that entry. ### Why not the built-in openai provider The built-in `openai` provider also accepts a custom `url`. In the package source it sends requests to `{url}/responses`, the OpenAI Responses API, while the `openai-compatible` driver sends them to `{url}/chat/completions`. Tokens serves both ([Chat Completions](/docs/chat-completions), [Responses](/docs/responses)). This page uses `openai-compatible` because it is the driver the docs describe for third-party gateways and it sends the Chat Completions format, which is the most widely supported one. ## Make a first call An anonymous agent needs no class. Save this as a route or run it in `php artisan tinker`: ```php use function Laravel\Ai\agent; $response = agent( instructions: 'Answer in one short paragraph.', )->prompt( 'When should I use a queue instead of running code in the request?', provider: 'tokens', model: 'deepseek/deepseek-v4.1-flash', ); echo $response->text; ``` `$response->text` is the reply. `(string) $response` gives the same text. `provider` accepts a plain string for a provider you named in `config/ai.php`. You can drop `model:` if `models.text.default` is set. ## Check that it works First list the models the key can use, outside Laravel: ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` Then, inside the app: ```bash php artisan tinker ``` ```php \Laravel\Ai\agent(instructions: 'Reply with one word.')->prompt('Say ready.', provider: 'tokens')->text ``` A short reply means the URL, key and model are right. A `404 model_not_found` means the id is wrong; a `401` means the key did not arrive. See Troubleshooting below. ## An agent class For anything reused, generate an agent class and pin its provider and model with attributes: ```bash php artisan make:agent ReviewCoach ``` ```php title="app/Ai/Agents/ReviewCoach.php" namespace App\Ai\Agents; use Laravel\Ai\Attributes\Model; use Laravel\Ai\Attributes\Provider; use Laravel\Ai\Attributes\Timeout; use Laravel\Ai\Contracts\Agent; use Laravel\Ai\Promptable; use Stringable; #[Provider('tokens')] #[Model('deepseek/deepseek-v4.1-flash')] #[Timeout(120)] class ReviewCoach implements Agent { use Promptable; public function instructions(): Stringable|string { return 'You review pull request descriptions and suggest one improvement.'; } } ``` ```php $response = (new \App\Ai\Agents\ReviewCoach)->prompt('Fix login redirect'); echo $response->text; ``` The `Provider`, `Model` and `Timeout` attributes are in the Laravel docs' agent configuration section. `Provider` accepts a `Lab` enum value, a string or an array. ## Stream responses `stream()` returns a `StreamableAgentResponse`. Returning it from a route sends it to the browser as server-sent events: ```php title="routes/web.php" use function Laravel\Ai\agent; Route::get('/ask', function () { return agent(instructions: 'Answer briefly.') ->stream('Explain job batching in Laravel.', provider: 'tokens'); }); ``` To handle the events yourself, for example in an Artisan command, loop over the stream: ```php use Laravel\Ai\Streaming\Events\Error; use Laravel\Ai\Streaming\Events\TextDelta; $stream = agent(instructions: 'Answer briefly.') ->stream('Explain job batching in Laravel.', provider: 'tokens'); foreach ($stream as $event) { if ($event instanceof TextDelta) { echo $event->delta; } elseif ($event instanceof Error) { // $event->type holds the error code, $event->message the text fwrite(STDERR, "{$event->type}: {$event->message}\n"); } } ``` The Laravel docs also describe a `then()` callback that runs when the whole response has been streamed; its `StreamedAgentResponse` carries `text`, `events` and `usage`. The driver asks Tokens for `stream_options.include_usage`, so token counts arrive at the end of the stream. See [Streaming](/docs/streaming) for the wire format. ## Tools Tools are classes with a description, a schema and a `handle` method. Generate one: ```bash php artisan make:tool GetWeather ``` ```php title="app/Ai/Tools/GetWeather.php" namespace App\Ai\Tools; use Illuminate\Contracts\JsonSchema\JsonSchema; use Laravel\Ai\Contracts\Tool; use Laravel\Ai\Tools\Request; use Stringable; class GetWeather implements Tool { public function description(): Stringable|string { return 'Current weather for a city.'; } public function handle(Request $request): Stringable|string { return $request['city'].': light rain, 29C'; // call your real weather source here } public function schema(JsonSchema $schema): array { return [ 'city' => $schema->string()->required(), ]; } } ``` Pass it to an anonymous agent, or return it from the `tools()` method of an agent class: ```php use App\Ai\Tools\GetWeather; use function Laravel\Ai\agent; $response = agent( instructions: 'Use the weather tool when asked about weather.', tools: [new GetWeather], )->prompt('Is it raining in Dhaka?', provider: 'tokens'); echo $response->text; ``` The SDK sends the tool definitions in the Chat Completions format and runs the loop for you. Tool calling only works on models that support it, so check the model page in [/models](/models) and read [Tool calling](/docs/tool-calling). ## Timeouts The agent timeout defaults to 60 seconds. Raise it per call, per class, or both: ```php $response = $agent->prompt('...', provider: 'tokens', timeout: 180); ``` For a class, use `#[Timeout(180)]`. The value is the HTTP client timeout in seconds for the request. In Laravel's HTTP client it is Guzzle's `timeout` option, which Guzzle describes as the total time for the request. A long streamed answer can therefore be cut at that limit, so set it above your longest expected generation. Other limits to know: - The gateway waits up to 600 seconds for response headers from the upstream, so long generations are fine on its side. - PHP itself can stop a web request first. Check `max_execution_time` in `php.ini` (the PHP default for web requests is 30 seconds), or run long jobs in a queue. ## Handle errors The SDK maps some HTTP statuses to its own exceptions. For Tokens: | Status | Exception | Tokens codes | | --- | --- | --- | | 402 | `Laravel\Ai\Exceptions\InsufficientCreditsException` | `insufficient_credits`, `no_funding`, `outstanding_debt`, `member_cap_reached` | | 429 | `Laravel\Ai\Exceptions\RateLimitedException` | `rate_limited`, `concurrency_limit`, `window_exhausted`, `model_limit_reached`, `rate_limit_exceeded` | | 502, 503, 504 | `Laravel\Ai\Exceptions\ProviderOverloadedException` | `upstream_unreachable`, `no_upstream_available`, `upstream_timeout` | | connection failure | `Laravel\Ai\Exceptions\ProviderConnectionException` | none; the request never reached Tokens | | other 4xx and 5xx | `Illuminate\Http\Client\RequestException` | `invalid_api_key`, `model_not_found`, `model_not_allowed_on_key` and the rest | The Tokens `code` is in the response body. The SDK exceptions keep the original error as the previous exception: ```php use Illuminate\Http\Client\RequestException; use Laravel\Ai\Exceptions\AiException; try { $text = agent(instructions: 'Say hi.')->prompt('hi', provider: 'tokens')->text; } catch (AiException|RequestException $e) { $request = $e instanceof RequestException ? $e : $e->getPrevious(); $response = $request instanceof RequestException ? $request->response : null; logger()->warning('Tokens call failed', [ 'status' => $response?->status(), 'code' => $response?->json('error.code'), 'request_id' => $response?->header('x-tokens-request-id'), ]); throw $e; } ``` Log `x-tokens-request-id` with every failure. Support can trace a request from it ([Errors](/docs/errors)). Retrying `429 window_exhausted` does not help until the window resets; `Retry-After` says when. ## What the openai-compatible driver covers In the package source the driver implements text generation (with streaming and tools), embeddings and transcription. Image generation and text-to-speech go through other providers in the SDK, not this one. Embeddings need `models.embeddings.default` in the provider config; Tokens serves embeddings ([Embeddings](/docs/embeddings)), but this page does not walk through that setup. ## Troubleshooting | Symptom | Cause and fix | | --- | --- | | `The [tokens] openai-compatible provider requires a 'url' to be configured` | `url` is empty. Check the `tokens` entry and clear the config cache. | | `... requires a default text model` | Add `models.text.default`, or pass `model:` on the call. | | `404 model_not_found` | The model id is wrong. Copy it from [/models](/models). | | `404` on every call, error mentions a path | The URL is missing `/v1`. Use `https://tokens.bd/v1`. | | `401 missing_api_key` or `invalid_api_key` | `TOKENS_API_KEY` is empty in the running process, often because config is cached. Run `php artisan config:clear`. | | `RateLimitedException` | Read `error.code` as shown above. `rate_limited` and `concurrency_limit` clear after `Retry-After`; `window_exhausted` clears at the plan reset. | | `InsufficientCreditsException` | Top up in [billing](/dashboard/billing). | | `cURL error 28: Operation timed out` | The `timeout` value is lower than the response time. Raise it. | | Streamed text stops partway | The total timeout was reached, or PHP's `max_execution_time` ended the request. Raise both. | Every Tokens error code is in [Errors](/docs/errors), and [Troubleshooting](/docs/troubleshooting) has the general fixes. --- # PHP > Use Tokens from PHP and Laravel: the openai-php/client package, the openai-php/laravel package and plain cURL. Base URL, a first call, streaming, tool calls, timeouts and Tokens error codes. Section: SDKs & Libraries. Page: https://tokens.bd/docs/php Tokens speaks the OpenAI Chat Completions protocol, so any PHP code that can send an HTTP request can use it. This page covers three ways: the community package `openai-php/client` in plain PHP, `openai-php/laravel` inside a Laravel app, and raw cURL with no packages. For Laravel's own agent package, see [Laravel AI SDK](/docs/laravel-ai-sdk). :::note[What was checked] Based on the README and source of `openai-php/client` and `openai-php/laravel`, both version 0.21.0 (released 17 September 2026, checked October 2026), and on the PHP and Guzzle documentation. The code was checked against the documentation and source, not run end to end against Tokens. These packages are community-maintained, not published by OpenAI or Tokens. ::: ## What you need - PHP 8.2 or newer (both packages require `php ^8.2`) and [Composer](https://getcomposer.org/). - A Tokens key from [API keys](/docs/api-keys), exported as `TOKENS_API_KEY`. - A model id from [/models](/models). The base URL for every example is `https://tokens.bd/v1`. It includes `/v1`. In `openai-php/client` you pass it to `withBaseUri()` exactly like that: the library appends a `/` and then the resource path, so requests go to `https://tokens.bd/v1/chat/completions`. Do not add a trailing slash or `/chat/completions` yourself. Write the full URL with `https://`. If you leave the scheme off, the library adds `https://` for you. ## openai-php/client ### Install ```bash composer require openai-php/client guzzlehttp/guzzle export TOKENS_API_KEY="tok_live_your_key" ``` The package needs a PSR-18 HTTP client. The README says to allow the `php-http/discovery` Composer plugin or to install a client such as Guzzle yourself. Installing Guzzle as above is the simplest. ### Create the client and make a call ```php title="hello.php" withApiKey((string) getenv('TOKENS_API_KEY')) ->withBaseUri('https://tokens.bd/v1') ->withHttpClient(new GuzzleHttp\Client([ 'connect_timeout' => 10, 'timeout' => 300, ])) ->make(); $response = $client->chat()->create([ 'model' => 'deepseek/deepseek-v4.1-flash', 'messages' => [ ['role' => 'system', 'content' => 'Answer in one short paragraph.'], ['role' => 'user', 'content' => 'When should I use a queue instead of running code in the request?'], ], 'max_tokens' => 400, ]); echo $response->choices[0]->message->content, PHP_EOL; echo "{$response->usage->promptTokens} in, {$response->usage->completionTokens} out", PHP_EOL; ``` ```bash php hello.php ``` `OpenAI::factory()`, `withApiKey()`, `withBaseUri()`, `withHttpClient()` and `make()` are the factory methods from the README. `withHttpHeader()` adds a header to every request if you need one. Do not use `OpenAI::client($key)` for Tokens: it has no base URI argument and would send your key to OpenAI. ### Choose the model `model` takes the Tokens id exactly as listed in [/models](/models), for example `deepseek/deepseek-v4.1-flash`. To check ids from code: ```php foreach ($client->models()->list()->data as $model) { echo $model->id, PHP_EOL; } ``` A wrong id returns `404 model_not_found`. ### Stream responses Use `createStreamed()` and loop over the result. Ask for usage, or the stream carries no token counts: ```php $stream = $client->chat()->createStreamed([ 'model' => 'deepseek/deepseek-v4.1-flash', 'messages' => [ ['role' => 'user', 'content' => 'Write a haiku about merge conflicts.'], ], 'stream_options' => ['include_usage' => true], ]); foreach ($stream as $chunk) { $text = $chunk->choices[0]->delta->content ?? null; if ($text !== null) { echo $text; flush(); } if ($chunk->usage !== null) { echo PHP_EOL, "{$chunk->usage->promptTokens} in, {$chunk->usage->completionTokens} out", PHP_EOL; } } ``` With `include_usage` on, the last chunk has an empty `choices` list and carries only `usage`. The `?? null` handles that. `usage` is `null` on every other chunk. See [Streaming](/docs/streaming). If the key or a limit fails during a stream, the library throws `OpenAI\Exceptions\ErrorException` from inside the `foreach`, so put the loop inside your `try` block. ### Tool calls Tool calling works on models that support it. Check the model's page in [/models](/models) first. ```php $tools = [[ 'type' => 'function', 'function' => [ 'name' => 'get_weather', 'description' => 'Current weather for a city', 'parameters' => [ 'type' => 'object', 'properties' => ['city' => ['type' => 'string']], 'required' => ['city'], ], ], ]]; $messages = [['role' => 'user', 'content' => 'Is it raining in Dhaka?']]; $first = $client->chat()->create([ 'model' => 'deepseek/deepseek-v4.1-flash', 'messages' => $messages, 'tools' => $tools, ]); $message = $first->choices[0]->message; if ($message->toolCalls !== []) { $messages[] = [ 'role' => 'assistant', 'content' => $message->content, 'tool_calls' => array_map(fn ($call) => [ 'id' => $call->id, 'type' => 'function', 'function' => ['name' => $call->function->name, 'arguments' => $call->function->arguments], ], $message->toolCalls), ]; foreach ($message->toolCalls as $call) { $args = json_decode($call->function->arguments, true); $result = ['city' => $args['city'], 'condition' => 'light rain', 'temp_c' => 29]; // your real lookup here $messages[] = [ 'role' => 'tool', 'tool_call_id' => $call->id, 'content' => json_encode($result), ]; } $final = $client->chat()->create([ 'model' => 'deepseek/deepseek-v4.1-flash', 'messages' => $messages, 'tools' => $tools, ]); echo $final->choices[0]->message->content, PHP_EOL; } ``` The request and response shapes are in [Tool calling](/docs/tool-calling). ### Timeouts for long requests The README says the default timeout depends on the HTTP client you use, and the way to raise it is to pass a configured client to `withHttpClient()`. With Guzzle: | Option | Meaning (from the Guzzle docs) | | --- | --- | | `connect_timeout` | Seconds to wait while connecting. Default 0, which waits forever. | | `timeout` | Total time for the whole request, in seconds. Default 0, which waits forever. | | `read_timeout` | Timeout for individual reads on a streamed body. Defaults to the `default_socket_timeout` ini setting. | Set `timeout` above your longest non-streamed answer, as in the first example. For streamed answers, a total `timeout` also covers the time you spend reading the stream, so a long answer can be cut off. For streams, use a client that has no total limit and relies on `read_timeout`: ```php $streamClient = OpenAI::factory() ->withApiKey((string) getenv('TOKENS_API_KEY')) ->withBaseUri('https://tokens.bd/v1') ->withHttpClient(new GuzzleHttp\Client([ 'connect_timeout' => 10, 'timeout' => 0, 'read_timeout' => 120, ])) ->make(); ``` On the Tokens side, the gateway waits up to 600 seconds for response headers from the upstream. Streaming means you are not holding a silent connection that long. PHP can also stop a long request before your client does. The web server setting `max_execution_time` in `php.ini` defaults to 30 seconds for web requests (the command line has no limit). For long generations in a web app, stream the answer, or run the call in a queue job. ### Handle errors The library reads Tokens' OpenAI-style error body, so `ErrorException` carries the Tokens `code`: ```php use OpenAI\Exceptions\ErrorException; use OpenAI\Exceptions\TransporterException; try { $client->chat()->create([ 'model' => 'deepseek/deepseek-v4.1-flash', 'messages' => [['role' => 'user', 'content' => 'hi']], ]); } catch (ErrorException $e) { $requestId = $e->response->getHeaderLine('x-tokens-request-id'); error_log(sprintf('%d %s %s request=%s', $e->getStatusCode(), $e->getErrorCode(), $e->getErrorMessage(), $requestId)); if ($e->getErrorCode() === 'insufficient_credits') { // top up at https://tokens.bd/dashboard/billing } } catch (TransporterException $e) { error_log('Network problem, timeout or server error: ' . $e->getMessage()); } ``` From the library source, with Guzzle as the HTTP client: a 4xx response with a JSON error body (401, 402, 403, 404 and 429 included) becomes an `ErrorException`. Network failures, timeouts and 5xx responses become a `TransporterException`, and the Guzzle exception sits in `$e->getPrevious()`, which has the response if one arrived. The library also defines `RateLimitException` and `ServerException` for HTTP clients that do not throw on error statuses, and each has a public `$response`. Log `x-tokens-request-id` with every failure so [support](/docs/support) can trace the request. The codes you are likely to meet are `invalid_api_key` (401), `model_not_allowed_on_key` and `tier_permission_denied` (403), `insufficient_credits` (402), and `rate_limited`, `concurrency_limit` and `window_exhausted` (429). [Errors](/docs/errors) lists them all. ## openai-php/laravel The Laravel package wraps the same client in a service provider and an `OpenAI` facade. It requires PHP 8.2+ and Laravel `^11.29`, `^12.12` or `^13.0`. ### Install ```bash composer require openai-php/laravel php artisan openai:install ``` The second command creates `config/openai.php` and appends blank `OPENAI_API_KEY` and `OPENAI_ORGANIZATION` lines to `.env`. Tokens does not use an organization, so delete the `OPENAI_ORGANIZATION` line. ### Configure The README documents these variables: ```ini title=".env" OPENAI_API_KEY=tok_live_your_key OPENAI_BASE_URL=https://tokens.bd/v1 OPENAI_REQUEST_TIMEOUT=300 ``` | Variable | Config key | Notes | | --- | --- | --- | | `OPENAI_API_KEY` | `api_key` | Your Tokens key. | | `OPENAI_BASE_URL` | `base_uri` | Defaults to `api.openai.com/v1`. Use `https://tokens.bd/v1`, with `/v1`. | | `OPENAI_REQUEST_TIMEOUT` | `request_timeout` | Seconds, default 30. Raise it for long answers. | :::warning[Use Tokens-specific names if anything else reads OPENAI_API_KEY] `OPENAI_API_KEY` is the variable name many packages read. For example, the built-in `openai` provider in `laravel/ai` reads it and sends it to OpenAI by default. If you put a Tokens key there, another package could send it to the wrong place. Safer: edit `config/openai.php` to read your own names and leave `OPENAI_*` unset. ::: ```php title="config/openai.php" return [ 'api_key' => env('TOKENS_API_KEY'), 'base_uri' => env('TOKENS_BASE_URL', 'https://tokens.bd/v1'), 'request_timeout' => env('TOKENS_REQUEST_TIMEOUT', 300), ]; ``` The published file also lists `organization` and `project`. Leave them out or leave them `null`. After changing `.env` or config on a server that caches config, run `php artisan config:clear`. ### Make a call ```php use OpenAI\Laravel\Facades\OpenAI; $response = OpenAI::chat()->create([ 'model' => 'deepseek/deepseek-v4.1-flash', 'messages' => [ ['role' => 'user', 'content' => 'Explain job batching in Laravel in two sentences.'], ], ]); echo $response->choices[0]->message->content; ``` The facade exposes the same `chat()`, `models()` and other resources as the client above, so streaming, tool calls and error handling are identical to the previous section. The README's own example uses `OpenAI::responses()`; Tokens also serves `/v1/responses` ([Responses](/docs/responses)), but the Chat Completions form above is the one this page checked. ### Stream from a route ```php title="routes/web.php" use Illuminate\Support\Facades\Route; use OpenAI\Laravel\Facades\OpenAI; Route::get('/ask', function () { return response()->stream(function () { $stream = OpenAI::chat()->createStreamed([ 'model' => 'deepseek/deepseek-v4.1-flash', 'messages' => [['role' => 'user', 'content' => 'Explain queues in Laravel.']], ]); foreach ($stream as $chunk) { $text = $chunk->choices[0]->delta->content ?? null; if ($text !== null) { echo $text; if (ob_get_level() > 0) { ob_flush(); } flush(); } } }, 200, ['Content-Type' => 'text/plain; charset=utf-8', 'Cache-Control' => 'no-cache']); }); ``` ### Timeouts in Laravel The package builds its Guzzle client with `request_timeout` as Guzzle's `timeout` option. That is a total limit, so it also bounds a streamed answer. Set it above your longest answer. If you need Guzzle's `read_timeout` instead, build the client with the factory (previous section) and register it in your own service provider in place of the package's. ## Raw cURL No Composer packages needed. This uses PHP's `curl` extension. ```php title="curl-hello.php" true, CURLOPT_HTTPHEADER => [ 'Authorization: Bearer ' . getenv('TOKENS_API_KEY'), 'Content-Type: application/json', ], CURLOPT_POSTFIELDS => json_encode([ 'model' => 'deepseek/deepseek-v4.1-flash', 'messages' => [['role' => 'user', 'content' => 'Say hello in five words.']], 'max_tokens' => 100, ], JSON_THROW_ON_ERROR), CURLOPT_RETURNTRANSFER => true, CURLOPT_CONNECTTIMEOUT => 10, CURLOPT_TIMEOUT => 300, CURLOPT_HEADERFUNCTION => function ($ch, string $header) use (&$requestId): int { if (stripos($header, 'x-tokens-request-id:') === 0) { $requestId = trim(substr($header, strlen('x-tokens-request-id:'))); } return strlen($header); }, ]); $body = curl_exec($ch); if ($body === false) { fwrite(STDERR, 'cURL error ' . curl_errno($ch) . ': ' . curl_error($ch) . PHP_EOL); exit(1); } $status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE); curl_close($ch); $data = json_decode($body, true, flags: JSON_THROW_ON_ERROR); if ($status >= 400) { $error = $data['error'] ?? []; fwrite(STDERR, sprintf("%d %s %s request=%s\n", $status, $error['code'] ?? '', $error['message'] ?? '', $requestId)); exit(1); } echo $data['choices'][0]['message']['content'], PHP_EOL; ``` `CURLOPT_TIMEOUT` is the total time allowed for the request, and `0` (the default) means no limit. `CURLOPT_CONNECTTIMEOUT` covers only the connection. To stream, send `"stream": true` and read the body as it arrives with `CURLOPT_WRITEFUNCTION`. The body is a series of `data: {...}` lines that end with `data: [DONE]`: ```php title="curl-stream.php" true, CURLOPT_HTTPHEADER => [ 'Authorization: Bearer ' . getenv('TOKENS_API_KEY'), 'Content-Type: application/json', ], CURLOPT_POSTFIELDS => json_encode([ 'model' => 'deepseek/deepseek-v4.1-flash', 'messages' => [['role' => 'user', 'content' => 'Write a haiku about merge conflicts.']], 'stream' => true, 'stream_options' => ['include_usage' => true], ], JSON_THROW_ON_ERROR), CURLOPT_CONNECTTIMEOUT => 10, CURLOPT_TIMEOUT => 0, // no total limit; streams can be long CURLOPT_WRITEFUNCTION => function ($ch, string $data) use (&$buffer, &$errorBody): int { if (curl_getinfo($ch, CURLINFO_RESPONSE_CODE) >= 400) { $errorBody .= $data; // an error comes back as plain JSON, not as a stream return strlen($data); } $buffer .= $data; while (($pos = strpos($buffer, "\n")) !== false) { $line = trim(substr($buffer, 0, $pos)); $buffer = substr($buffer, $pos + 1); if (!str_starts_with($line, 'data:')) { continue; } $payload = trim(substr($line, 5)); if ($payload === '[DONE]') { continue; } $chunk = json_decode($payload, true); $text = $chunk['choices'][0]['delta']['content'] ?? ''; if ($text !== '') { echo $text; flush(); } } return strlen($data); }, ]); $ok = curl_exec($ch); $status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE); if ($ok === false) { fwrite(STDERR, 'cURL error ' . curl_errno($ch) . ': ' . curl_error($ch) . PHP_EOL); exit(1); } if ($status >= 400) { fwrite(STDERR, "HTTP $status: $errorBody" . PHP_EOL); exit(1); } curl_close($ch); echo PHP_EOL; ``` The write function must return the number of bytes it received, or cURL aborts the transfer. The `choices[0]` lookup uses `??` because the final chunk has no choices when `include_usage` is on. ## Check that it works Run any of the programs above. A short answer printed means the key, the URL and the model id are right. To check without PHP: ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` ## Limits - `openai-php/client` models only the OpenAI API. Anything Tokens adds on top (for example the usage endpoint described in [Models and usage](/docs/models-and-usage)) is not a method in the package. Call it with cURL or Guzzle. - Tokens limits requests per minute and concurrent requests per account. A queue worker pool that runs many jobs at once can hit `429 concurrency_limit`. See [Rate limits](/docs/rate-limits). - Do not call Tokens from browser JavaScript. Tokens sends no CORS headers and the key would be visible. Make the call from PHP and return the result. - PHP-FPM and web servers buffer output. Streaming to a browser may need `ob_flush()`, `flush()` and, on nginx, response buffering turned off for that route. ## Troubleshooting | Symptom | Cause and fix | | --- | --- | | `404 unsupported_endpoint` | The base URI is wrong. Use `https://tokens.bd/v1` with `/v1` and no `/chat/completions` on the end. | | `404 model_not_found` | The model id is wrong. Copy it from [/models](/models). | | `401 missing_api_key` or `invalid_api_key` | The key was empty or wrong. `getenv()` returns `false` if the variable is not set in the PHP process. PHP-FPM often does not see shell variables, so use `.env`, a config file or `clear_env = no`. In Laravel, run `php artisan config:clear`. | | `To use stream requests you must provide an stream handler closure via the OpenAI factory` (exception message) | You passed a custom HTTP client that is not Guzzle or Symfony. Use Guzzle, or add `withStreamHandler()` as the README describes. | | `cURL error 28: Operation timed out` or `TransporterException` after a long wait | Your own timeout is shorter than the response. Raise `timeout`, or stream. | | The page stops partway through a stream | The total timeout or `max_execution_time` ended the request. Raise both. | | `402 insufficient_credits` | Top up in [billing](/dashboard/billing). | | `429 rate_limited` or `concurrency_limit` | Wait `Retry-After` seconds or lower your parallelism. `window_exhausted` means the plan window is used up; retrying will not help until it resets. | Every Tokens error code is in [Errors](/docs/errors), with the fix for each in [Troubleshooting](/docs/troubleshooting). :::warning Keep the key in an environment variable or a secrets manager, never in a committed file. If it leaks, rotate it in [/dashboard/keys](/dashboard/keys). The old secret stops working immediately. ::: --- # Authentication > Base URLs, the two supported auth headers, the tok_live_ key format, and what the 401 and 403 error codes mean. Section: API Reference. Page: https://tokens.bd/docs/authentication Every request to the Tokens API authentication layer needs one platform key, sent in a header. This page covers the base URLs, the headers we accept, the key format, and what each authentication error means. ## Base URLs Tokens exposes two surfaces on the same host. Which one you use depends on the client, not the model. | Client style | Base URL | Typical tools | | -------------------- | ---------------------- | ------------------------------------------------------------------- | | OpenAI-compatible | `https://tokens.bd/v1` | OpenAI SDKs, Cursor, Cline, Aider, OpenCode, Codex CLI | | Anthropic-compatible | `https://tokens.bd` | Anthropic SDKs, Claude Code (they append `/v1/messages` themselves) | If a tool asks for an "API base" or "base URL" and is OpenAI-flavored, include `/v1`. If it is Anthropic-flavored, leave `/v1` off, because the SDK adds it. Getting this wrong is the most common cause of a 404 on the first request. Tools that need to discover the endpoints can read them from a public, unauthenticated config endpoint: ```bash curl https://tokens.bd/api/gateway/config ``` ```json { "openaiBaseUrl": "https://tokens.bd/v1", "anthropicBaseUrl": "https://tokens.bd" } ``` ## Send your API key in a header We accept the key in either of two headers. Use whichever your client sends by default. | Header | Format | Sent by | | --------------- | -------------------------- | ---------------------------------------------------------------- | | `Authorization` | `Bearer tok_live_your_key` | OpenAI SDKs, Claude Code with `ANTHROPIC_AUTH_TOKEN`, most tools | | `x-api-key` | `tok_live_your_key` | Anthropic SDKs | If both are present, the `Authorization: Bearer` header wins. An `Authorization` header with any scheme other than `Bearer` is ignored, and the gateway then looks for `x-api-key`. :::code-tabs ```bash title="Bearer" curl https://tokens.bd/v1/models \ -H "Authorization: Bearer $TOKENS_API_KEY" ``` ```bash title="x-api-key" curl https://tokens.bd/v1/models \ -H "x-api-key: $TOKENS_API_KEY" ``` ::: ## Key format Keys look like `tok_live_` followed by 48 hexadecimal characters. The full secret is shown once, when you create the key in [the dashboard](/dashboard/keys). We store only a hash, so a lost key cannot be recovered; create a new one instead. Key creation, spend caps, allowed-model lists and rotation are covered in [API keys](/docs/api-keys). ## Keep the key in an environment variable All examples in these docs read the key from `TOKENS_API_KEY`: ```bash export TOKENS_API_KEY="tok_live_your_key" ``` On Windows PowerShell: ```powershell $env:TOKENS_API_KEY = "tok_live_your_key" ``` Then read it in code rather than pasting it: :::code-tabs ```python title="Python" import os from openai import OpenAI client = OpenAI( base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], ) ``` ```typescript title="Node.js" import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY, }); ``` ::: :::warning[Do not commit keys] A key in a committed `.env`, config file or notebook is a leaked key. Add those files to `.gitignore`. If a key does leak, rotate it in the dashboard; the old secret stops working immediately. ::: ## Server-side only: no browser calls The API does not send CORS headers on its responses, so `fetch` from a web page will fail in the browser even with a valid key. This is deliberate: a key shipped to a browser is readable by anyone who opens dev tools. Call Tokens from your backend, a serverless function, a CLI, or a coding agent, and have your frontend talk to that. ## What 401 and 403 mean Authentication failures return the standard error body with a machine-readable `code`: ```json { "error": { "message": "Invalid API key.", "type": "authentication_error", "code": "invalid_api_key", "param": null, "request_id": "8f0c7a4e-2b1d-4c55-9a51-3f7e2d9b6c10" } } ``` | Status | Code | Meaning | Fix | | ------ | ---------------------------- | -------------------------------------------------------------------- | ---------------------------------------------------------------------------- | | 401 | `missing_api_key` | No `Authorization: Bearer` or `x-api-key` header arrived | Check the env var is set in the shell that runs the tool | | 401 | `invalid_api_key` | The key does not match any key we issued | Check for truncation or stray quotes; copy the key again or create a new one | | 403 | `key_inactive` | The key was revoked or rotated | Use the current secret or create a new key | | 403 | `key_expired` | The key had an expiry date that has passed | Create a new key | | 403 | `account_suspended` | The account is suspended | Contact [support](/docs/support) | | 403 | `model_not_allowed_on_key` | The key has an allowed-models list that excludes this model | Use a listed model or a different key | | 403 | `monthly_spend_cap_exceeded` | The key reached its monthly spend cap | Wait for the next month or use another key | | 403 | `tier_permission_denied` | Your plan does not include this model and you have no wallet balance | See [plans and wallet](/docs/plans-and-wallet) | A 401 or 403 with code `upstream_auth_error` is different: it means the gateway's own credentials for an upstream provider were refused, not yours. Your key is fine; retry later or try another model, and include the request id if you contact support. The full list of codes, including billing and rate-limit errors, is in [errors](/docs/errors). To confirm a key works end to end, the [quickstart](/docs/quickstart) has a one-line test request. --- # Chat Completions > POST /v1/chat/completions: request fields, a full request and response, the usage object, and how the gateway treats max_tokens, n and streaming. Section: API Reference. Page: https://tokens.bd/docs/chat-completions `POST https://tokens.bd/v1/chat/completions` is the OpenAI-compatible chat completions API, and the endpoint most SDKs and coding agents use. The request and response follow OpenAI's format; the gateway checks your key, plan and limits, then forwards the body to the upstream provider for the model you named. ## Make a chat completions request :::code-tabs ```bash title="cURL" curl https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "messages": [ {"role": "system", "content": "You are a concise senior engineer."}, {"role": "user", "content": "What does HTTP 429 mean?"} ], "temperature": 0.2, "max_tokens": 300 }' ``` ```python title="Python" import os from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) resp = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", messages=[ {"role": "system", "content": "You are a concise senior engineer."}, {"role": "user", "content": "What does HTTP 429 mean?"}, ], temperature=0.2, max_tokens=300, ) print(resp.choices[0].message.content) print(resp.usage) ``` ```typescript title="Node.js" import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY }); const resp = await client.chat.completions.create({ model: "deepseek/deepseek-v4.1-flash", messages: [ { role: "system", content: "You are a concise senior engineer." }, { role: "user", content: "What does HTTP 429 mean?" }, ], temperature: 0.2, max_tokens: 300, }); console.log(resp.choices[0].message.content, resp.usage); ``` ::: The `Content-Type: application/json` header matters. Without it the gateway can't read the body, and you get a 400 saying the request must specify a `model` field even though it does. ## Request fields | Field | Type | Notes | | ---------------------------------------------------------------- | ---------------- | ------------------------------------------------------------------------------------------------ | | `model` | string | Required. A catalog id such as `deepseek/deepseek-v4.1-flash`. List yours with `GET /v1/models`. | | `messages` | array | Required. Objects with `role` (`system`, `user`, `assistant`, `tool`) and `content`. | | `temperature` | number | Usually 0 to 2. Some reasoning models ignore or reject it. | | `max_tokens` | integer | Upper bound on generated tokens. Set it; see the note on reservations below. | | `max_completion_tokens` | integer | Newer OpenAI name for the same limit. Some models require it instead of `max_tokens`. | | `stream` | boolean | `true` returns Server-Sent Events. See [streaming](/docs/streaming). | | `stream_options.include_usage` | boolean | With `stream: true`, adds a final chunk carrying `usage`. Off unless you ask. | | `tools` | array | Function definitions. See [tool calling](/docs/tool-calling). | | `tool_choice` | string or object | `"auto"`, `"none"`, `"required"`, or a specific function. Support varies by model. | | `response_format` | object | `{"type": "json_object"}` or a `json_schema` object, where the model supports it. | | `n` | integer | Number of choices, 1 to 4. Values outside that range return 400 `invalid_request`. | | `stop`, `top_p`, `seed`, `presence_penalty`, `frequency_penalty` | various | Passed through as sent. | :::note[Parameters depend on the upstream model] Apart from `model` and `n`, the gateway passes these fields to the provider serving the model without checking them, with two exceptions. For OpenAI's reasoning models (the `o`-series and GPT-5 and later) it renames `max_tokens` to `max_completion_tokens` and removes `temperature` and `top_p` unless they equal 1, because those models reject them. And when the model is served by an Anthropic-only provider, fields that format has no place for are dropped: `response_format`, `reasoning_effort`, `n`, `seed`, `logprobs` and the penalty fields. See [Structured output](/docs/structured-output) and [Reasoning](/docs/reasoning). If a model doesn't support `response_format`, `tools` or a given `temperature`, the provider decides what happens: it may ignore the field or reject the request. Check the model's page in [the catalog](/models) before relying on a feature. ::: The request body can be up to 10 MB. Larger bodies return 413. ## Example response ```json { "id": "chatcmpl-a1b2c3", "object": "chat.completion", "created": 1790000000, "model": "deepseek/deepseek-v4.1-flash", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "429 Too Many Requests: the server is rate limiting you. Back off and retry after the Retry-After interval." }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 27, "completion_tokens": 24, "total_tokens": 51 } } ``` The body comes from the upstream provider, so the exact `id` format, the `model` string and any extra fields (such as `reasoning_content` or `system_fingerprint`) vary by model. ## The usage object `usage` reports what the provider counted: | Field | Meaning | | ------------------------------------- | ----------------------------------------------------------------------------- | | `prompt_tokens` | Input tokens, including any cached ones | | `completion_tokens` | Output tokens, including reasoning tokens on models that report them that way | | `total_tokens` | Sum of the two | | `prompt_tokens_details.cached_tokens` | Input tokens served from the provider's cache, when reported | Billing uses these upstream counts, with cached input priced at the model's cache-read rate where one exists. If a provider sends no usage at all, the gateway estimates from the request and response size. Per-request costs appear in [usage analytics](/docs/usage-and-alerts); per-model rates are on [the model catalog](/models) and [pricing](/pricing). ## How max_tokens affects admission Before forwarding, the gateway reserves the worst-case cost of the request against your plan credits or wallet, using `max_tokens` (or `max_completion_tokens`) for the output side. If you leave it unset, the reservation assumes 8,192 output tokens. Two practical consequences: - A key close to its monthly spend cap can be refused for a large `max_tokens` while a small one still gets through. - If your balance can't cover the requested `max_tokens`, the gateway may lower it to what the balance covers (never below 16). You'll see `finish_reason: "length"` on a shorter answer. Top up in [billing](/dashboard/billing) or lower `max_tokens` yourself. You are charged for actual usage, not the reservation. ## Errors Gateway errors use the OpenAI error shape with a `code` you can branch on, and every response carries an `x-tokens-request-id` header. The full table is in [errors](/docs/errors), and per-minute and concurrency limits are in [rate limits](/docs/rate-limits). --- # Messages (Anthropic API) > Call POST /v1/messages with Anthropic SDKs or curl: base URL, headers, a full example, streaming events, error shapes, and which headers are not forwarded. Section: API Reference. Page: https://tokens.bd/docs/messages Tokens serves the Anthropic Messages API at `POST https://tokens.bd/v1/messages`. This is the endpoint Claude Code and the Anthropic SDKs call, so you can point them at Tokens by changing the base URL and key. You can request any model in your catalog through it, not only Claude models. ## Base URL and headers Anthropic SDKs and Claude Code add `/v1/messages` themselves, so their base URL is the bare host: `https://tokens.bd`. | Header | Value | Required | | ------------------------------ | ------------------------------------------------- | ------------------------------- | | `x-api-key` or `Authorization` | `tok_live_your_key` or `Bearer tok_live_your_key` | Yes | | `anthropic-version` | `2023-06-01` | Recommended; forwarded upstream | | `anthropic-beta` | Beta flags, comma separated | Optional; forwarded upstream | | `content-type` | `application/json` | Yes | The SDKs send `x-api-key` and `anthropic-version` for you. Claude Code with `ANTHROPIC_AUTH_TOKEN` sends a Bearer header instead; both work. ## Send a Messages API request with curl ```bash curl https://tokens.bd/v1/messages \ -H "x-api-key: $TOKENS_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "max_tokens": 512, "system": "You are a concise senior engineer.", "messages": [ {"role": "user", "content": "Explain idempotency keys in two sentences."} ] }' ``` `max_tokens` is required by the Messages API. A typical response: ```json { "id": "msg_01AbCdEf", "type": "message", "role": "assistant", "model": "deepseek/deepseek-v4.1-flash", "content": [ { "type": "text", "text": "An idempotency key is a client-chosen id sent with a request..." } ], "stop_reason": "end_turn", "stop_sequence": null, "usage": { "input_tokens": 31, "output_tokens": 58 } } ``` ## Use the Anthropic Python SDK ```python import os import anthropic client = anthropic.Anthropic( base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"], ) message = client.messages.create( model="deepseek/deepseek-v4.1-flash", max_tokens=512, messages=[{"role": "user", "content": "Explain idempotency keys in two sentences."}], ) print(message.content[0].text) print(message.usage) ``` The TypeScript SDK works the same way: `new Anthropic({ baseURL: "https://tokens.bd", apiKey: process.env.TOKENS_API_KEY })`. For Claude Code, the settings live in `~/.claude/settings.json`; see the setup steps in the [quickstart](/docs/quickstart) or the snippet at [Connect your agent](/dashboard/connect). ## Native and translated models on /v1/messages Each model is served by one or more upstream providers. When a provider speaks the Messages API, your request goes to it as is: thinking blocks, prompt caching, `anthropic-beta` features and the exact usage numbers come back untouched. Tokens prefers such a provider whenever one serves the model. When the model is only available from a provider that speaks the OpenAI format, the gateway translates your Messages request into a chat completions request and translates the answer back. `system`, `messages` (text and images), `max_tokens`, `stop_sequences`, `temperature`, `top_p`, `tools` and `tool_choice` are carried across. Anthropic-only content (prompt caching markers, extended thinking blocks, documents) may be dropped in translation, so test those features against the specific model before depending on them. ## Streaming events Set `"stream": true` and the response is Server-Sent Events in Anthropic's format: ```text event: message_start data: {"type":"message_start","message":{"id":"msg_01AbCdEf","type":"message","role":"assistant","content":[],"model":"deepseek/deepseek-v4.1-flash","stop_reason":null,"usage":{"input_tokens":31,"output_tokens":0}}} event: content_block_start data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}} event: content_block_delta data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"An idempotency"}} event: content_block_stop data: {"type":"content_block_stop","index":0} event: message_delta data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"output_tokens":58}} event: message_stop data: {"type":"message_stop"} ``` Tool calls stream as a `content_block_start` with `"type": "tool_use"` followed by `input_json_delta` deltas. With the Python SDK, `client.messages.stream(...)` handles the event parsing. More on disconnects and timeouts in [streaming](/docs/streaming). ## Error shapes on /v1/messages Every error on `/v1/messages` and `/v1/messages/count_tokens` uses Anthropic's error format, whether it came from Tokens (bad key, no balance, plan limits) or from the upstream provider: ```json { "type": "error", "error": { "type": "billing_error", "message": "Prepaid wallet balance is insufficient for this request. Top up your wallet in the dashboard: https://tokens.bd/dashboard/billing", "code": "insufficient_credits" }, "request_id": "8f0c7a4e-2b1d-4c55-9a51-3f7e2d9b6c10" } ``` `error.type` follows Anthropic's list (`invalid_request_error`, `authentication_error`, `billing_error`, `permission_error`, `not_found_error`, `request_too_large`, `rate_limit_error`, `api_error`, `overloaded_error`), so the SDKs and Claude Code raise the right exception and retry where they should. Errors from Tokens also carry `error.code`, which tells you exactly which limit was hit; see [errors](/docs/errors). Upstream error messages are replaced with a generic one so internal provider details don't leak. Read `x-tokens-request-id` from the response headers when you contact support. ## Count tokens `POST /v1/messages/count_tokens` takes the same body as `/v1/messages` (without `max_tokens`) and returns `{"input_tokens": N}`. It is never billed and never runs the model, but it is not free of limits: it counts toward your per-minute request limit, takes a concurrency slot, and needs a funded account (a used-up usage window or a key's allowed-models list blocks it like any other request). See [Token counting](/docs/token-counting). When a provider that serves the model counts tokens natively, you get its exact count. Otherwise Tokens answers with an estimate and adds the response header `x-tokens-estimated: true`. Use it for rough budgeting; the `usage` object in each response is what you are billed on. ## anthropic-beta The `anthropic-beta` header is forwarded to providers that speak the Messages API, so beta features (newer tool types, extended context and so on) work where the provider supports them. On models served through translation the header has no effect. ## List models in Anthropic format `GET /v1/models` answers in Anthropic's format when the request carries `anthropic-version` (the Anthropic SDKs send it), so `client.models.list()` works: ```json { "data": [ { "type": "model", "id": "deepseek/deepseek-v4.1-flash", "display_name": "DeepSeek V4.1 Flash", "created_at": "2026-09-01T00:00:00.000Z" } ], "has_more": false, "first_id": "deepseek/deepseek-v4.1-flash", "last_id": "deepseek/deepseek-v4.1-flash" } ``` --- # Responses API > POST /v1/responses, the OpenAI Responses API: when to use it instead of chat completions, request and response examples, max_output_tokens and streaming events. Section: API Reference. Page: https://tokens.bd/docs/responses `POST https://tokens.bd/v1/responses` passes through the OpenAI Responses API. It exists mainly for clients built on it, Codex CLI being the obvious one. If you are writing new code and have no reason to prefer it, chat completions is the more widely supported choice across upstream models. ## When to use the Responses API | Use `/v1/responses` when | Use `/v1/chat/completions` when | | ------------------------------------------------------------------------- | ----------------------------------------------- | | Your tool only speaks Responses (Codex CLI with `wire_api = "responses"`) | You want the broadest model compatibility | | You already have code written against `client.responses.create` | You use frameworks that expect chat completions | | You want typed output items and event names | You need `n` > 1 or other chat-only fields | Support for this endpoint depends on the upstream that serves the model. The gateway forwards the request as is; if the provider behind a model doesn't implement Responses, the call fails with a 400 or 404 from upstream. When that happens, use chat completions for that model. ## Responses API request example :::code-tabs ```bash title="cURL" curl https://tokens.bd/v1/responses \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "instructions": "You are a concise senior engineer.", "input": "Give me one reason to pin dependency versions.", "max_output_tokens": 200 }' ``` ```python title="Python" import os from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) resp = client.responses.create( model="deepseek/deepseek-v4.1-flash", instructions="You are a concise senior engineer.", input="Give me one reason to pin dependency versions.", max_output_tokens=200, ) print(resp.output_text) print(resp.usage) ``` ::: A trimmed response: ```json { "id": "resp_abc123", "object": "response", "status": "completed", "model": "deepseek/deepseek-v4.1-flash", "output": [ { "type": "message", "role": "assistant", "content": [ { "type": "output_text", "text": "Reproducible builds: the same commit installs the same code everywhere." } ] } ], "usage": { "input_tokens": 24, "output_tokens": 15, "total_tokens": 39 } } ``` Fields commonly used in the request: | Field | Notes | | ---------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `model` | Required. Any id from `GET /v1/models`. | | `input` | A string, or an array of input items (messages, tool outputs). | | `instructions` | System-style instructions. | | `max_output_tokens` | Output limit. The gateway uses it for the cost reservation and, on a low balance, may lower it (minimum 16). | | `tools`, `tool_choice` | Function tools, where the model supports them. Hosted tools such as web search or file search are provider features; don't assume they work through the gateway. | | `stream` | `true` for Server-Sent Events. | ## Conversation state Tokens doesn't store prompts or responses, so do not count on `previous_response_id` or `store: true` to keep state between calls. Whether they work depends on the upstream, and a retry or failover can land on a different source that has never seen the earlier response. Send the full conversation in `input` each time; that works everywhere. ## Stream Responses API events With `"stream": true`, the response is Server-Sent Events with typed events. The ones most clients care about: | Event | Contains | | ---------------------------------------- | -------------------------------------------- | | `response.created` | The response object, status `in_progress` | | `response.output_text.delta` | A text fragment in `delta` | | `response.function_call_arguments.delta` | A fragment of tool-call arguments | | `response.completed` | The final response object, including `usage` | Unlike chat completions, you don't need `stream_options` to get usage: it arrives in `response.completed`. ```python stream = client.responses.create( model="deepseek/deepseek-v4.1-flash", input="Write a haiku about cache invalidation.", stream=True, ) for event in stream: if event.type == "response.output_text.delta": print(event.delta, end="", flush=True) elif event.type == "response.completed": print("\n", event.response.usage) ``` Disconnect handling and timeouts work the same as for the other endpoints; see [streaming](/docs/streaming). ## Codex CLI uses /v1/responses Codex CLI talks to Tokens through this endpoint. The provider block in `~/.codex/config.toml` looks like this, with the key read from `TOKENS_API_KEY`: ```toml title="~/.codex/config.toml" model = "deepseek/deepseek-v4.1-flash" model_provider = "tokens" [model_providers.tokens] name = "Tokens" base_url = "https://tokens.bd/v1" env_key = "TOKENS_API_KEY" wire_api = "responses" ``` The Tokens CLI can write this for you; the [quickstart](/docs/quickstart) has the one-line setup. If Codex reports errors for a specific model, check whether that model works on `/v1/responses` with the curl example above before digging into Codex settings. Errors follow the shape described in [errors](/docs/errors), and limits are in [rate limits](/docs/rate-limits). --- # Models and Usage Endpoints > GET /v1/models lists the models your key can call. GET /v1/tokens/usage returns your plan, usage windows, wallet balance and key limits. Plus notes on embeddings and legacy completions. Section: API Reference. Page: https://tokens.bd/docs/models-and-usage Two read-only endpoints let scripts and tools answer "which models can I call?" and "how much do I have left?" without a browser session. Both use the same API key as inference and neither is billed. ## List models with GET /v1/models ```bash curl https://tokens.bd/v1/models \ -H "Authorization: Bearer $TOKENS_API_KEY" ``` ```json { "object": "list", "data": [ { "id": "deepseek/deepseek-v4.1-flash", "object": "model", "created": 1788000000, "owned_by": "tokens", "permission": [], "root": "deepseek/deepseek-v4.1-flash", "parent": null } ] } ``` | Field | Meaning | | ------------------------------ | ------------------------------------------------------------------ | | `id` | The exact string to send as `model`. Always `provider/model` form. | | `object` | Always `"model"`. | | `created` | Unix timestamp of when the model was added to the catalog. | | `owned_by` | Always `"tokens"`, whichever lab built the model. | | `root`, `parent`, `permission` | Present for OpenAI SDK compatibility; `root` equals `id`. | The response does not include prices or context windows. Those are on [the model catalog](/models), one page per model. The catalog has no capability flags (vision, tool calling, reasoning): see [Model catalog](/docs/model-catalog) for how to find out what a model supports. ### Why a model is missing from the list The list is filtered for the key that calls it, so two keys on the same account can see different lists: 1. **Key allow-list.** If the key was created with an allowed-models list, only those models appear. 2. **Plan and wallet.** Each model is enabled for certain plan tiers and for pay-as-you-go. A model appears if your active plan's tier includes it, or if pay-as-you-go is allowed for it and your wallet balance is above zero. 3. **Catalog status.** Only active catalog models are listed. An empty `data` array usually means no active plan and no wallet balance. Subscribe or top up in [billing](/dashboard/billing); details in [plans and wallet](/docs/plans-and-wallet). ## Check remaining usage with GET /v1/tokens/usage This endpoint is specific to Tokens. It returns the current plan, plan usage windows, wallet balance and the calling key's limits. The `tokens.mjs usage` CLI command reads it, and it's handy in a status bar or a pre-flight check in a long-running agent. ```bash curl https://tokens.bd/v1/tokens/usage \ -H "Authorization: Bearer $TOKENS_API_KEY" ``` ```json { "object": "tokens.usage", "plan": { "name": "Pro Monthly", "tier": "monthly", "periodEnd": "2026-11-02T08:15:00.000Z" }, "windows": [ { "type": "session_5h", "label": "5-Hour Session", "unit": "usd", "limit": 5, "used": 1.284, "remaining": 3.716, "percentUsed": 26, "resetsAt": "2026-10-03T14:40:00.000Z" } ], "wallet": { "balanceUsd": 12.5 }, "key": { "monthlySpendCapUsd": 20, "allowedModels": null } } ``` The plan name and numbers above are illustrative; yours depend on your plan. | Field | Type | Meaning | | -------------------------------------- | ---------------- | ---------------------------------------------------------------------------- | | `plan` | object or null | Active subscription, or `null` on pay-as-you-go only | | `plan.tier` | string | Plan tier, such as `weekly` or `monthly` | | `plan.periodEnd` | ISO 8601 or null | When the current subscription period ends | | `windows[]` | array | Plan usage windows that have usage in their current period | | `windows[].type` | string | `session_5h` (rolling 5 hours), `weekly` or `monthly` | | `windows[].unit` | string | `usd` (credits, expressed in USD) or `requests` | | `windows[].limit`, `used`, `remaining` | number | In the window's unit | | `windows[].percentUsed` | integer | 0 to 100 | | `windows[].resetsAt` | ISO 8601 | When the window resets | | `wallet` | object or null | `balanceUsd`: prepaid wallet balance in USD, or `null` if there is no wallet | | `key.monthlySpendCapUsd` | number or null | The calling key's monthly cap, `null` if uncapped | | `key.allowedModels` | array or null | The key's allow-list, `null` if any model is allowed | :::note A window that hasn't been used yet in its current period is left out of `windows`, so an empty array on a fresh plan or right after a reset is normal. ::: When a window is used up, inference requests return 429 `window_exhausted` until `resetsAt`. The dashboard can also notify you at 50, 75, 90 and 100 percent of a window; see [usage and alerts](/docs/usage-and-alerts). The response is sent with `Cache-Control: no-store`. Polling it once a minute is plenty; it counts toward neither your bill nor your per-minute request limit. ## Embeddings: POST /v1/embeddings The route exists and follows OpenAI's embeddings format, but it only works for catalog models that are embedding models. Chat models will fail on it. Check [the model catalog](/models) for an embedding model before building on this endpoint; if none is listed, there is no embedding model available for your account yet. ## Legacy completions: POST /v1/completions The old prompt-in, text-out completions endpoint is passed through for tools that still use it. It only works when the upstream behind a model supports it, and many chat models don't. Use [chat completions](/docs/chat-completions) for anything new. ## Endpoints that don't exist Any other path under `/v1` returns 404 with code `unsupported_endpoint`. That includes images, audio, files, batches, assistants, fine-tuning and moderations. See [errors](/docs/errors) for the full code list. --- # Streaming > How Server-Sent Events streaming works through the Tokens gateway for chat completions and messages: the event format, usage chunks, timeouts, and handling disconnects. Section: API Reference. Page: https://tokens.bd/docs/streaming Set `"stream": true` on chat completions, messages, responses or legacy completions and the gateway streams Server-Sent Events (SSE) from the upstream provider to you as they arrive. The gateway doesn't buffer or rewrite the content; it reads token counts off the stream for billing and passes the bytes through. ## SSE format for chat completions Each event is a `data:` line with a JSON chunk, separated by a blank line. The stream ends with `data: [DONE]`. ```text data: {"id":"chatcmpl-a1b2","object":"chat.completion.chunk","created":1790000000,"model":"deepseek/deepseek-v4.1-flash","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]} data: {"id":"chatcmpl-a1b2","object":"chat.completion.chunk","created":1790000000,"model":"deepseek/deepseek-v4.1-flash","choices":[{"index":0,"delta":{"content":"Retry"},"finish_reason":null}]} data: {"id":"chatcmpl-a1b2","object":"chat.completion.chunk","created":1790000000,"model":"deepseek/deepseek-v4.1-flash","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]} data: {"id":"chatcmpl-a1b2","object":"chat.completion.chunk","created":1790000000,"model":"deepseek/deepseek-v4.1-flash","choices":[],"usage":{"prompt_tokens":18,"completion_tokens":42,"total_tokens":60}} data: [DONE] ``` Reasoning models may also send `delta.reasoning_content`, and tool calls arrive as `delta.tool_calls` fragments (see [tool calling](/docs/tool-calling)). ## Get token usage while streaming with include_usage The chunk with `"choices": []` and a `usage` object only appears if you ask for it: ```json { "stream": true, "stream_options": { "include_usage": true } } ``` Without it, you get no usage chunk, matching OpenAI's behavior. Billing doesn't depend on this flag: the gateway always meters the stream. The flag only controls whether you see the numbers. If you set it, guard against the empty `choices` array in your loop. Code that does `chunk.choices[0]` unconditionally will crash on the last chunk. ## SSE format for Anthropic messages `/v1/messages` streams named events: `message_start`, `content_block_start`, `content_block_delta`, `content_block_stop`, `message_delta` and `message_stop`. Input token usage is in `message_start`, output usage in `message_delta`, so there is no flag to set. The full sequence is shown in [messages](/docs/messages). The Responses API has its own event names, listed in [responses](/docs/responses). ## Stream in Python and Node.js :::code-tabs ```python title="Python" import os from openai import OpenAI client = OpenAI( base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], timeout=600, # seconds; long reasoning answers can take minutes ) stream = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": "Explain exponential backoff with jitter."}], stream=True, stream_options={"include_usage": True}, ) finished = False for chunk in stream: if chunk.choices: choice = chunk.choices[0] if choice.delta.content: print(choice.delta.content, end="", flush=True) if choice.finish_reason: finished = True if chunk.usage: print("\n", chunk.usage) if not finished: print("\n[stream ended without a finish_reason: treat the answer as incomplete]") ``` ```typescript title="Node.js" import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY, timeout: 600_000, // ms }); const controller = new AbortController(); // Cancel after 2 minutes, or wire this to a user's "stop" button. const timer = setTimeout(() => controller.abort(), 120_000); const stream = await client.chat.completions.create( { model: "deepseek/deepseek-v4.1-flash", messages: [{ role: "user", content: "Explain exponential backoff with jitter." }], stream: true, stream_options: { include_usage: true }, }, { signal: controller.signal } ); try { for await (const chunk of stream) { const delta = chunk.choices[0]?.delta?.content; if (delta) process.stdout.write(delta); if (chunk.usage) console.log("\n", chunk.usage); } } finally { clearTimeout(timer); } ``` ::: For raw HTTP, `curl -N` disables output buffering so you can watch events arrive: ```bash curl -N https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"deepseek/deepseek-v4.1-flash","stream":true,"messages":[{"role":"user","content":"Count to five."}]}' ``` ## Errors before and during a stream Errors that happen before the first byte, such as a bad key, no balance, rate limits or an upstream that refused the request, come back as a normal JSON error with the matching HTTP status, not as an SSE event. Check the status code before you start parsing events. The [errors](/docs/errors) page lists every code. Once the stream has started, the status is already 200. If the upstream fails midway, the connection closes without `[DONE]` (or without `message_stop` on messages). Treat a stream that ends without a `finish_reason` or stop event as incomplete. The `x-tokens-request-id` response header arrives with the headers, before any content. Log it at the start of each request so you have it if the stream dies later. ## Timeouts for long requests Long requests are fine. The gateway waits up to 600 seconds for an upstream to start responding, which covers reasoning models that think for minutes before the first token, and up to 300 seconds between chunks once a stream is flowing. If the upstream doesn't start in time, you get 504 `upstream_timeout`. When a model has a backup provider behind it, the gateway moves to the backup after about 30 seconds without a first token, so the full wait only applies when the last provider is the slow one. You never see the switch. Your client's timeout needs to be at least as generous. Set it explicitly, as in the examples above, rather than relying on library defaults or a proxy in front of your app that cuts idle connections after 30 or 60 seconds. ## Handle client disconnects If your client disconnects or aborts mid-stream, the gateway cancels the upstream request and bills for what was generated up to that point: the input plus the output already streamed. Nothing further is charged. Before retrying a broken stream, remember that a retry sends and bills the full prompt again. For long agent prompts, that adds up. Retry streams that failed with a retryable error before any content arrived; for streams that broke partway, decide whether the partial output is usable first. Retry rules by error code are in [errors](/docs/errors). Each open stream counts toward your account's concurrency limit until it finishes. Streams you abandon without closing hold a slot, so close or abort them explicitly. See [rate limits](/docs/rate-limits). --- # Tool Calling > Use OpenAI-style function tools and Anthropic-style tools through the Tokens gateway, with a complete runnable tool loop in Python and tips on model support. Section: API Reference. Page: https://tokens.bd/docs/tool-calling Tool calling (function calling) lets a model ask your code to run a function and then use the result. Tokens carries tool definitions and tool calls between you and the provider on chat completions and messages, so the same code you would write against OpenAI or Anthropic works here. Whether it works well depends on the model. ## How tool calling works 1. You send the conversation plus a list of `tools`, each with a name, description and JSON Schema for its arguments. 2. The model either answers in text or returns one or more tool calls with arguments. 3. Your code runs each function and sends the results back as `tool` messages. 4. Repeat until the model answers without calling a tool. The gateway never executes tools. It only carries the messages. ## OpenAI-style tools on /v1/chat/completions A tool definition: ```json { "type": "function", "function": { "name": "get_current_time", "description": "Get the current time in an IANA timezone, e.g. Asia/Dhaka.", "parameters": { "type": "object", "properties": { "timezone": { "type": "string", "description": "IANA timezone name" } }, "required": ["timezone"] } } } ``` When the model calls it, the assistant message has `tool_calls` instead of (or alongside) `content`, and `finish_reason` is `"tool_calls"`: ```json { "role": "assistant", "content": null, "tool_calls": [ { "id": "call_01", "type": "function", "function": { "name": "get_current_time", "arguments": "{\"timezone\": \"Asia/Dhaka\"}" } } ] } ``` `arguments` is a JSON string, not an object. Parse it, and expect that a model can occasionally produce invalid JSON. ## Complete tool loop in Python This runs as is with `pip install openai` and `TOKENS_API_KEY` set (on Windows, also `pip install tzdata` for timezone data). The tools use only the standard library, so there is nothing else to configure. ```python title="tool_loop.py" import json import os from datetime import datetime from zoneinfo import ZoneInfo from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) MODEL = "deepseek/deepseek-v4.1-flash" def get_current_time(timezone: str) -> dict: return {"timezone": timezone, "time": datetime.now(ZoneInfo(timezone)).isoformat()} def add(a: float, b: float) -> dict: return {"result": a + b} FUNCTIONS = {"get_current_time": get_current_time, "add": add} TOOLS = [ { "type": "function", "function": { "name": "get_current_time", "description": "Get the current time in an IANA timezone, e.g. Asia/Dhaka.", "parameters": { "type": "object", "properties": {"timezone": {"type": "string"}}, "required": ["timezone"], }, }, }, { "type": "function", "function": { "name": "add", "description": "Add two numbers.", "parameters": { "type": "object", "properties": {"a": {"type": "number"}, "b": {"type": "number"}}, "required": ["a", "b"], }, }, }, ] messages = [ {"role": "user", "content": "What time is it in Dhaka and in London? Also, what is 1250.5 + 349.5?"} ] for _ in range(8): # hard stop so a confused model can't loop forever resp = client.chat.completions.create( model=MODEL, messages=messages, tools=TOOLS, tool_choice="auto", max_tokens=1024 ) msg = resp.choices[0].message messages.append(msg.model_dump(exclude_none=True)) if not msg.tool_calls: print(msg.content) break for call in msg.tool_calls: fn = FUNCTIONS.get(call.function.name) try: args = json.loads(call.function.arguments or "{}") result = fn(**args) if fn else {"error": f"unknown tool {call.function.name}"} except Exception as exc: # report errors back to the model instead of crashing result = {"error": str(exc)} messages.append( {"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)} ) else: print("Stopped after 8 rounds without a final answer.") ``` A few details that save debugging time: - Append the assistant message that contains `tool_calls` before the `tool` results. Most providers reject a `tool` message whose `tool_call_id` doesn't match a preceding call. - Models can return several tool calls in one turn. Answer all of them before the next request. - Errors go back to the model as tool results. It can often recover, for example by fixing a bad timezone name. - Every round is a separate billed request that resends the whole conversation, so long loops cost more than the final answer suggests. ## Control tool use with tool_choice | Value | Effect | | --------------------------------------------------- | -------------------------------------------------- | | `"auto"` | Model decides (the default when tools are present) | | `"none"` | Model must answer in text | | `"required"` | Model must call at least one tool | | `{"type": "function", "function": {"name": "add"}}` | Model must call that function | Support for `"required"` and forced functions varies by model and provider. Some ignore it; some return a 400. If you need a forced call, test it on the exact model first. ## Anthropic-style tools on /v1/messages On `/v1/messages`, use Anthropic's format: tools have `name`, `description` and `input_schema`, the model replies with `tool_use` content blocks, and you answer with `tool_result` blocks in a user message. ```python import os import anthropic client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"]) tools = [{ "name": "add", "description": "Add two numbers.", "input_schema": { "type": "object", "properties": {"a": {"type": "number"}, "b": {"type": "number"}}, "required": ["a", "b"], }, }] messages = [{"role": "user", "content": "What is 1250.5 + 349.5?"}] resp = client.messages.create( model="deepseek/deepseek-v4.1-flash", max_tokens=1024, tools=tools, messages=messages ) if resp.stop_reason == "tool_use": call = next(b for b in resp.content if b.type == "tool_use") total = call.input["a"] + call.input["b"] messages += [ {"role": "assistant", "content": resp.content}, {"role": "user", "content": [ {"type": "tool_result", "tool_use_id": call.id, "content": str(total)} ]}, ] resp = client.messages.create( model="deepseek/deepseek-v4.1-flash", max_tokens=1024, tools=tools, messages=messages ) print(resp.content[0].text) ``` `tool_choice` takes Anthropic's forms here: `{"type": "auto"}`, `{"type": "any"}` or `{"type": "tool", "name": "add"}`. When a non-Claude model is served through `/v1/messages`, the gateway translates tools and tool calls to and from the OpenAI format; see [messages](/docs/messages) for what survives translation. The `anthropic-beta` header is forwarded to providers that speak the Messages API natively, so server-side tools that need it work where that provider supports them. ## Tips for tool calling - **Not every model supports tools.** Check the model's page in [the catalog](/models) before building an agent on it. Small or older models may ignore tools or invent arguments. - **Keep schemas simple.** Flat objects with clear descriptions get more reliable arguments than deeply nested schemas. - **Streaming works.** Tool call arguments arrive as fragments in `delta.tool_calls` (or `input_json_delta` on messages); concatenate them before parsing. See [streaming](/docs/streaming). - **Validate before executing.** Treat arguments as untrusted input, especially for tools that touch files, shells or money. If tool calls fail with a 400 on one model but work on another, the model or its provider doesn't accept that tool feature. The [troubleshooting](/docs/troubleshooting) page covers other common failures. --- # Errors > Every error code the Tokens API returns, what it means and what to do, plus the error JSON shape, request ids for support tickets, and which errors to retry. Section: API Reference. Page: https://tokens.bd/docs/errors This page lists every API error code the Tokens gateway returns, what each one means, and whether to retry. Branch on the `code` field in your code; messages are for humans and can change. ## Error JSON shape Errors raised by the gateway on OpenAI-style endpoints (`/v1/chat/completions`, `/v1/completions`, `/v1/responses`, `/v1/embeddings`, `/v1/models`) look like OpenAI's: ```json { "error": { "message": "Prepaid wallet balance is insufficient for this request. Top up your wallet in the dashboard: https://tokens.bd/dashboard/billing", "type": "insufficient_quota", "code": "insufficient_credits", "param": null, "request_id": "8f0c7a4e-2b1d-4c55-9a51-3f7e2d9b6c10" } } ``` | Field | Notes | | ------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `code` | Stable, machine-readable. Use this. | | `type` | Broad class: `authentication_error`, `permission_denied_error`, `insufficient_quota`, `rate_limit_error`, `invalid_request_error`, `api_error`, `server_error` or `tokens_error`. | | `message` | Human-readable. Billing errors include a link to [billing](/dashboard/billing). | | `request_id` | Same value as the `x-tokens-request-id` header. Also in the `x-tokens-request-id` header on every response. | On `/v1/messages` and `/v1/messages/count_tokens`, every error uses Anthropic's shape instead, whether Tokens or the upstream provider raised it: `{"type": "error", "error": {"type": "...", "message": "...", "code": "..."}, "request_id": "..."}`. Tokens adds `code` to Anthropic's shape, so you can still branch on the codes below. Upstream error messages are replaced with a generic message so provider internals don't leak; the HTTP status and `code` still tell you the category. ## Error code reference | Status | Code | Meaning | What to do | | ------- | ------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ | | 400 | `invalid_request` | Missing `model` (or no `Content-Type: application/json`), `n` outside 1 to 4, or the upstream rejected the request body | Fix the request. For upstream rejections, check the model supports the parameters you sent | | 400 | `model_not_available` | The model is in the catalog but can't be served right now | Pick another model from `GET /v1/models` | | 400 | `endpoint_not_supported_for_model` | None of the providers serving this model can answer this endpoint (for example `/v1/responses`, `/v1/completions` or embeddings on a chat model) | Use `/v1/chat/completions`, or pick another model | | 401 | `missing_api_key` | No `Authorization: Bearer` or `x-api-key` header | Set `TOKENS_API_KEY` in the environment that runs the tool | | 401 | `invalid_api_key` | Key not recognized | Recopy the key or create a new one | | 401/403 | `upstream_auth_error` | The upstream provider refused the gateway's credentials. Not your key | Retry later or switch model; report it with the request id | | 402 | `insufficient_credits` | Plan credits and wallet can't cover the request | Top up or renew in [billing](/dashboard/billing), or lower `max_tokens` | | 402 | `no_funding` | No active plan and no wallet balance | Subscribe or add funds | | 402 | `outstanding_debt` | Earlier usage left a negative balance | Top up to clear it | | 402 | `member_cap_reached` | Team accounts (if enabled): your member monthly cap is reached | Ask your team admin | | 403 | `key_inactive` | Key revoked or rotated | Use the current secret | | 403 | `key_expired` | Key past its expiry date | Create a new key | | 403 | `account_suspended` | Account suspended | Contact [support](/docs/support) | | 403 | `model_not_allowed_on_key` | Model not in this key's allowed list | Use an allowed model or another key | | 403 | `monthly_spend_cap_exceeded` | Key's monthly spend cap reached (the check includes this request's worst-case cost) | Lower `max_tokens`, wait for the new month, or use another key | | 403 | `tier_permission_denied` | Your plan doesn't include this model and you have no wallet balance | Upgrade, or add wallet funds for pay-as-you-go | | 404 | `model_not_found` | Model id unknown or inactive (also returned when the upstream doesn't know the model) | Check the exact id with `GET /v1/models` | | 404 | `unsupported_endpoint` | Path or method isn't one of the supported endpoints | See [models and usage](/docs/models-and-usage) for the endpoint list | | 404 | `anthropic_protocol_disabled` | The Messages API (`/v1/messages`) is switched off on this deployment | Use Chat Completions instead | | 413 | `request_entity_too_large` | Body over 10 MB | Trim context or attachments | | 429 | `rate_limited` | Requests-per-minute limit reached | Wait `Retry-After` seconds | | 429 | `concurrency_limit` | Too many requests in flight on your account | Wait `Retry-After` (2 s) or reduce parallelism | | 429 | `window_exhausted` | A plan usage window (5-hour, weekly or monthly) is used up | Wait for the reset (`Retry-After`) or upgrade the plan | | 429 | `model_limit_reached` | This model's allowance on your plan is used up for the billing period; your other models still work | Switch models, or wait for the reset (`Retry-After`, shown as a date in the message) | | 429 | `rate_limit_exceeded` | The upstream provider rate limited us after failover was exhausted | Retry with backoff; honor `Retry-After` if present | | 500 | `lookup_failed`, `admission_error`, `catalog_error`, `usage_unavailable` | Internal error on our side | Retry once or twice with backoff, then contact support | | 502 | `upstream_unreachable` | Couldn't connect to any upstream for this model | Retry with backoff; check [status](/status) | | 5xx | `upstream_error` | The upstream returned a server error (status passed through) | Retry with backoff | | 503 | `no_upstream_available` | No upstream is configured for this model right now | Try another model; check [status](/status) | | 503 | `model_not_priced` | The model has no price configured, so it can't be billed | Try another model and report it | | 504 | `upstream_timeout` | Upstream didn't respond within 600 s, or the connection broke after the request was sent | Retry once; consider a smaller request | The gateway already fails over to another upstream source on 429, 502, 503, 504 and connection errors before returning anything to you. An upstream error you see means every configured source for that model failed or the last one did. ## Request ids and support tickets Every response, success or error, carries two headers: | Header | Value | | --------------------- | ------------------------------------------------------------------------------------ | | `x-tokens-request-id` | The gateway's id for this request. Always generated by us. | | `x-request-id` | Your own `x-request-id` if you sent one, otherwise the same as `x-tokens-request-id` | Sending your own `x-request-id` lets you correlate our id with your logs. To read the header with the OpenAI Python SDK: ```python raw = client.chat.completions.with_raw_response.create( model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": "ping"}], ) print(raw.headers.get("x-tokens-request-id")) completion = raw.parse() ``` With curl, add `-i` to print headers. When you open a ticket in [support](/dashboard/support), include the `x-tokens-request-id`, the time (with timezone), the endpoint, the model and the status code. Never include the API key. We store usage metadata by request id, not prompt content, so the id is what lets us find your request. More in [support](/docs/support). ## Retry guidance | Retry | Codes | | ---------------------------------------- | ---------------------------------------------------------------------- | | Yes, after `Retry-After` | `rate_limited`, `concurrency_limit`, `rate_limit_exceeded` | | Yes, with exponential backoff | `upstream_unreachable`, `upstream_error`, `upstream_timeout`, all 500s | | Only after the reset time, not in a loop | `window_exhausted`, `model_limit_reached` (`Retry-After` can be days) | | No, fix something first | All 400, 401, 402, 403, 404 and 413 errors | Backoff that works in practice: start around 1 second, double each attempt, add random jitter, cap at 30 seconds, and stop after 4 or 5 attempts. If `Retry-After` is present, wait at least that long. A retry is a new request and is billed if it succeeds, and the OpenAI and Anthropic SDKs already retry some of these errors on their own, so check `max_retries` before stacking your own loop on top. A full backoff example is in [rate limits](/docs/rate-limits), and agent-specific fixes are in [troubleshooting](/docs/troubleshooting). --- # Rate Limits > Requests per minute, concurrency, plan usage windows and per-key monthly spend caps: how each limit works, the errors they return, and how to back off correctly. Section: API Reference. Page: https://tokens.bd/docs/rate-limits Four separate limits can stop a request before it reaches a model: requests per minute, concurrent requests, plan usage windows, and a key's monthly spend cap. Upstream providers have their own rate limits on top. This page explains each one, the error it returns, and how a client should react. ## Rate limits at a glance | Limit | Scope | Default | Error | `Retry-After` | | ----------------------- | --------------------------- | --------------------------------------- | -------------------------------- | --------------------------------------- | | Requests per minute | Account (all keys combined) | 60 RPM, or your plan's value | 429 `rate_limited` | Seconds until the next minute | | Concurrent requests | Account | Plan's limit; 10 with a plan, 3 without | 429 `concurrency_limit` | 2 | | Usage windows | Subscription | Set by the plan | 429 `window_exhausted` | Seconds until the window resets | | Monthly spend cap | One key | None unless set at key creation | 403 `monthly_spend_cap_exceeded` | Not sent | | Upstream provider limit | Provider | Set by the provider | 429 `rate_limit_exceeded` | Passed through if the provider sent one | Your plan's exact numbers are shown in [billing](/dashboard/billing) and explained in [plans and wallet](/docs/plans-and-wallet). ## Requests per minute The per-minute limit counts requests per account in fixed one-minute buckets that start on the clock minute. Every key on the account shares the same bucket, so creating more keys doesn't raise it. The dashboard playground has its own separate limit of 10 RPM and doesn't use up your API allowance. When you go over, you get 429 `rate_limited` with `Retry-After` set to the seconds remaining in the current minute, so the wait is never more than 60 seconds. The error message mentions "this key", but the limit is per account. Requests rejected for insufficient funds (402 `insufficient_credits` or `no_funding`) or an exhausted usage window still count toward the minute. A client that retries a 402 in a tight loop will also hit the per-minute limit. Don't retry 402s at all. `GET /v1/models` and `GET /v1/tokens/usage` don't count. ## Concurrency limit Concurrency is the number of requests in flight at once on your account. A streaming request occupies a slot from admission until the stream finishes, so a coding agent that opens several streams in parallel, or a script that fans out with `asyncio.gather`, can hit this before the per-minute limit. The 429 `concurrency_limit` response carries `Retry-After: 2`. The fix is usually to cap parallelism on your side, for example with a semaphore sized below your plan's limit, rather than retrying harder. Close or abort streams you no longer need; an abandoned stream still holds its slot until it ends. ## Usage windows Subscription plans can define usage windows, each limited in credits (expressed in USD) or in requests: | Window | `type` in the API | Resets | | -------------- | ----------------- | ------------------------------------------------ | | 5-hour session | `session_5h` | 5 hours after the request that opened the window | | Weekly | `weekly` | Mondays at 00:00 UTC | | Monthly | `monthly` | At the end of the subscription period | When any window is used up, requests return 429 `window_exhausted` and `Retry-After` is the number of seconds until that window resets. That can be hours, so don't retry automatically: show the reset time to the user, or stop the job. Check the remaining amounts before starting a long agent run: ```bash curl -s https://tokens.bd/v1/tokens/usage -H "Authorization: Bearer $TOKENS_API_KEY" ``` The response shape is documented in [models and usage](/docs/models-and-usage). Alerts at 50, 75, 90 and 100 percent are covered in [usage and alerts](/docs/usage-and-alerts). ## Monthly spend caps per key A key can have a monthly spend cap in USD, set when you create it. Spend is counted per calendar month (UTC). Before each request the gateway checks the key's spend so far plus the worst-case cost of the new request, based on its `max_tokens` (or 8,192 output tokens if unset). Near the cap, a request with a large `max_tokens` can be refused while a smaller one would pass. The error is 403 `monthly_spend_cap_exceeded`, not a 429, because waiting a few seconds won't help. Caps can't be edited after creation; if you need a higher one, create a new key. If you only change one setting on a key that goes into a shared CI system or an agent you don't watch, make it the spend cap. More in [API keys](/docs/api-keys). ## Retry-After and rate limit headers There are no `X-RateLimit-Limit` or `X-RateLimit-Remaining` headers. The only rate-limit header is `Retry-After`, in seconds, on 429 responses. To see remaining budget ahead of time, poll `GET /v1/tokens/usage`. ## Client-side backoff example The OpenAI and Anthropic SDKs retry some 429 and 5xx errors on their own (`max_retries`, 2 by default). The example below turns that off and handles it explicitly, so `window_exhausted` isn't retried and `Retry-After` is honored. :::code-tabs ```python title="Python" import os import random import time import openai from openai import OpenAI client = OpenAI( base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], max_retries=0, ) RETRYABLE = {429, 500, 502, 503, 504} def create_with_backoff(max_attempts: int = 5, **kwargs): delay = 1.0 for attempt in range(1, max_attempts + 1): try: return client.chat.completions.create(**kwargs) except openai.APIStatusError as e: if ( e.status_code not in RETRYABLE or e.code == "window_exhausted" or attempt == max_attempts ): raise try: wait = float(e.response.headers.get("retry-after", delay)) except ValueError: wait = delay request_id = e.response.headers.get("x-tokens-request-id") print(f"{e.status_code} {e.code} (request {request_id}), retrying in {wait:.1f}s") except openai.APIConnectionError: if attempt == max_attempts: raise wait = delay time.sleep(min(wait, 60) + random.uniform(0, 0.5 * delay)) delay = min(delay * 2, 30) resp = create_with_backoff( model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": "Say hello."}], max_tokens=50, ) print(resp.choices[0].message.content) ``` ```typescript title="Node.js" import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY, maxRetries: 0, }); const RETRYABLE = new Set([429, 500, 502, 503, 504]); const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms)); function retryAfterMs(err: InstanceType): number | undefined { const h: unknown = err.headers; // a Headers object or a plain record, depending on SDK version const value = h instanceof Headers ? h.get("retry-after") : (h as Record | undefined)?.["retry-after"]; return value ? Number(value) * 1000 : undefined; } async function createWithBackoff( body: OpenAI.Chat.ChatCompletionCreateParamsNonStreaming, maxAttempts = 5 ) { let delay = 1000; for (let attempt = 1; ; attempt++) { try { return await client.chat.completions.create(body); } catch (err) { if (!(err instanceof OpenAI.APIError) || attempt >= maxAttempts) throw err; const connectionError = err instanceof OpenAI.APIConnectionError; if ( !connectionError && (!RETRYABLE.has(err.status ?? 0) || err.code === "window_exhausted") ) { throw err; } const wait = (!connectionError && retryAfterMs(err)) || delay; await sleep(Math.min(wait, 60_000) + Math.random() * 0.5 * delay); delay = Math.min(delay * 2, 30_000); } } } const resp = await createWithBackoff({ model: "deepseek/deepseek-v4.1-flash", messages: [{ role: "user", content: "Say hello." }], max_tokens: 50, }); console.log(resp.choices[0].message.content); ``` ::: Every successful retry is a billed request, so keep the attempt count low. The full list of retryable and non-retryable codes is in [errors](/docs/errors). --- # Embeddings > POST /v1/embeddings turns text into vectors for search, retrieval and clustering. How to find an embedding model in the catalog, the request and response, how it is billed, and the errors you can hit. Section: API Reference. Page: https://tokens.bd/docs/embeddings `POST https://tokens.bd/v1/embeddings` is the OpenAI-compatible embeddings endpoint. You send text and get back one vector per input, which you can store and compare for semantic search, retrieval for RAG, deduplication or clustering. The gateway checks your key, plan and limits, then forwards the body to the upstream provider for the model you named. This endpoint only works with **embedding models**. A chat model such as `deepseek/deepseek-v4.1-flash` is not one, and sending it here fails. Read the next section first. ## Find an embedding model Tokens does not add a capability flag to `GET /v1/models`, so the list has ids and nothing else. To find an embedding model: 1. Open [the model catalog](/models) and search for `embed`. The search matches a model's name, id and provider. 2. Open the model's page to check its context window and its input price per million tokens. 3. Confirm your key can call it: the id must appear in `GET /v1/models` for that key ([why a model can be missing](/docs/models-and-usage)). If no model in the catalog is an embedding model, none is available to your account yet and `/v1/embeddings` has nothing to serve. Keep your current embedding provider until one is listed. This page does not name a model because the catalog changes; the examples below read the id from an environment variable instead: ```bash export EMBEDDING_MODEL="the-id-from-the-catalog" ``` ## Create embeddings :::code-tabs ```bash title="cURL" curl https://tokens.bd/v1/embeddings \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d @- < item.embedding); console.log(vectors.length, "vectors of", vectors[0].length, "numbers"); console.log(resp.usage); ``` ::: The `Content-Type: application/json` header matters. Without it the gateway can't read the body and answers 400 `invalid_request` saying the request must specify a `model` field, even though it does. ## Request fields The body follows OpenAI's embeddings format. Checked against OpenAI's API reference in October 2026. | Field | Type | Notes | | ----------------- | ---------------- | -------------------------------------------------------------------------------------------------------------------------- | | `model` | string | Required. The id of an embedding model, exactly as the catalog shows it. | | `input` | string or array | Required. One string, or an array of strings to embed in one request. OpenAI's format also allows token ids. | | `encoding_format` | string | `"float"` (default in the raw API) or `"base64"`. Support is up to the model's provider. | | `dimensions` | integer | Shorter output vectors. In OpenAI's own API only some models accept it; whether yours does is decided by its provider. | | `user` | string | An end-user identifier, passed to the provider. | :::note[Parameters depend on the upstream model] The gateway reads `model` and nothing else from this body. The other fields go to the provider unchanged, so the limits that matter (maximum tokens per input, how many inputs per request, whether `dimensions` is allowed, the vector length) are the provider's, and they differ between models. For OpenAI's own embedding models, OpenAI documents 8,192 tokens per input, at most 2,048 items in an input array and 300,000 tokens across one request. Do not assume the same numbers for another model. ::: The request body can be up to 10 MB in total. Larger bodies return 413 `request_entity_too_large`. ### The Python SDK asks for base64 by default If you leave `encoding_format` out, the OpenAI Python SDK sends `"base64"` and decodes the result itself (checked in the SDK source, October 2026). That works only when the model's provider supports base64. If a call fails or returns odd vectors, set `encoding_format="float"` as the examples above do. Setting it explicitly in the Node.js SDK costs nothing either. ## Example response Vectors are shortened here. A real one has hundreds or thousands of numbers, so the response body is large compared with the request. ```json { "object": "list", "data": [ { "object": "embedding", "index": 0, "embedding": [0.0123, -0.0456, 0.0789] }, { "object": "embedding", "index": 1, "embedding": [0.0311, -0.0127, 0.0644] } ], "model": "the-id-from-the-catalog", "usage": { "prompt_tokens": 14, "total_tokens": 14 } } ``` `data` has one entry per input, in the same order; `index` is the position of the input. The body comes from the provider, so extra fields can appear and the vector length depends on the model. There is no `completion_tokens` because an embedding request produces no text. ## Compare two vectors Vectors from the same model can be compared with cosine similarity. A small example in Python: ```python import math def cosine(a, b): dot = sum(x * y for x, y in zip(a, b)) return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b))) print(cosine(vectors[0], vectors[1])) ``` Never compare or index together vectors from different models, or from the same model with different `dimensions`. Store the model id and `dimensions` next to each vector so you know when to re-embed. ## Billing and limits - **Metering.** Embeddings are billed on the `usage.prompt_tokens` the provider reports, at the model's input price per million tokens. There is no output side. Per-model prices are in [the model catalog](/models) and [pricing](/pricing). If a provider sends no usage, the gateway estimates from the request and response size. - **No `max_tokens`.** Unlike chat, there is no output reservation. Admission reserves the worst-case cost from the request size only, so a request is refused for funds only when your balance can't cover that estimate. You pay for actual usage. - **Rate limits.** Every embeddings request counts as one request toward your per-minute limit and holds a concurrency slot while it runs, whatever its size. Put many inputs into one `input` array instead of sending one request per text, within the provider's limit for the model. See [rate limits](/docs/rate-limits). - **Failed requests are not billed.** A response with a status of 400 or above from the provider is returned to you without a charge. - **Not streamed.** Embeddings have no `stream` mode. ## Using a chat model on this endpoint Sending a chat model here is the most common mistake. What you see depends on the model's provider, not on a fixed Tokens rule: | Response | Cause | | ------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------- | | 400 `invalid_request`, "rejected by the upstream provider" | The provider took the request and refused it because the model can't produce embeddings. | | 404 `model_not_found` | The provider doesn't have an embeddings route for that model. Also returned for an id that isn't in the catalog. | | 400 `endpoint_not_supported_for_model` | The model is served only through a provider that speaks the Anthropic Messages protocol, which has no embeddings. Nothing was sent. | | 502 `upstream_unreachable` | The model has several providers and each one failed or refused. | In every case switch to an embedding model from the catalog. The upstream message is replaced with a generic one, so use the `x-tokens-request-id` header when you contact support. The full list is in [errors](/docs/errors). ## Errors Gateway errors use the OpenAI error shape with a `code` you can branch on, and every response carries an `x-tokens-request-id` header. The ones you meet most often here: | Status | Code | What to do | | ------ | -------------------------------------- | -------------------------------------------------------------------------------- | | 400 | `invalid_request` | Check `model` and `input`, and that you sent `Content-Type: application/json`. | | 402 | `insufficient_credits`, `no_funding` | Top up or renew in [billing](/dashboard/billing). | | 403 | `model_not_allowed_on_key` | The key has an allow-list that excludes this model. Use another key. | | 404 | `model_not_found` | Check the id with `GET /v1/models`, or the model isn't an embedding model. | | 413 | `request_entity_too_large` | Send fewer inputs per request. | | 429 | `rate_limited`, `concurrency_limit` | Wait for `Retry-After`; batch inputs into fewer requests. | For a bulk indexing job, add retries with backoff as shown in [rate limits](/docs/rate-limits), and keep the number of parallel requests under your plan's concurrency limit. ## Related - [Chat completions](/docs/chat-completions) for text generation. - [Token counting](/docs/token-counting) for estimating input size before you embed a large corpus. - [Models and usage](/docs/models-and-usage) for `GET /v1/models` and `GET /v1/tokens/usage`. --- # Legacy Completions > POST /v1/completions is the old prompt-in, text-out endpoint. Request fields, streaming, billing, which models accept it, and how to move the same call to chat completions. Section: API Reference. Page: https://tokens.bd/docs/legacy-completions `POST https://tokens.bd/v1/completions` is the original OpenAI text completions endpoint: you send a `prompt` string and the model continues it. Tokens passes it through for older tools and scripts that still call it. For anything new, use [chat completions](/docs/chat-completions), which every chat model supports and which most coding agents expect. :::warning[Not every model serves this endpoint] Tokens forwards the request to the model's provider, and the provider decides whether the model can do plain text completion. Many chat models can't, and there is no list of the ones that can. Test your model with a small request (below) before you build on it. ::: ## Make a completions request :::code-tabs ```bash title="cURL" curl https://tokens.bd/v1/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "prompt": "A one-line definition of HTTP 429:", "max_tokens": 40, "temperature": 0 }' ``` ```python title="Python" import os from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) resp = client.completions.create( model="deepseek/deepseek-v4.1-flash", prompt="A one-line definition of HTTP 429:", max_tokens=40, temperature=0, ) print(resp.choices[0].text) print(resp.usage) ``` ```typescript title="Node.js" import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY }); const resp = await client.completions.create({ model: "deepseek/deepseek-v4.1-flash", prompt: "A one-line definition of HTTP 429:", max_tokens: 40, temperature: 0, }); console.log(resp.choices[0].text, resp.usage); ``` ::: If this returns an error, read [Which models work](#which-models-work). The `Content-Type: application/json` header matters: without it the gateway can't read the body and answers 400 `invalid_request` about a missing `model` field. ## Request fields The body follows OpenAI's completions format. Checked against OpenAI's API reference in October 2026. | Field | Type | Notes | | --------------------------------------------- | -------------------- | ---------------------------------------------------------------------------------------------------- | | `model` | string | Required. A catalog id. List yours with `GET /v1/models`. | | `prompt` | string or array | Required in OpenAI's format. A string, or an array of strings. Token-id arrays are in the format too. | | `max_tokens` | integer | Upper bound on generated tokens. Set it; see the note on reservations below. | | `temperature`, `top_p` | number | Sampling. Change one or the other, as OpenAI recommends. | | `n` | integer | Number of completions per prompt. The gateway accepts 1 to 4 and returns 400 `invalid_request` otherwise. | | `stop` | string or array | Up to 4 stop sequences in OpenAI's format. | | `stream` | boolean | `true` returns Server-Sent Events. | | `stream_options.include_usage` | boolean | With `stream: true`, adds a final chunk with `usage`. | | `suffix`, `echo`, `logprobs`, `best_of` | various | OpenAI's format has them. Whether a model honors them is up to its provider. | | `seed`, `presence_penalty`, `frequency_penalty`, `logit_bias`, `user` | various | Passed through as sent. | :::note[Parameters depend on the upstream model] Apart from `model` and `n`, the gateway doesn't validate or rewrite these fields. It passes them to the provider serving the model, which may ignore a field or reject the request. In OpenAI's own documentation, for example, `suffix` works only with one model. Check the model's page in [the catalog](/models) and test. ::: The request body can be up to 10 MB. Larger bodies return 413 `request_entity_too_large`. ## Example response Ids and numbers are illustrative. ```json { "id": "cmpl-a1b2c3", "object": "text_completion", "created": 1790000000, "model": "deepseek/deepseek-v4.1-flash", "choices": [ { "index": 0, "text": " The server is rate limiting you; wait and retry.", "finish_reason": "stop", "logprobs": null } ], "usage": { "prompt_tokens": 12, "completion_tokens": 11, "total_tokens": 23 } } ``` The generated text is in `choices[].text`, not `choices[].message.content` as in chat completions. `finish_reason` is `stop` when the model ended on its own or hit a stop sequence, and `length` when it reached `max_tokens`. The body comes from the provider, so extra fields such as `system_fingerprint` vary by model. ## Streaming Set `"stream": true` and the response is Server-Sent Events. Each event carries a partial `choices[].text`, and the stream ends with `data: [DONE]`: ```text data: {"id":"cmpl-a1b2c3","object":"text_completion","choices":[{"index":0,"text":" The server","finish_reason":null}]} data: {"id":"cmpl-a1b2c3","object":"text_completion","choices":[{"index":0,"text":" is rate limiting you.","finish_reason":"stop"}]} data: [DONE] ``` For billing, the gateway asks the provider to include a usage chunk at the end of a streamed completion. That chunk has an empty `choices` array and a `usage` object, so write your stream parser to accept it. Disconnects, timeouts and proxy buffering behave as described in [streaming](/docs/streaming). ## Billing and limits - **Metering.** Usage is billed on the `prompt_tokens` and `completion_tokens` the provider reports, at the model's input and output prices. If a provider sends no usage, the gateway estimates from the request and response size. - **Reservation.** Before forwarding, the gateway reserves the worst-case cost of the request, using `max_tokens` for the output side. If you leave it out, the reservation assumes 8,192 output tokens, even though the provider's own default may be much smaller. A key close to its monthly spend cap or a balance near zero can refuse a request that omits `max_tokens` while a request with a small one passes. If your balance can't cover the requested `max_tokens`, the gateway may lower it to what the balance covers. You are charged for actual usage, not the reservation. - **Rate limits.** A completions request counts toward your per-minute limit and concurrency like any other inference request. See [rate limits](/docs/rate-limits). - **Failed requests are not billed.** A provider response with a status of 400 or above is returned without a charge. ## Which models work The gateway forwards the call as is. It does not translate a completions request into a chat request, and it does not check whether the model supports plain completion. What you get back depends on the model's provider: | Response | Likely cause | | -------------------------------------------------- | ------------------------------------------------------------------------------------------------------ | | 200 with text | The provider serves `/completions` for this model. | | 400 `invalid_request` | The provider refused the request, for example because the model is chat-only or a field isn't allowed. | | 404 `model_not_found` | The provider has no completions route for this model. Also returned for an id not in the catalog. | | 400 `endpoint_not_supported_for_model` | The model is served only through a provider that speaks the Anthropic Messages protocol. Nothing was sent. | | 502 `upstream_unreachable` | The model has several providers and each one failed or refused. | The upstream message is replaced with a generic one, so keep the `x-tokens-request-id` header if you need support to look at a failure. Full list in [errors](/docs/errors). ## Move to chat completions If a completions call fails or you are writing new code, the same request as a chat completion is one change of shape: ```python # Before: /v1/completions resp = client.completions.create(model=model, prompt="A one-line definition of HTTP 429:", max_tokens=40) text = resp.choices[0].text # After: /v1/chat/completions resp = client.chat.completions.create( model=model, messages=[{"role": "user", "content": "A one-line definition of HTTP 429:"}], max_tokens=40, ) text = resp.choices[0].message.content ``` A chat model answers a question or instruction, where a completions model continues text, so reword prompts that relied on continuation (for example "The three causes are:" becomes "List the three causes."). Features that exist only on completions, such as `suffix` and `echo`, have no chat equivalent. ## Related - [Chat completions](/docs/chat-completions), the endpoint to use for new work. - [Responses](/docs/responses) and [Messages](/docs/messages) for the other inference endpoints. - [Token counting](/docs/token-counting) for estimating prompt size and reading `usage`. --- # Token Counting > POST /v1/messages/count_tokens returns the input size of a request without running the model. How exact it is, what it costs, and how to count or estimate tokens for chat, completions, responses and embeddings. Section: API Reference. Page: https://tokens.bd/docs/token-counting Token counts decide what a request costs, whether it fits a model's context window and how much of your plan it uses. Tokens gives you two ways to know them: a counting endpoint that tells you the input size before you send, and a `usage` object in every response that tells you what was billed. This page covers both, and how to estimate when you have neither. ## Count tokens with POST /v1/messages/count_tokens `POST https://tokens.bd/v1/messages/count_tokens` takes the same body as [`/v1/messages`](/docs/messages) and returns the number of input tokens. It never runs the model and generates no text. It speaks the Anthropic format, so Anthropic SDKs use the bare host `https://tokens.bd` as their base URL. :::code-tabs ```bash title="cURL" curl -i https://tokens.bd/v1/messages/count_tokens \ -H "x-api-key: $TOKENS_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "system": "You are a concise senior engineer.", "messages": [ {"role": "user", "content": "Explain idempotency keys in two sentences."} ] }' ``` ```python title="Python" import os import anthropic client = anthropic.Anthropic( base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"], ) count = client.messages.count_tokens( model="deepseek/deepseek-v4.1-flash", system="You are a concise senior engineer.", messages=[{"role": "user", "content": "Explain idempotency keys in two sentences."}], ) print(count.input_tokens) ``` ```typescript title="Node.js" import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic({ baseURL: "https://tokens.bd", apiKey: process.env.TOKENS_API_KEY, }); const count = await client.messages.countTokens({ model: "deepseek/deepseek-v4.1-flash", system: "You are a concise senior engineer.", messages: [{ role: "user", content: "Explain idempotency keys in two sentences." }], }); console.log(count.input_tokens); ``` ::: The response is one field: ```json { "input_tokens": 31 } ``` The number is illustrative. `model` is required, as on every Tokens endpoint, and `max_tokens` is not needed. You can send `system`, `messages` with text and images, and `tools`; the count includes all of them. Authentication is the same as on `/v1/messages`: `x-api-key` or `Authorization: Bearer`. As elsewhere, send `content-type: application/json`. ## Exact count or estimate The model decides how tokens are counted, so Tokens answers in one of two ways: | What happens | How you can tell | | ------------------------------------------------------------------------------------------------------------- | ---------------------------------------------- | | A provider that serves the model speaks the Anthropic Messages protocol, so the request goes to it and you get its own count. | No special header. | | No provider for the model speaks that protocol, or the one that does is down or answers 404, 405, 429 or a 5xx error. Tokens counts locally. | Response header `x-tokens-estimated: true`. | A model served only through an OpenAI-format provider has no native counting, so it always gets the estimate. Check the header (`curl -i` prints it) before you trust a number to the token. The local estimate is rough by design. It counts the text in `system`, `messages` and `tools` at about four characters per token, adds a flat 1,600 tokens for each image or document block and a few tokens for each message, and returns at least 1. Those constants can change. It is fine for "will this fit" and "about how big is this", and it is not an exact count. Code, JSON and non-English text such as Bengali are the cases where four characters per token is least accurate, because they usually need more tokens per character than English prose. Measure your own data (see below) before you size a budget. Counting follows Anthropic's own endpoint when a provider forwards it. Based on Anthropic's token counting documentation, checked October 2026: the count is itself an estimate that can differ from the billed number by a small amount, it counts system prompts, tools, images and PDFs, and server tools, the MCP connector and `url` or `file` image and document sources are rejected (send images and PDFs as base64). Counts also depend on the model's tokenizer, so count against the model id you will send. ## What counting costs and what it counts toward - **Not billed.** A count never settles a charge, uses no plan credits and writes no usage record. - **Still checked like a request.** Counting goes through the same admission step as inference, so it needs a valid key, a model your key and plan can call, and an account with a plan or wallet balance. It counts toward your per-minute request limit and holds a concurrency slot while it runs. A used-up usage window (429 `window_exhausted`) or a key restricted to other models blocks it too. Unlike `GET /v1/models` and `GET /v1/tokens/usage`, it is not free of rate limits. See [rate limits](/docs/rate-limits). - **Same size limit.** The body can be up to 10 MB, images included. - **Needs the Messages endpoint.** If the gateway has turned off the Anthropic Messages endpoint, counting returns 404 `anthropic_protocol_disabled` and you should use the methods below. Errors on this endpoint use Anthropic's error shape, with the Tokens `code` inside, as described in [Messages](/docs/messages). Codes are in [errors](/docs/errors). ## Count tokens for other endpoints There is no counting endpoint for chat completions, completions, responses or embeddings. Use these instead. ### Read the usage that comes back Every successful inference response tells you what the provider counted, and that is what Tokens bills: | Endpoint | Where the numbers are | | ----------------------- | ------------------------------------------------------------------------------------------------------------ | | `/v1/chat/completions` | `usage.prompt_tokens`, `usage.completion_tokens`, `usage.total_tokens`; cached input in `usage.prompt_tokens_details.cached_tokens` when reported | | `/v1/completions` | `usage.prompt_tokens`, `usage.completion_tokens`, `usage.total_tokens` | | `/v1/embeddings` | `usage.prompt_tokens`, `usage.total_tokens` | | `/v1/responses` | `usage.input_tokens`, `usage.output_tokens` | | `/v1/messages` | `usage.input_tokens`, `usage.output_tokens`, plus cache fields when the provider reports them | For streams, chat completions only include `usage` when you set `stream_options: {"include_usage": true}`; Anthropic-format streams carry it in the `message_start` and `message_delta` events. Details are in [streaming](/docs/streaming). Per-request costs appear in the [usage analytics](/docs/usage-and-alerts), and remaining plan windows and wallet balance are in `GET /v1/tokens/usage` ([models and usage](/docs/models-and-usage)). ### Measure with a one-token request To learn the exact prompt size for a chat model before running a long generation, send the real prompt with `max_tokens` set to 1 and read `usage.prompt_tokens`: ```bash curl https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "messages": [{"role": "user", "content": "Explain idempotency keys in two sentences."}], "max_tokens": 1 }' ``` This is billed: you pay for the input tokens and at most one output token. It also counts as a request. Some models reject or ignore very small `max_tokens` values, so raise it slightly if you get a 400. For a cheaper check on the Messages format, use the counting endpoint above. ### Estimate with a rule of thumb For a quick budget, divide the number of characters by four for English prose. The gateway does the same for its own pre-flight check: before forwarding, it reserves the worst-case cost using the request body size divided by four for input, and `max_tokens` (8,192 if you leave it out) for output. That reservation is a safety margin and not a bill: you are charged for the `usage` the provider reports. It does mean a large `max_tokens` can make a request fail on funds or a spend cap even when the real answer would be short, so set `max_tokens` close to what you need. See [chat completions](/docs/chat-completions). A local tokenizer library gives exact counts only for the family of models it was built for. For other models it is another estimate, so use it for sizing, and use the `usage` object for truth. ## Check that a prompt fits before sending A request fits when the prompt tokens plus `max_tokens` stay inside the model's context window, which is on the model's page in [the catalog](/models). This Python helper counts first and then sets `max_tokens` from the space left: ```python import os import anthropic client = anthropic.Anthropic( base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"], ) CONTEXT_WINDOW = 128_000 # read this from the model's catalog page messages = [{"role": "user", "content": open("big-file.txt").read()}] used = client.messages.count_tokens(model="deepseek/deepseek-v4.1-flash", messages=messages).input_tokens room = CONTEXT_WINDOW - used if room < 1_000: raise SystemExit(f"Prompt is {used} tokens; too close to the context window.") reply = client.messages.create( model="deepseek/deepseek-v4.1-flash", max_tokens=min(4_000, room), messages=messages, ) print(reply.usage) ``` Replace `CONTEXT_WINDOW` with the real value for your model. When the count is an estimate, leave extra room, for example 10 percent. ## Related - [Messages](/docs/messages) for the endpoint whose body counting takes. - [Chat completions](/docs/chat-completions), [Embeddings](/docs/embeddings) and [Legacy completions](/docs/legacy-completions) for the `usage` object on each. - [Plans and wallet](/docs/plans-and-wallet) for how tokens turn into credits and cost. --- # Vision: sending images to models > Send images to a model on Chat Completions (image_url parts) or Messages (image blocks): URL and base64 forms, size limits, what the gateway translates, how image input is billed, and how to check a model accepts images. Section: API Reference. Page: https://tokens.bd/docs/vision Vision means sending an image to a model and asking about it: a screenshot of an error, a diagram, a photo of a whiteboard. On Tokens images travel inside the normal chat request, on `POST https://tokens.bd/v1/chat/completions` as `image_url` content parts and on `POST https://tokens.bd/v1/messages` as `image` content blocks. The gateway carries the image to the provider and bills what the provider reports. Whether the model can read the image is up to the model. Tokens has no separate image endpoint. Image generation, file uploads and the Anthropic Files API are not served (see [Models and usage](/docs/models-and-usage)). This page is about image **input** only. ## Check that the model accepts images Not every model reads images. The model catalog at [/models](/models) records the context window, prices and a description for each model, but it has no field that says "supports images". To find out: - Read the model's description on its [/models](/models) page and the maker's own documentation for the model. - See the vision section of [Choosing a model](/docs/choosing-a-model), which lists several models by what they accept. - Send one small test image (the examples below) and ask "What is in this image?". A model that reads images answers about the picture. A text-only model does not always fail loudly. The provider decides: it can return a 400 (which reaches you as `invalid_request`, see [errors](/docs/errors)), or it can drop the image and answer from the text alone. If the answer ignores the picture, treat that as "this model does not take images". ## Chat Completions: image_url parts Put an array in the message `content` instead of a string. Each element is a `text` part or an `image_url` part. The `url` is either a public `https` URL or a base64 data URL (`data:;base64,`). ```json { "role": "user", "content": [ { "type": "text", "text": "What does this error screenshot say?" }, { "type": "image_url", "image_url": { "url": "data:image/png;base64,", "detail": "auto" } } ] } ``` `detail` is OpenAI's optional hint (`low`, `high` or `auto`; `auto` when omitted). Other providers may ignore it, and the gateway drops it when it has to translate the request for an Anthropic-style provider (see below). :::code-tabs ```bash title="cURL (public URL)" curl https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "max_tokens": 300, "messages": [ { "role": "user", "content": [ {"type": "image_url", "image_url": {"url": "https://YOUR-HOST/path/to/image.png"}}, {"type": "text", "text": "Describe this image in two sentences."} ] } ] }' ``` ```python title="Python (local file)" import base64 import os from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) with open("screenshot.png", "rb") as f: b64 = base64.b64encode(f.read()).decode("ascii") resp = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", max_tokens=300, messages=[ { "role": "user", "content": [ {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{b64}"}}, {"type": "text", "text": "Describe this image in two sentences."}, ], } ], ) print(resp.choices[0].message.content) print(resp.usage) ``` ```typescript title="Node.js (local file)" import fs from "node:fs"; import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY }); const b64 = fs.readFileSync("screenshot.png").toString("base64"); const resp = await client.chat.completions.create({ model: "deepseek/deepseek-v4.1-flash", max_tokens: 300, messages: [ { role: "user", content: [ { type: "image_url", image_url: { url: `data:image/png;base64,${b64}` } }, { type: "text", text: "Describe this image in two sentences." }, ], }, ], }); console.log(resp.choices[0].message.content, resp.usage); ``` ::: Replace `https://YOUR-HOST/path/to/image.png` with a link to an image you control. The `media type` in a data URL must match the file (`image/png`, `image/jpeg`, `image/webp`, `image/gif`). :::note[Who fetches a URL image] When you send an `https` URL, Tokens does not download the image. It forwards the URL and the provider fetches it, so the link must be public and reachable from the provider. Some providers accept only base64 data URLs. If a URL fails and the same image works as base64, use base64. ::: ## Messages: image content blocks On `/v1/messages` an image is a content block with a `source`. Two source types work through Tokens: ```json { "role": "user", "content": [ { "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "" } }, { "type": "text", "text": "What does this error screenshot say?" } ] } ``` ```json { "type": "image", "source": { "type": "url", "url": "https://YOUR-HOST/path/to/image.png" } } ``` `data` is the raw base64 string with no `data:` prefix. Anthropic's third source, `{"type": "file", "file_id": "..."}`, needs their Files API, which Tokens does not serve. Send the image itself. :::code-tabs ```bash title="cURL (public URL)" curl https://tokens.bd/v1/messages \ -H "x-api-key: $TOKENS_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "max_tokens": 300, "messages": [ { "role": "user", "content": [ {"type": "image", "source": {"type": "url", "url": "https://YOUR-HOST/path/to/image.png"}}, {"type": "text", "text": "Describe this image in two sentences."} ] } ] }' ``` ```python title="Python (local file)" import base64 import os import anthropic client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"]) with open("screenshot.png", "rb") as f: b64 = base64.standard_b64encode(f.read()).decode("ascii") message = client.messages.create( model="deepseek/deepseek-v4.1-flash", max_tokens=300, messages=[ { "role": "user", "content": [ {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": b64}}, {"type": "text", "text": "Describe this image in two sentences."}, ], } ], ) print(message.content[0].text) print(message.usage) ``` ```typescript title="Node.js (local file)" import fs from "node:fs"; import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic({ baseURL: "https://tokens.bd", apiKey: process.env.TOKENS_API_KEY }); const b64 = fs.readFileSync("screenshot.png").toString("base64"); const message = await client.messages.create({ model: "deepseek/deepseek-v4.1-flash", max_tokens: 300, messages: [ { role: "user", content: [ { type: "image", source: { type: "base64", media_type: "image/png", data: b64 } }, { type: "text", text: "Describe this image in two sentences." }, ], }, ], }); console.log(message.content[0], message.usage); ``` ::: Anthropic's guidance is to put images before the text that asks about them, and to label several images in the text ("Image 1:", "Image 2:") so you can refer to them. Earlier images in a conversation stay visible to the model, but you must resend them in every request, because the API is stateless. ## Base64 from the command line A base64 image is too large to type into a `curl -d '...'` argument. Write the request to a file and send the file. On macOS and Linux, with `jq` installed: ```bash base64 screenshot.png | tr -d '\n' > screenshot.b64 jq -n --rawfile img screenshot.b64 '{ model: "deepseek/deepseek-v4.1-flash", max_tokens: 300, messages: [{ role: "user", content: [ {type: "image_url", image_url: {url: ("data:image/png;base64," + $img)}}, {type: "text", text: "Describe this image in two sentences."} ] }] }' > request.json curl https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d @request.json ``` ## What the gateway does with images On a native request (the provider speaks the same format as your request) the gateway forwards the body and changes only the `model` field (plus a usage-reporting option on streamed Chat Completions). It does not decode, resize or inspect the image. Some models are only available from a provider that speaks the other format. Then the gateway translates, and images are handled like this: | Your request | Provider speaks | What happens to images | | ----------------------------------------------- | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Messages (`image` block) | OpenAI format | Base64 becomes a data URL and a URL stays a URL, as `image_url` parts of the user message. Images in `assistant` turns are dropped. Images inside a `tool_result` are not sent as images (see warning). | | Chat Completions (`image_url` part) | Anthropic format | A data URL becomes a base64 `image` block, any other URL becomes a URL `image` block. `detail` is dropped. Only images in `user` messages are carried. | :::warning[Images inside tool results] On a Messages request that the gateway translates to the OpenAI format, a `tool_result` whose content includes an `image` block is flattened to text, and the image block is written into that text as JSON. The model does not see a picture and the base64 data counts as input tokens. If an agent returns screenshots from tools, use a model that is served in Anthropic format, or have the tool describe the image in text. ::: The same translation rules, and the other fields that are carried across, are described in [Messages](/docs/messages). ## Size limits Two limits matter, and the smaller one wins. **Tokens:** the whole request body can be at most 10 MB, otherwise the gateway answers `413 request_entity_too_large`. Base64 makes data about a third larger than the file, so all images in one request together can be roughly 7 MB of image files at most, and less once the text and earlier conversation turns are counted. Every request resends the full conversation, including earlier images. **The provider:** each provider has its own limits on image formats, dimensions, count and size per image. As checked in October 2026: - Anthropic: JPEG, PNG, GIF (first frame only) and WebP; up to 8000 x 8000 pixels per image; 10 MB per image (base64); up to 100 images per request on models with a 200K-token context window and 600 on others. Above 20 images in one request a stricter per-image pixel limit applies, and Anthropic suggests keeping each side to 2000 px. Source: [Anthropic vision documentation](https://platform.claude.com/docs/en/build-with-claude/vision). - OpenAI: PNG, JPEG, WebP and non-animated GIF; up to 1,500 images per request. Source: [OpenAI images and vision guide](https://developers.openai.com/api/docs/guides/images-vision). OpenAI's 512 MB payload limit is above Tokens' 10 MB, so the Tokens limit applies first. - Other makers publish their own limits. Check the maker's documentation for the model you use. Resize large photos before sending. A phone photo is often 4000 pixels wide, and models downscale it anyway, so you pay for upload size and latency without gaining detail. Keep text in screenshots legible: very aggressive JPEG compression makes small text unreadable. ## How image input is billed An image is converted to input tokens by the provider, and the provider reports the total in `usage`. Tokens bills that reported input count at the model's input price, like any other input. Roughly, providers count images in patches: Anthropic documents `ceil(width / 28) x ceil(height / 28)` tokens per image before downscaling to a cap (about 1,568 tokens on its standard tier, up to 4,784 on its high-resolution tier). OpenAI documents patch-based and tile-based counts that depend on the model and on `detail`. Other makers differ, so read `usage` from a test request to see the real cost for your model. Admission works from a different number. Before forwarding a request the gateway reserves its worst-case cost against your balance, and it estimates the input side as one token per four characters of the request body. A large base64 image is a lot of characters, so the reservation for a request with a multi-megabyte image can be far above what the provider ends up charging. You are charged the real usage, not the reservation, but a low balance or a key close to its spend cap can be refused with `insufficient_credits` or `monthly_spend_cap_exceeded` for a request that would have cost very little. Resizing the image fixes both the cost and the refusal. See the reservation notes in [Chat Completions](/docs/chat-completions). ## Troubleshooting | Symptom | Likely cause | Fix | | ----------------------------------------------------- | ---------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- | | `400 invalid_request` mentioning images or content | The model does not accept image input, or the media type is not one it supports | Try a model that takes images; check the data URL's media type matches the file. | | The answer ignores the picture | The provider dropped the image for a text-only model | Switch to a model that reads images. | | `413 request_entity_too_large` | Body over 10 MB, usually base64 images or a long history of images | Resize or compress, send fewer images, or drop old turns. | | `402 insufficient_credits` with a small balance | The input estimate counts the base64 characters | Resize the image, or top up in [billing](/dashboard/billing). | | URL image fails, base64 works | The provider could not fetch the URL, or only accepts base64 | Use base64, or make the URL public. | | Messages request with tool screenshots, model is blind | `tool_result` images are flattened on translated models | See the warning above. | If the error does not say what is wrong, the upstream message is replaced with a generic one. Send the `x-tokens-request-id` to [support](/docs/support), and see [errors](/docs/errors) for the codes. --- # Structured output: JSON mode and JSON Schema > Get machine-readable JSON from a model through Tokens: JSON mode, JSON Schema with response_format, forced tool calls, and Anthropic's output_config, with examples, validation advice and what the gateway drops. Section: API Reference. Page: https://tokens.bd/docs/structured-output Structured output means asking a model for JSON your code can parse, instead of prose. There are four ways to do it through Tokens, and which one works depends on the model and on how the gateway reaches it: | Method | Endpoint | How strong the guarantee is | | ----------------------------------------------- | -------------------------- | ------------------------------------------------------------------- | | JSON mode: `response_format: {"type": "json_object"}` | Chat Completions | Valid JSON, but no promise about the keys | | JSON Schema: `response_format: {"type": "json_schema", ...}` | Chat Completions | Output follows your schema (with `strict: true`), if the model supports it | | A forced tool call whose arguments are your schema | Chat Completions, Messages | Arguments follow the schema as well as the model manages; works on more models | | `output_config.format` with a JSON Schema | Messages | Output follows your schema, on models that support it | None of this is enforced by Tokens. The gateway does not read, validate or repair the JSON. It forwards the request fields and returns the provider's answer, so the guarantees above are the provider's, and a model that does not support a feature either ignores the field or rejects the request. Always parse and validate the result in your own code. ## What the gateway does and drops - On **Chat Completions** to a provider that speaks the OpenAI format, `response_format` is forwarded unchanged. - On **Messages** to a provider that speaks the Anthropic format, `output_config`, `tools` and `tool_choice` are forwarded unchanged. - When a model is only available from a provider that speaks the **other** format, the gateway translates the request, and the translation carries only a fixed list of fields. A Chat Completions request to an Anthropic-format provider does not carry `response_format` (nor `n`, `seed`, `logprobs` or the penalty fields). A Messages request to an OpenAI-format provider does not carry `output_config`. In both cases the field has no effect and nothing tells you so; the model just answers in free text. - Tools survive translation in both directions, which is why a forced tool call is the most portable method. See [Tool calling](/docs/tool-calling) and [Messages](/docs/messages). If a structured request comes back as prose, the likely causes are a model without support, or a translated path that dropped the field. Try the tool-call method on the same model. ## How to find out whether a model supports it The model catalog at [/models](/models) records context window, prices and a description, but no field for JSON mode or JSON Schema support. To check: 1. Read the maker's documentation for the model (for example, whether it lists "JSON output" or "structured outputs"). 2. Send a small request with the field and look at the result: a 400 means the model or provider rejects it, free-form prose means the field was ignored, valid JSON means it works for that request. 3. Run the test several times. Unconstrained JSON mode can pass once and fail on the next prompt. [Choosing a model](/docs/choosing-a-model) lists the models Tokens compares most often. ## JSON mode JSON mode makes the model return a valid JSON object, with no guarantee about which keys it contains. Tell the model in a message to produce JSON and describe the shape you want. OpenAI's documentation says you must instruct the model to produce JSON in some message in the conversation; without that the model can fill the response with whitespace until it hits the token limit, and the API can reject the request when the word "JSON" is missing from the context. :::code-tabs ```bash title="cURL" curl https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "max_tokens": 400, "response_format": {"type": "json_object"}, "messages": [ {"role": "system", "content": "Reply with a JSON object with the keys \"language\" and \"summary\"."}, {"role": "user", "content": "def add(a, b): return a + b"} ] }' ``` ```python title="Python" import json import os from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) resp = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", max_tokens=400, response_format={"type": "json_object"}, messages=[ {"role": "system", "content": 'Reply with a JSON object with the keys "language" and "summary".'}, {"role": "user", "content": "def add(a, b): return a + b"}, ], ) choice = resp.choices[0] if choice.finish_reason == "length": raise RuntimeError("Cut off by max_tokens: the JSON is incomplete") data = json.loads(choice.message.content) print(data) ``` ```typescript title="Node.js" import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY }); const resp = await client.chat.completions.create({ model: "deepseek/deepseek-v4.1-flash", max_tokens: 400, response_format: { type: "json_object" }, messages: [ { role: "system", content: 'Reply with a JSON object with the keys "language" and "summary".' }, { role: "user", content: "def add(a, b): return a + b" }, ], }); const choice = resp.choices[0]; if (choice.finish_reason === "length") throw new Error("Cut off by max_tokens: the JSON is incomplete"); console.log(JSON.parse(choice.message.content ?? "{}")); ``` ::: Check `finish_reason` before parsing. If it is `length`, the model hit `max_tokens` in the middle of the object and the JSON is cut off. Reasoning models spend part of that limit on thinking first, so give them more room; see [Reasoning and thinking models](/docs/reasoning). ## JSON Schema with response_format JSON Schema output constrains the answer to a schema you supply. On Chat Completions the request looks like this (the shape is OpenAI's, checked October 2026 in [OpenAI's structured outputs guide](https://developers.openai.com/api/docs/guides/structured-outputs) and [migration guide](https://developers.openai.com/api/docs/guides/migrate-to-responses)): ```json { "response_format": { "type": "json_schema", "json_schema": { "name": "ticket", "strict": true, "schema": { "type": "object", "properties": { "title": { "type": "string" }, "severity": { "type": "string", "enum": ["low", "medium", "high"] }, "needs_followup": { "type": "boolean" } }, "required": ["title", "severity", "needs_followup"], "additionalProperties": false } } } } ``` With `strict: true` OpenAI applies these rules to the schema, and rejects a schema that breaks them: - The root must be an object, and it cannot be an `anyOf`. - Every property must be listed in `required`. To make a field optional, allow `null`: `{"type": ["string", "null"]}`. - Every object needs `"additionalProperties": false`. - At most 5,000 properties in total and 10 levels of nesting. Other makers support a subset of JSON Schema, and the subset differs. If a provider rejects your schema, simplify it: flat objects, simple types, `enum` for fixed choices. :::code-tabs ```python title="Python" import json import os from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) schema = { "type": "object", "properties": { "title": {"type": "string"}, "severity": {"type": "string", "enum": ["low", "medium", "high"]}, "needs_followup": {"type": "boolean"}, }, "required": ["title", "severity", "needs_followup"], "additionalProperties": False, } resp = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", max_tokens=400, messages=[ {"role": "user", "content": "Users report that the login page returns a 500 after the last deploy."} ], response_format={ "type": "json_schema", "json_schema": {"name": "ticket", "strict": True, "schema": schema}, }, ) choice = resp.choices[0] if getattr(choice.message, "refusal", None): print("The model refused:", choice.message.refusal) elif choice.finish_reason != "stop": print("Incomplete answer, finish_reason =", choice.finish_reason) else: ticket = json.loads(choice.message.content) print(ticket["severity"], ticket["title"]) ``` ```typescript title="Node.js" import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY }); const schema = { type: "object", properties: { title: { type: "string" }, severity: { type: "string", enum: ["low", "medium", "high"] }, needs_followup: { type: "boolean" }, }, required: ["title", "severity", "needs_followup"], additionalProperties: false, }; const resp = await client.chat.completions.create({ model: "deepseek/deepseek-v4.1-flash", max_tokens: 400, messages: [ { role: "user", content: "Users report that the login page returns a 500 after the last deploy." }, ], response_format: { type: "json_schema", json_schema: { name: "ticket", strict: true, schema } }, }); const choice = resp.choices[0]; if (choice.finish_reason !== "stop") { console.log("Incomplete or refused answer:", choice.finish_reason, choice.message.refusal); } else { console.log(JSON.parse(choice.message.content ?? "{}")); } ``` ```bash title="cURL" curl https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "max_tokens": 400, "messages": [ {"role": "user", "content": "Users report that the login page returns a 500 after the last deploy."} ], "response_format": { "type": "json_schema", "json_schema": { "name": "ticket", "strict": true, "schema": { "type": "object", "properties": { "title": {"type": "string"}, "severity": {"type": "string", "enum": ["low", "medium", "high"]}, "needs_followup": {"type": "boolean"} }, "required": ["title", "severity", "needs_followup"], "additionalProperties": false } } } }' ``` ::: Handle three outcomes before you parse: a refusal (OpenAI reports it in `message.refusal`), an incomplete answer (`finish_reason` of `length` or `content_filter`), and an answer that parses but fails your own checks. The OpenAI and Anthropic SDKs have helpers that build the schema from a Pydantic model or a Zod schema and parse the reply; they send the same JSON as above, so they work through Tokens when the model supports the feature. :::note[Responses API] The Responses API takes the same schema under `text.format` instead of `response_format`, with `name`, `strict` and `schema` one level up and no nested `json_schema` key. See [Responses API](/docs/responses). Support depends on the provider behind the model. ::: ## Structured output with tools A forced tool call is the most portable method, because tool calling is supported by many more models than `response_format`, and it survives the gateway's format translation. Define one function whose `parameters` are the schema you want, force the model to call it, and read the arguments. ```python title="Chat Completions: forced function" import json import os from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) tools = [ { "type": "function", "function": { "name": "record_ticket", "description": "Record a support ticket extracted from the user's message.", "strict": True, "parameters": { "type": "object", "properties": { "title": {"type": "string"}, "severity": {"type": "string", "enum": ["low", "medium", "high"]}, }, "required": ["title", "severity"], "additionalProperties": False, }, }, } ] resp = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", max_tokens=400, messages=[{"role": "user", "content": "The login page returns a 500 after the last deploy."}], tools=tools, tool_choice={"type": "function", "function": {"name": "record_ticket"}}, ) call = resp.choices[0].message.tool_calls[0] ticket = json.loads(call.function.arguments) print(ticket) ``` On Chat Completions `arguments` is a JSON string, so parse it. `strict` inside the `function` object is OpenAI's switch for exact schema adherence and has the same schema rules as above; models and providers that do not know it may ignore it. Nothing here is run by Tokens, which only carries the call. Support for `tool_choice` set to a specific function varies by model: some ignore it, some return a 400. See [Tool calling](/docs/tool-calling). OpenAI's documentation says Chat Completions does not support tool calling with a `reasoning_effort` other than `none` on GPT-5.4 and later models. If a forced tool call fails on an OpenAI reasoning model, set `reasoning_effort` to `none` or use the Responses API. On Messages the same idea uses Anthropic's tool format, and the arguments come back already parsed as an object in a `tool_use` block: ```python title="Messages: forced tool" import os import anthropic client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"]) tools = [ { "name": "record_ticket", "description": "Record a support ticket extracted from the user's message.", "input_schema": { "type": "object", "properties": { "title": {"type": "string"}, "severity": {"type": "string", "enum": ["low", "medium", "high"]}, }, "required": ["title", "severity"], }, } ] message = client.messages.create( model="deepseek/deepseek-v4.1-flash", max_tokens=400, tools=tools, tool_choice={"type": "tool", "name": "record_ticket"}, messages=[{"role": "user", "content": "The login page returns a 500 after the last deploy."}], ) block = next(b for b in message.content if b.type == "tool_use") print(block.input) ``` Anthropic notes that forcing a tool with `tool_choice` of `any` or `tool` cannot be combined with its manual extended thinking (`thinking.type: "enabled"`), and that a few of its newest models do not accept forced tool use even with adaptive thinking. If you need both thinking and structured output on Claude models, read [Reasoning and thinking models](/docs/reasoning) first. ## JSON Schema on Messages with output_config Anthropic's Messages API has its own JSON Schema output, set in `output_config.format` (checked October 2026 in [Anthropic's structured outputs guide](https://platform.claude.com/docs/en/build-with-claude/structured-outputs)). The reply is text in a `text` block that holds the JSON. Anthropic also offers `strict: true` on a tool definition to guarantee that tool names and inputs match the schema. ```bash title="cURL" curl https://tokens.bd/v1/messages \ -H "x-api-key: $TOKENS_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "max_tokens": 400, "messages": [ {"role": "user", "content": "The login page returns a 500 after the last deploy."} ], "output_config": { "format": { "type": "json_schema", "schema": { "type": "object", "properties": { "title": {"type": "string"}, "severity": {"type": "string", "enum": ["low", "medium", "high"]} }, "required": ["title", "severity"], "additionalProperties": false } } } }' ``` ```python title="Python" import json import os import anthropic client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"]) response = client.messages.create( model="deepseek/deepseek-v4.1-flash", max_tokens=400, messages=[{"role": "user", "content": "The login page returns a 500 after the last deploy."}], output_config={ "format": { "type": "json_schema", "schema": { "type": "object", "properties": { "title": {"type": "string"}, "severity": {"type": "string", "enum": ["low", "medium", "high"]}, }, "required": ["title", "severity"], "additionalProperties": False, }, } }, ) if response.stop_reason in ("refusal", "max_tokens"): raise RuntimeError(f"No valid JSON: stop_reason = {response.stop_reason}") text = next(b.text for b in response.content if b.type == "text") print(json.loads(text)) ``` This needs a recent Anthropic SDK that knows `output_config`; if your SDK version rejects the argument, upgrade it. Anthropic lists the Claude models that support the feature, and it supports only a subset of JSON Schema: no recursive schemas, no numeric or string length constraints, `additionalProperties: false` on objects, and unsupported features return a 400. A `stop_reason` of `refusal` still returns a 200 and is billed, and the output may not match the schema; `max_tokens` can leave it incomplete. Remember the limits of this path on Tokens. It works when the model is served by a provider that speaks the Anthropic format. When the gateway has to translate your Messages request to the OpenAI format, `output_config` is dropped (see above). Use the forced-tool method if you need one code path for every model. ## Streaming and structured output Structured output works with `"stream": true`, but the JSON is only valid when the stream is complete. Concatenate the `content` deltas (or the tool-call `arguments` fragments, or Anthropic's `input_json_delta` pieces) and parse once at the end. Do not parse partial text. See [streaming](/docs/streaming). ## Cost and limits - The schema, tool definitions and your instructions are input tokens on every request. A large schema costs on every call. - Tokens bills the usage the provider reports, at the model's input and output prices. There is no surcharge for structured output. - `max_tokens` bounds the whole answer. A truncated JSON object does not parse, so set it above the longest answer you expect, and add room if the model reasons first. - The gateway reserves the worst-case cost from `max_tokens` before it forwards the request. See the reservation notes in [Chat Completions](/docs/chat-completions). ## Make it reliable - Validate every reply against your own schema in code (Pydantic, Zod, `jsonschema`), even when you used strict mode. - Keep schemas flat and give each field a short `description`; the model reads it. - Set low `temperature` for extraction tasks. Some reasoning models reject or ignore it; see [Reasoning and thinking models](/docs/reasoning). - On a parse failure, retry once with the error message appended, then fail loudly. Each retry is a new billed request. - Test the exact model you will run in production. A schema that works on one model can be rejected by another. --- # Reasoning and thinking models > Use reasoning models through Tokens: reasoning_effort on Chat Completions, thinking and effort on Messages, how reasoning text is returned and streamed, how reasoning tokens are counted and billed, and what the gateway rewrites or drops. Section: API Reference. Page: https://tokens.bd/docs/reasoning Reasoning models (OpenAI calls them reasoning models, Anthropic calls the feature thinking) work through a problem before they write the answer. That extra work improves results on hard coding, math and planning tasks. It also costs tokens and time: the model produces thinking tokens you may never see, and you are billed for them. This page covers what you can send, what comes back, how Tokens counts and bills the thinking tokens, and where the gateway changes your request. The parameter names and the meaning of each value belong to the provider that serves the model. Tokens forwards them and does not interpret them, so read the sections marked as provider behaviour as a summary of the maker's documentation (checked October 2026), not as a promise from Tokens. ## The short version | You call | Control reasoning with | Reasoning text comes back as | | --------------------------- | ------------------------------------------------------------------ | -------------------------------------------------------------- | | `/v1/chat/completions` | `reasoning_effort` | `reasoning_content` on the message or the stream delta, when the model exposes it | | `/v1/messages` | `thinking` and `output_config.effort` (older Claude models: `thinking.budget_tokens`) | `thinking` content blocks and `thinking_delta` stream events | | `/v1/responses` | `reasoning.effort` and `reasoning.summary` | A `reasoning` output item with a `summary` list | Which of these a given model honours depends on the model. Some models always reason and have no switch, some take an effort level, some take a token budget, and some ignore every one of these fields. ## Find out whether a model reasons The model catalog at [/models](/models) lists context window, prices and a description for each model. It has no flag for "reasoning" or "thinking". To find out: - Read the model's page and the maker's documentation. [Choosing a model](/docs/choosing-a-model) notes a few models that always think, such as Kimi K3 and GLM-5.3, and says that always-on thinking adds output tokens to every call. - Send a test request and look at the answer. A reasoning model usually returns a `reasoning_content` field, a `thinking` block, or reasoning token details inside `usage`. A much larger `completion_tokens` than the visible answer also shows it. ```python import os from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) resp = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", max_tokens=2000, messages=[{"role": "user", "content": "A train leaves at 09:10 and arrives at 11:45. How long is the trip?"}], ) message = resp.choices[0].message print("visible answer:", message.content) print("reasoning text:", getattr(message, "reasoning_content", None)) print("usage:", resp.usage) ``` `reasoning_content` is not part of OpenAI's own API. It is a convention some OpenAI-compatible providers use for the model's reasoning text, and the gateway passes it through when the provider sends it. ## Chat Completions: reasoning_effort `reasoning_effort` is OpenAI's parameter for how much a reasoning model thinks. Per OpenAI's reasoning guide the accepted values depend on the model and include `none`, `minimal`, `low`, `medium`, `high`, `xhigh` and `max`; most current OpenAI models default to `medium` when you omit it. Some models reject some values: OpenAI documents that GPT-6 Astra rejects `none` with a 400, and that GPT-6.1 Sol rejects both `none` and `minimal`. Lower effort gives faster and cheaper answers. Higher effort suits hard problems. :::code-tabs ```bash title="cURL" curl https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "max_completion_tokens": 4000, "reasoning_effort": "low", "messages": [ {"role": "user", "content": "Find the bug: for i in range(len(xs)+1): total += xs[i]"} ] }' ``` ```python title="Python" import os from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) resp = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", max_completion_tokens=4000, reasoning_effort="low", messages=[ {"role": "user", "content": "Find the bug: for i in range(len(xs)+1): total += xs[i]"} ], ) print(resp.choices[0].message.content) print(resp.usage) ``` ```typescript title="Node.js" import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY }); const resp = await client.chat.completions.create({ model: "deepseek/deepseek-v4.1-flash", max_completion_tokens: 4000, reasoning_effort: "low", messages: [{ role: "user", content: "Find the bug: for i in range(len(xs)+1): total += xs[i]" }], }); console.log(resp.choices[0].message.content, resp.usage); ``` ::: What happens to the field depends on the model and the provider behind it. If the model does not take `reasoning_effort`, the provider either ignores it or answers with a 400, which reaches you as `invalid_request` ([errors](/docs/errors)). Some makers use their own field names for the same idea. The gateway does not validate Chat Completions fields, so a provider-specific field in the body is forwarded as it is; with the OpenAI SDKs you can add one with `extra_body`. Whether the provider honours it is the maker's decision, so check their documentation. Two OpenAI details from its documentation that matter in practice: reasoning models count their thinking against the output limit, and Chat Completions does not support tool calling with a `reasoning_effort` other than `none` starting with GPT-5.4. For tool use with those models, use `reasoning_effort: "none"` or the Responses API. ### What the gateway rewrites for OpenAI reasoning models OpenAI's o-series and GPT-5 and later models refuse `max_tokens` and any `temperature` or `top_p` other than 1. Many clients send them anyway, so for a Chat Completions request to one of these models (matched by the model's name) the gateway edits the body before forwarding: - `max_tokens` is renamed to `max_completion_tokens` (if you sent both, yours is kept). - `temperature` and `top_p` are removed unless they are exactly 1. This applies on `/v1/chat/completions` only, not on Messages or Responses, and not to other makers' models. Models from other makers get your parameters as sent, and some of them reject or ignore `temperature` while they reason. If your balance cannot cover the `max_tokens` or `max_completion_tokens` you asked for, the gateway can lower it (never below 16) as described in [Chat Completions](/docs/chat-completions). On a reasoning model that cut can use up the whole allowance on thinking and leave an empty or short answer with `finish_reason: "length"`. Set a limit your balance covers, or top up in [billing](/dashboard/billing). ## Messages: thinking and effort On `/v1/messages` Anthropic's models take a `thinking` object and, on current models, an `output_config.effort` level. The gateway forwards both unchanged to a provider that speaks the Messages API. Per Anthropic's documentation (checked October 2026, [thinking](https://platform.claude.com/docs/en/build-with-claude/thinking), [extended thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking), [effort](https://platform.claude.com/docs/en/build-with-claude/effort)): - **Current Claude models (4.7 and later, including the 5.x line):** use `thinking: {"type": "adaptive"}` and control depth with `output_config: {"effort": "low" | "medium" | "high" | "xhigh" | "max"}`. Which levels exist depends on the model. On several 5.x models thinking is already on without any `thinking` field. The old form `thinking: {"type": "enabled", "budget_tokens": N}` returns a 400 on these models. - **Claude 4.6:** the `enabled` form still works but is deprecated. - **Claude 4.5 and earlier:** only the `enabled` form exists. `budget_tokens` is a target that must be at least 1,024 and less than `max_tokens`. - **Seeing the thinking text:** on many current models the `thinking` field of each block comes back empty by default (`display` is `omitted`). Send `"display": "summarized"` inside the `thinking` object to get a summary of the reasoning. Anthropic does not return the raw chain of thought under any setting. :::code-tabs ```python title="Python: current Claude models" import os import anthropic client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"]) message = client.messages.create( model="deepseek/deepseek-v4.1-flash", max_tokens=8000, thinking={"type": "adaptive", "display": "summarized"}, output_config={"effort": "medium"}, messages=[{"role": "user", "content": "Plan a safe rollout for a database column rename."}], ) for block in message.content: if block.type == "thinking": print("thinking summary:", block.thinking) elif block.type == "text": print("answer:", block.text) print(message.usage) ``` ```python title="Python: Claude 4.5 and earlier" import os import anthropic client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"]) message = client.messages.create( model="deepseek/deepseek-v4.1-flash", max_tokens=8000, thinking={"type": "enabled", "budget_tokens": 4000}, # at least 1024, below max_tokens messages=[{"role": "user", "content": "Plan a safe rollout for a database column rename."}], ) for block in message.content: if block.type == "thinking": print("thinking:", block.thinking) elif block.type == "text": print("answer:", block.text) ``` ```bash title="cURL" curl https://tokens.bd/v1/messages \ -H "x-api-key: $TOKENS_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "max_tokens": 8000, "thinking": {"type": "adaptive", "display": "summarized"}, "output_config": {"effort": "medium"}, "messages": [ {"role": "user", "content": "Plan a safe rollout for a database column rename."} ] }' ``` ::: `output_config` needs a recent Anthropic SDK. These examples show the request shape for Claude models; use the form your model's maker documents. Which form a model takes is not recorded in the Tokens catalog. A response with thinking has `thinking` blocks before the `text` block: ```json { "content": [ { "type": "thinking", "thinking": "The rename needs a two-step deploy...", "signature": "EosnCkYICxIM..." }, { "type": "text", "text": "Roll it out in three steps: add the new column..." } ], "stop_reason": "end_turn", "usage": { "input_tokens": 24, "output_tokens": 912 } } ``` The `signature` is encrypted data that lets the model continue its own reasoning. Anthropic requires thinking blocks (and any `redacted_thinking` blocks, which hold encrypted content with no readable text) to be sent back unchanged when you return tool results in a tool-use loop. Keep the whole assistant `content` array, not just the text; the loop in [Tool calling](/docs/tool-calling) shows `{"role": "assistant", "content": resp.content}`. Anthropic also documents that manual extended thinking only allows `tool_choice` of `auto` or `none`. ## What survives translation Each model is served by one or more providers, and the gateway prefers one that speaks the same format as your request. When a model is only available in the other format, the gateway translates (see [Messages](/docs/messages)), and reasoning is handled like this: | Direction | Reasoning controls in your request | Reasoning in the answer | | -------------------------------------------- | --------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ | | Messages request, OpenAI-format provider | `thinking` and `output_config` are **not carried** | A `reasoning_content` the provider sends is **not** turned into a `thinking` block; it is dropped | | Chat Completions request, Anthropic-format provider | `reasoning_effort` is **not carried** | `thinking` blocks become `message.reasoning_content` (and `delta.reasoning_content` when streaming). `redacted_thinking` blocks and `signature` values are dropped | Consequences: - On a translated path you cannot switch reasoning on or change its level with those fields. Whether the model reasons is then the model's default. - The usage numbers still include the tokens the model spent thinking, so you are billed for reasoning you cannot see. - A Chat Completions request cannot carry Anthropic's thinking signatures, so a tool-use loop that needs thinking blocks passed back should use `/v1/messages`. If you need a particular reasoning setting to apply, call the endpoint in the model's native format, and confirm with a test request that the answer carries reasoning (a `thinking` block or `reasoning_content`) or that `usage` changes with the setting. ## Responses API On `POST https://tokens.bd/v1/responses` OpenAI's parameter is `reasoning`, with `effort` and, for a readable summary, `summary`. The gateway forwards the body as it is, and support depends on the provider behind the model (see [Responses API](/docs/responses)). ```python import os from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) resp = client.responses.create( model="deepseek/deepseek-v4.1-flash", input="Find the bug: for i in range(len(xs)+1): total += xs[i]", reasoning={"effort": "low", "summary": "auto"}, max_output_tokens=4000, ) print(resp.output_text) print(resp.usage) ``` OpenAI documents that raw reasoning tokens are never returned. With `summary` set, the response has a `reasoning` output item whose `summary` list holds a readable summary. OpenAI also says to pass reasoning items back to the model together with function-call outputs when you continue a tool loop. ## Streaming reasoning With `"stream": true` the reasoning arrives before the answer, and the stream stays quiet for a while if the model thinks first. **Chat Completions.** Providers that expose reasoning text send it as `delta.reasoning_content`, followed by `delta.content` for the answer. Translated Anthropic thinking is delivered the same way. A client that ignores unknown fields shows only the answer. :::code-tabs ```python title="Python" import os from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) stream = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", max_tokens=4000, stream=True, stream_options={"include_usage": True}, messages=[{"role": "user", "content": "Why does 0.1 + 0.2 != 0.3 in floating point?"}], ) for chunk in stream: if chunk.usage: print("\nusage:", chunk.usage) if not chunk.choices: continue delta = chunk.choices[0].delta reasoning = getattr(delta, "reasoning_content", None) if reasoning: print(reasoning, end="", flush=True) # reasoning text, if the model sends it if delta.content: print(delta.content, end="", flush=True) ``` ```typescript title="Node.js" import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY }); const stream = await client.chat.completions.create({ model: "deepseek/deepseek-v4.1-flash", max_tokens: 4000, stream: true, stream_options: { include_usage: true }, messages: [{ role: "user", content: "Why does 0.1 + 0.2 != 0.3 in floating point?" }], }); for await (const chunk of stream) { if (chunk.usage) console.log("\nusage:", chunk.usage); const delta = chunk.choices[0]?.delta as | { content?: string | null; reasoning_content?: string | null } | undefined; if (delta?.reasoning_content) process.stdout.write(delta.reasoning_content); if (delta?.content) process.stdout.write(delta.content); } ``` ::: **Messages.** Thinking arrives as `content_block_delta` events with a `thinking_delta`, then one `signature_delta` just before the block closes, and the text blocks follow. If `display` is `omitted`, the `thinking_delta` events carry an empty string and only the signature arrives. ```python import os import anthropic client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"]) stream = client.messages.create( model="deepseek/deepseek-v4.1-flash", max_tokens=8000, stream=True, thinking={"type": "adaptive", "display": "summarized"}, messages=[{"role": "user", "content": "Why does 0.1 + 0.2 != 0.3 in floating point?"}], ) for event in stream: if event.type == "content_block_delta": if event.delta.type == "thinking_delta": print(event.delta.thinking, end="", flush=True) elif event.delta.type == "text_delta": print(event.delta.text, end="", flush=True) ``` More on stream formats, disconnects and timeouts is in [streaming](/docs/streaming). ## Token counts and billing Reasoning tokens are output tokens. Both OpenAI and Anthropic document that the tokens a model spends thinking are billed as output tokens, and they count toward the output limit (`max_completion_tokens` or `max_tokens`) even when the reasoning text is not returned to you. - Tokens bills what the provider reports in `usage`: input, output, and cache reads and writes, each at the model's catalog price for that kind of token ([prices](/models)). Reasoning tokens are part of the output count, so they are priced at the model's output rate. There is no separate reasoning price. - The visible answer can be a small part of the bill. A response with a 300-token answer and 1,200 thinking tokens is billed as 1,500 output tokens. Usage analytics ([usage](/docs/usage-and-alerts)) show output tokens as billed. - Summaries and omitted thinking do not reduce the bill. Anthropic documents that you are charged for the full thinking tokens, not the summary, and that `display: "omitted"` only reduces latency. - Some providers report the split inside `usage`, for example `completion_tokens_details.reasoning_tokens` on Chat Completions, `output_tokens_details.reasoning_tokens` on Responses and `output_tokens_details.thinking_tokens` on Messages. The fields come from the provider and are passed through; they are informational, and the output total is what is billed. - If a provider reports no usage at all, the gateway estimates from the response size, counting `reasoning_content` text along with the answer. - Anthropic documents that its newer models keep thinking blocks from earlier turns in context and bill them as input on the later turns. A long tool loop with thinking resends them each round, so cost grows with the number of rounds. Before forwarding, the gateway reserves the worst case against your balance using your output limit (8,192 tokens if you set none), as explained in [Chat Completions](/docs/chat-completions). A large limit for a reasoning model, such as 32,000 or more, reserves a large amount even if the model finishes early. You are charged for actual usage. ## How much to allow - Set the output limit high enough for thinking plus the answer. If the model uses it all on thinking, you get a cut-off or empty answer with `finish_reason: "length"` (Chat Completions) or `stop_reason: "max_tokens"` (Messages). OpenAI suggests reserving at least 25,000 tokens for reasoning and output when you start out on its reasoning models, then tuning down. - Start at low or medium effort and raise it only if answers are not good enough. Higher effort means more thinking tokens and more latency. - For routine work (renames, formatting, simple edits) a non-reasoning model or `reasoning_effort: "none"` where the model allows it is cheaper and faster. ## Latency and timeouts A model that thinks before it answers can be silent for a long time before the first token. Stream the response so your client can show progress and so the connection stays active. The gateway also watches for silence. When several providers serve a model, it waits for a streaming provider's first event for 30 seconds by default (the operator can change this), then tries the next provider if one is configured; a non-streaming request gets at least 120 seconds. The last provider in the chain has no early cut-off. A provider that sends nothing for 600 seconds ends in `504 upstream_timeout` ([errors](/docs/errors)). If a slow reasoning model times out on non-streaming calls, switch to streaming and set your own client timeout above the longest answer you expect. ## Troubleshooting | Symptom | Likely cause | Fix | | -------------------------------------------------------- | --------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------- | | 400 `invalid_request` after adding `reasoning_effort` | The model or provider does not accept that field or value | Remove it, or use a value the maker documents for that model. | | 400 on `thinking` with `budget_tokens` | The Claude model only supports adaptive thinking | Use `thinking: {"type": "adaptive"}` and `output_config.effort`. | | 400 on `max_tokens` or `temperature` | An OpenAI reasoning model on a path where the gateway does not rewrite them | On Chat Completions the gateway rewrites them; on Messages or Responses use the model's own parameter names. | | Empty answer, `finish_reason: "length"` | Thinking used the whole output limit | Raise `max_tokens` or `max_completion_tokens`, or lower the effort. | | No reasoning text in the response | The model does not expose it, `display` is `omitted`, or the path dropped it | Set `display: "summarized"` on Claude models, or see the translation table above. | | `thinking` block missing on Messages for a non-Claude model | The model is served in OpenAI format and `reasoning_content` is dropped | Read the reasoning on Chat Completions instead. | | Bill larger than the visible answer | Thinking tokens are output tokens | Lower the effort or budget, or use a smaller model. | | 504 `upstream_timeout` on a long non-streaming call | The model thought for longer than the upstream wait | Stream the response. | | Tool loop fails with a thinking-block error on Claude | Thinking blocks were not passed back unchanged | Resend the full assistant `content`, including `thinking` and `redacted_thinking` blocks, on `/v1/messages`. | --- # Prompt caching > How prompt caching works through Tokens: what the provider does, what the gateway passes through, how cache reads and writes show up in usage, and how they are priced and billed. Section: API Reference. Page: https://tokens.bd/docs/prompt-caching Prompt caching lets a provider reuse the work it already did on the start of your prompt. A request that repeats a long system prompt, tool list or conversation history is billed at a lower rate for the repeated part. For a coding agent, which resends almost the same context on every turn, it is the largest cost lever there is. The cache belongs to the upstream provider. Tokens does not run a cache of its own. What Tokens does is forward your request, read the cache counts the provider reports, and bill those tokens at the cache prices set for the model. This page separates the two: what is the provider's rule, and what Tokens does. ## Who does what | Question | Answer | | ------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------- | | Who decides whether a request hits the cache? | The provider that serves it. Tokens cannot force a hit. | | Who sets cache lifetime and minimum prompt size? | The provider, per model. | | Does Tokens change cache markers in my request? | Not when the request goes to the provider in its own format. When the gateway translates between formats, cache markers are dropped. | | Does Tokens change the usage numbers? | No. The usage object in the response is the provider's. When the gateway translates formats it maps the cache fields (see below). | | Who prices cache reads and writes? | Tokens, per model. A model with no cache price set is billed at its normal input price. | | Where do I see what I paid? | [Usage](/dashboard/usage) shows cached tokens per request and the cost of the request. | ## Two kinds of caching **Automatic caching.** You change nothing. The provider detects that the start of your prompt matches an earlier request and reuses it. OpenAI's models work this way, and so does DeepSeek's. The provider's documentation says the cache works on a prefix: a request hits only if its beginning is identical to an earlier one. DeepSeek describes its cache as best effort and does not promise a hit. Prompts shorter than the provider's minimum are not cached; OpenAI lists 1,024 tokens for its newest models. **Explicit caching with `cache_control`.** On Anthropic's Messages API you mark the content you want cached with a `cache_control` block. Anthropic's rules, checked October 2026 in its [prompt caching documentation](https://platform.claude.com/docs/en/build-with-claude/prompt-caching): - A cache breakpoint covers everything before it, in this order: `tools`, `system`, then `messages`. - You can set up to four breakpoints per request. A top-level `cache_control` field places one automatically on the last cacheable block. - The default lifetime is 5 minutes, refreshed each time the cached content is used. `"ttl": "1h"` asks for one hour. - There is a minimum prompt length that depends on the model (512 to 4,096 tokens in Anthropic's table). A shorter prompt is processed normally, without an error and without caching. - A cache entry becomes available only after the first response begins, so parallel requests sent at the same moment do not share it. - Any change at one level, for example editing a tool definition, invalidates that level and everything after it. These are Anthropic's rules for its own models. Other providers that accept `cache_control` may differ. Check the provider's documentation for the model you use. ## Using cache_control through /v1/messages Send `cache_control` exactly as you would to Anthropic. When the model is served by a provider that speaks the Messages API, the request body goes to it as written. ```python import os import anthropic client = anthropic.Anthropic( base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"], ) # Must be longer than the model's minimum cacheable length. project_notes = open("project-notes.txt", encoding="utf-8").read() def ask(question: str) -> None: message = client.messages.create( model="deepseek/deepseek-v4.1-flash", max_tokens=300, system=[ { "type": "text", "text": project_notes, "cache_control": {"type": "ephemeral"}, } ], messages=[{"role": "user", "content": question}], ) usage = message.usage print( "input:", usage.input_tokens, "| cache write:", usage.cache_creation_input_tokens or 0, "| cache read:", usage.cache_read_input_tokens or 0, "| output:", usage.output_tokens, ) ask("Summarize the notes in one sentence.") # first call: expect a cache write ask("List three risks mentioned in the notes.") # within the TTL: expect a cache read ``` If the second call shows a cache read, the markers reached a provider that honors them. If both calls show zero for cache write and cache read, one of these is true: the prompt is under the model's minimum, the serving provider does not support explicit caching, or the request was translated (next section). Testing with two identical calls is the only reliable check for a given model. ## What is passed through and what is dropped Each model is served by one or more upstream providers, and each provider speaks the OpenAI format, the Anthropic format, or both. Tokens sends your request in your own format when the provider supports it, and translates when it does not. | You call | Provider speaks | What happens to caching | | ----------------------- | -------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- | | `/v1/messages` | Anthropic format | Body forwarded as is. `cache_control` reaches the provider. Usage comes back in Anthropic fields. | | `/v1/messages` | OpenAI format only | Translated to chat completions. `cache_control` markers are dropped. Automatic caching at the provider can still apply. | | `/v1/chat/completions` | OpenAI format | Body forwarded as is. Automatic caching applies. Fields such as `prompt_cache_key` go to the provider if you send them. | | `/v1/chat/completions` | Anthropic format only | Translated to Messages. Cache markers are not added or carried. The Anthropic provider's usage is mapped back to OpenAI fields. | You cannot choose which provider serves a request. Providers are tried in priority order, and a request moves to the next one only when the first fails or does not answer in time. The next provider has not seen your prefix, so expect a miss on that request. ## Cache fields in the response The gateway reads these fields from the provider's usage object. Everything else in the response is passed through. | Format and field | What it means | How Tokens bills it | | ------------------------------------------------------------- | ---------------------------------------------------------------------------------------- | ---------------------- | | Chat completions: `prompt_tokens` | All input tokens, cached ones included | Minus cached, as input | | Chat completions: `prompt_tokens_details.cached_tokens` | Input tokens served from cache | Cache read | | Chat completions: `cache_creation_input_tokens` (top level) | Present when the gateway translated an Anthropic answer: input tokens written to cache | Cache write | | Messages: `input_tokens` | Input tokens that were neither read from nor written to cache | Input | | Messages: `cache_read_input_tokens` | Input tokens served from cache | Cache read | | Messages: `cache_creation_input_tokens` | Input tokens written to cache | Cache write | Two consequences: - The two formats count differently. In chat completions `prompt_tokens` already contains the cached tokens. In Messages, `input_tokens` does not. Tokens converts both into the same four buckets before pricing, so you do not need to adjust for the difference. - A provider that reports caching only in a field Tokens does not read, for example a custom hit counter, has its cached tokens billed as normal input. If the numbers on the [usage page](/dashboard/usage) show no cached tokens for a model that you know caches, tell [support](/docs/support) with a request id. On OpenAI-format streams the usage chunk only reaches you if you set `stream_options.include_usage`. Billing does not depend on it: the gateway meters the stream either way. See [streaming](/docs/streaming). ## How cache tokens are billed Every request is split into four buckets, and each has its own price per million tokens: ```text cost = input x input price + cache read x cache-read price + cache write x cache-write price + output x output price ``` The prices come from the model's entry in the Tokens catalog. Where a model has no cache price, that bucket is billed at the model's input price, so caching saves nothing for that model. A cache price of zero means the bucket is free. Plan discounts, if your plan has one, apply to the total. An example with made-up prices: input $2.00, cache read $0.20, output $8.00 per million tokens. A request has 10,000 input tokens, of which 9,880 come from the cache, and produces 300 output tokens. | Bucket | Tokens | Price per million | Cost | | ----------- | ------ | ----------------- | ------------ | | Input | 120 | $2.00 | $0.000240 | | Cache read | 9,880 | $0.20 | $0.001976 | | Output | 300 | $8.00 | $0.002400 | | **Total** | | | **$0.004616** | The same request with no cache hit costs 10,000 x $2.00 per million plus the same output, $0.0224. Real prices for each model are on [the model catalog](/models) and [pricing](/pricing), and the usage page shows what each request was actually charged. Providers that charge extra for writing a cache entry (Anthropic charges 1.25 times the input price for a 5-minute write and 2 times for a 1-hour write, on its own API) pass that on only if the Tokens catalog has a cache-write price for the model. Where it does not, writes are billed as input. Cached requests are still checked against your balance before they run. The gateway reserves a worst-case amount using your `max_tokens` and an estimate of the input priced at the normal input rate, because it cannot know in advance that the cache will hit. A request with a very large cached prefix can therefore be refused for a low balance even though its final cost would be small. Settlement then charges the real amount. More in [chat completions](/docs/chat-completions). ## See cache hits on the dashboard The activity table on [Usage](/dashboard/usage) shows tokens per request as `input · N cached · output`. The cached figure is cache reads and cache writes added together. The row's cost is the full charge for the request. A request that was estimated, because the provider sent no usage, carries an Estimated badge. To see the split between reads and writes for one call, read the `usage` object in the response itself and log it. ## Get more cache hits These follow from the providers' rules above. - **Put stable content first.** System prompt, tool definitions and long reference text go at the start, unchanged between requests. Variable content, such as the new user message, goes last. - **Keep the prefix byte-identical.** A timestamp, request id or random value near the top of the system prompt changes every request and defeats the cache for everything after it. Keep tool definitions in a fixed order. - **Append to history, do not rewrite it.** Editing or compacting earlier turns creates a new prefix. - **Keep the same model.** Caches are per model. - **Stay inside the lifetime.** With a 5-minute lifetime, a pause longer than that means the next request pays for a write again. For slow interactive use on Anthropic models, `"ttl": "1h"` can cost less overall, because a 1-hour write costs more than a 5-minute one. - **Send the first request alone.** Wait for the first response to begin before firing parallel requests that share the prefix. - **Make the prompt long enough.** Short prompts are below the provider's minimum and never cache. ## Troubleshooting **No cached tokens ever.** Run the two-call test above. Check, in order: the prompt length against the model's minimum, whether anything at the top of the prompt changes between calls, whether the calls are more than a few minutes apart, and whether you are using `/v1/messages` with a model that is only served in OpenAI format (markers dropped). If the test shows zero reads on a model you expect to cache, send the two request ids to [support](/docs/support). **Cached tokens appear, but the request cost about the same.** The model probably has no cache-read price in the catalog, so cached tokens are billed at the input price. Compare the cost on the usage page against the model's prices. **Cache write tokens higher than expected.** Each change near the start of the prompt writes a new entry. Look for content that varies per request, and for tool lists that change between turns. **`cache_control` returns a 400.** The provider rejected the request body. Anthropic returns an error if a request combines automatic caching with four explicit breakpoints. Check the provider's limits, then see [errors](/docs/errors). **A coding agent costs more than you expected.** The agent decides what goes in the prompt and where. When it compacts or rewrites earlier turns, the prefix changes and the next request misses the cache. Check your agent's settings for compaction, and see its page under [coding agents](/docs/claude-code). --- # Browser and mobile apps > Why a Tokens API key must never be in a web page or mobile app, and the pattern that works: your own backend in between. Working Next.js and Express examples with streaming, plus the mobile version. Section: API Reference. Page: https://tokens.bd/docs/browser-and-mobile A web page or mobile app cannot call Tokens directly. Anything shipped to a user's device can be read by that user, and a Tokens key spends your money. Browsers add a second block: Tokens does not send CORS headers on its responses, so a page's JavaScript cannot read the answer anyway. The fix is one small backend of your own. The app calls your backend, your backend checks who is asking and calls Tokens with the key, and the answer goes back to the app. This page shows that backend for Next.js and for Express, with streaming passed through, then the same idea for a mobile app. ## Why a key in client code fails - **The key is not secret.** A browser shows every request in its developer tools. A mobile app can be unpacked and its strings read. Environment variables that a framework copies into the client bundle, such as those starting with `NEXT_PUBLIC_` in Next.js, are public by definition. - **A leaked key is a spending problem.** Anyone with it can run requests until the key's spend cap or your balance is gone. - **CORS does not protect you, and does not help you either.** The gateway answers the browser's preflight request, but the real responses carry no `Access-Control-Allow-Origin` header, so the browser refuses to hand the answer to your page. Do not read that as a safety mechanism: the request can still leave the browser with your key in it. Treat a key in client code as already leaked. If a key has ever been in client code, revoke it in [API keys](/dashboard/keys) and create a new one. See [API keys](/docs/api-keys). ## The pattern ```text Browser or app --(your user's session)--> Your backend --(Tokens key)--> Tokens <------- streamed answer --------------- <---- stream ----- ``` Your backend does five things: 1. **Authenticates your user.** Use the session or token your app already has. An open endpoint is an open wallet. 2. **Validates the input.** Accept a message list or a prompt, check its size, and ignore everything else the client sends. 3. **Fixes the model and limits on the server.** The client does not choose `model`, `max_tokens` or tools. Otherwise a user can pick the most expensive model and the largest output. 4. **Calls Tokens with the key from an environment variable,** and passes the stream through without buffering. 5. **Hides upstream errors.** A billing or limit error is your problem, not your user's. Log the `x-tokens-request-id` and tell the user something short. Use a dedicated key for this service, with a monthly spend cap and an allowed-models list, so a bug or an abusive user can only cost a known amount. See [API keys](/docs/api-keys) and the [production checklist](/docs/production-checklist). ## Next.js route handler with streaming This is a route handler in the App Router. It runs only on the server, so `process.env.TOKENS_API_KEY` is never sent to the browser. ```typescript title="app/api/chat/route.ts" import { getSessionUserId } from "@/lib/session"; // your own auth export const runtime = "nodejs"; export const dynamic = "force-dynamic"; const TOKENS_URL = "https://tokens.bd/v1/chat/completions"; const MODEL = "deepseek/deepseek-v4.1-flash"; type ChatMessage = { role: "user" | "assistant"; content: string }; function parseMessages(value: unknown): ChatMessage[] | null { if (!Array.isArray(value) || value.length === 0 || value.length > 40) return null; const messages: ChatMessage[] = []; for (const item of value) { if (typeof item !== "object" || item === null) return null; const { role, content } = item as Record; if (role !== "user" && role !== "assistant") return null; if (typeof content !== "string" || content.length > 20_000) return null; messages.push({ role, content }); } return messages; } export async function POST(req: Request): Promise { const userId = await getSessionUserId(req); if (!userId) return Response.json({ error: "unauthorized" }, { status: 401 }); const payload: unknown = await req.json().catch(() => null); const messages = parseMessages((payload as { messages?: unknown } | null)?.messages); if (!messages) return Response.json({ error: "invalid_request" }, { status: 400 }); let upstream: Response; try { upstream = await fetch(TOKENS_URL, { method: "POST", headers: { Authorization: `Bearer ${process.env.TOKENS_API_KEY}`, "Content-Type": "application/json", "x-request-id": crypto.randomUUID(), }, body: JSON.stringify({ model: MODEL, messages, stream: true, max_tokens: 1024 }), signal: req.signal, // client left: cancel the upstream request too }); } catch { return Response.json({ error: "unavailable" }, { status: 502 }); } if (!upstream.ok || !upstream.body) { console.error("tokens call failed", { userId, status: upstream.status, requestId: upstream.headers.get("x-tokens-request-id"), }); await upstream.body?.cancel(); const headers = new Headers(); const retryAfter = upstream.headers.get("retry-after"); if (upstream.status === 429 && retryAfter) headers.set("Retry-After", retryAfter); return Response.json( { error: upstream.status === 429 ? "busy" : "unavailable" }, { status: upstream.status === 429 ? 429 : 502, headers } ); } return new Response(upstream.body, { headers: { "Content-Type": "text/event-stream; charset=utf-8", "Cache-Control": "no-cache, no-transform", "X-Accel-Buffering": "no", }, }); } ``` `getSessionUserId` stands for whatever your app already uses to identify a user (Auth.js, Clerk, Supabase Auth, your own cookie). It is the one line you replace. If your host limits how long a route may run, raise the limit for this route; reasoning models can take minutes. ### Read the stream in the browser The answer is the same Server-Sent Events text Tokens sends, so the client reads `data:` lines. Chunks can end in the middle of a line, so keep the unfinished part in a buffer. ```typescript title="chat-client.ts" export async function streamChat( messages: { role: "user" | "assistant"; content: string }[], onText: (text: string) => void, signal?: AbortSignal ): Promise { const res = await fetch("/api/chat", { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({ messages }), signal, }); if (!res.ok || !res.body) throw new Error(`Chat failed with status ${res.status}`); const reader = res.body.pipeThrough(new TextDecoderStream()).getReader(); let buffer = ""; for (;;) { const { done, value } = await reader.read(); if (done) return; buffer += value; const lines = buffer.split("\n"); buffer = lines.pop() ?? ""; for (const line of lines) { if (!line.startsWith("data:")) continue; const data = line.slice(5).trim(); if (data === "[DONE]") return; const delta = JSON.parse(data).choices?.[0]?.delta?.content; if (delta) onText(delta); } } } ``` Pass an `AbortSignal` from a "stop" button. The abort cancels your backend request, which cancels the Tokens request, and Tokens bills only what was generated up to that point. ## Node.js with Express The same backend as a plain Express app. This uses the global `fetch` of Node.js 18 or later and ES modules. ```javascript title="server.js" import express from "express"; import { randomUUID } from "node:crypto"; import { Readable } from "node:stream"; import { pipeline } from "node:stream/promises"; import { requireUser } from "./auth.js"; // your own auth middleware const TOKENS_URL = "https://tokens.bd/v1/chat/completions"; const MODEL = "deepseek/deepseek-v4.1-flash"; const app = express(); app.use(express.json({ limit: "100kb" })); function validMessages(messages) { return ( Array.isArray(messages) && messages.length > 0 && messages.length <= 40 && messages.every( (m) => (m?.role === "user" || m?.role === "assistant") && typeof m.content === "string" && m.content.length <= 20_000 ) ); } app.post("/api/chat", requireUser, async (req, res) => { const messages = req.body?.messages; if (!validMessages(messages)) return res.status(400).json({ error: "invalid_request" }); // The client closed the connection before the answer finished: stop the upstream request. const controller = new AbortController(); res.on("close", () => { if (!res.writableEnded) controller.abort(); }); let upstream; try { upstream = await fetch(TOKENS_URL, { method: "POST", headers: { Authorization: `Bearer ${process.env.TOKENS_API_KEY}`, "Content-Type": "application/json", "x-request-id": randomUUID(), }, body: JSON.stringify({ model: MODEL, messages, stream: true, max_tokens: 1024 }), signal: controller.signal, }); } catch { return res.status(502).json({ error: "unavailable" }); } if (!upstream.ok || !upstream.body) { console.error("tokens call failed", { userId: req.user.id, status: upstream.status, requestId: upstream.headers.get("x-tokens-request-id"), }); await upstream.body?.cancel(); const retryAfter = upstream.headers.get("retry-after"); if (upstream.status === 429 && retryAfter) res.set("Retry-After", retryAfter); return res .status(upstream.status === 429 ? 429 : 502) .json({ error: upstream.status === 429 ? "busy" : "unavailable" }); } res.status(200).set({ "Content-Type": "text/event-stream; charset=utf-8", "Cache-Control": "no-cache, no-transform", "X-Accel-Buffering": "no", }); res.flushHeaders(); try { await pipeline(Readable.fromWeb(upstream.body), res); } catch { // The client left or the upstream stream broke. Nothing more to send. } }); app.listen(3001, () => console.log("Listening on http://localhost:3001")); ``` `requireUser` is your authentication middleware; it sets `req.user`. The browser code from the previous section works against this server unchanged. If you run Express behind nginx, set `proxy_buffering off;` for this route. Without it the stream arrives in one block. Do not compress `text/event-stream` responses. More in [troubleshooting](/docs/troubleshooting). ## Mobile apps A mobile app does exactly the same thing: it talks to your backend and never to Tokens. The app sends the user's own session token. Your backend checks it, applies your per-user limits, and calls Tokens with the key. The simplest mobile design is a non-streaming call. Not every platform's default HTTP client delivers a response body piece by piece, and a plain request is easier to get right first. Add a second backend route that returns the whole answer as JSON: ```typescript title="app/api/ask/route.ts" import { getSessionUserId } from "@/lib/session"; // your own auth export const runtime = "nodejs"; export const dynamic = "force-dynamic"; export async function POST(req: Request): Promise { const userId = await getSessionUserId(req); if (!userId) return Response.json({ error: "unauthorized" }, { status: 401 }); const payload = (await req.json().catch(() => null)) as { prompt?: unknown } | null; const prompt = payload?.prompt; if (typeof prompt !== "string" || prompt.length === 0 || prompt.length > 20_000) { return Response.json({ error: "invalid_request" }, { status: 400 }); } let upstream: Response; try { upstream = await fetch("https://tokens.bd/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.TOKENS_API_KEY}`, "Content-Type": "application/json", "x-request-id": crypto.randomUUID(), }, body: JSON.stringify({ model: "deepseek/deepseek-v4.1-flash", messages: [{ role: "user", content: prompt }], max_tokens: 800, }), signal: AbortSignal.timeout(120_000), }); } catch { return Response.json({ error: "unavailable" }, { status: 502 }); } if (!upstream.ok) { console.error("tokens call failed", { userId, status: upstream.status, requestId: upstream.headers.get("x-tokens-request-id"), }); return Response.json({ error: "unavailable" }, { status: upstream.status === 429 ? 429 : 502 }); } const data = (await upstream.json()) as { choices?: { message?: { content?: string } }[] }; return Response.json({ text: data.choices?.[0]?.message?.content ?? "" }); } ``` The app calls that route with its own session token: :::code-tabs ```swift title="Swift (iOS)" import Foundation struct AskReply: Decodable { let text: String } func ask(_ prompt: String, sessionToken: String) async throws -> String { var request = URLRequest(url: URL(string: "https://api.example.com/api/ask")!) request.httpMethod = "POST" request.setValue("Bearer \(sessionToken)", forHTTPHeaderField: "Authorization") request.setValue("application/json", forHTTPHeaderField: "Content-Type") request.httpBody = try JSONEncoder().encode(["prompt": prompt]) request.timeoutInterval = 130 let (data, response) = try await URLSession.shared.data(for: request) guard let http = response as? HTTPURLResponse, http.statusCode == 200 else { throw URLError(.badServerResponse) } return try JSONDecoder().decode(AskReply.self, from: data).text } ``` ```typescript title="React Native" export async function ask(prompt: string, sessionToken: string): Promise { const res = await fetch("https://api.example.com/api/ask", { method: "POST", headers: { Authorization: `Bearer ${sessionToken}`, "Content-Type": "application/json", }, body: JSON.stringify({ prompt }), }); if (!res.ok) throw new Error(`Request failed with status ${res.status}`); const data = (await res.json()) as { text: string }; return data.text; } ``` ::: Replace `https://api.example.com` with your own backend's address. The app holds a short-lived session for your user, never a Tokens key, so a stolen phone or a decompiled app costs you one user's session, not your account. If you want streaming on mobile, first check that your HTTP client exposes the response body as it arrives. If it does not, keep the non-streaming route and show a loading state. ## Protect the backend itself Moving the key to a server moves the target. Anyone who can call your endpoint can spend through it. - **Rate-limit per user,** not per IP only, and cap how many requests run at once. Tokens limits requests per minute and concurrency for your whole account, so one noisy user can use up everyone's share. See [rate limits](/docs/rate-limits). - **Cap the size of what you accept,** as the examples do: number of messages, characters per message and `max_tokens`. - **Keep the model list on the server.** If users can pick a model, choose it from your own list. - **Do not expose upstream error text.** Return a short code and keep the details in your logs. - **Use a separate key with a spend cap.** If the backend is abused, the cap stops the spending. - **Do not trust the `Origin` header** as authentication. Any script can set it. ## Troubleshooting **The browser console shows a CORS error.** The page is calling Tokens directly. Change it to call your own route, as above. **The answer arrives all at once.** Something between your backend and the user buffers the stream: a proxy, a compression layer, or an HTTP client that waits for the full body. Check [troubleshooting](/docs/troubleshooting) and test the backend with `curl -N`. **The backend returns 502 and the log shows 401 or 403.** The key is wrong, revoked or restricted. See [API keys](/docs/api-keys). **The backend returns 502 and the log shows 402.** The balance or plan is used up. Top up in [billing](/dashboard/billing). Do not retry in a loop; see [errors](/docs/errors). **Users see 429 from your backend.** Your account hit its per-minute or concurrency limit, or a usage window ran out. The log has the code. See [rate limits](/docs/rate-limits). --- # Production checklist > What to set up before real users depend on the Tokens API: retries and backoff, timeouts, Retry-After, request ids, a key per environment, spend caps, alerts, key rotation, and what to do on 402 and 429. Section: API Reference. Page: https://tokens.bd/docs/production-checklist A script that works on your laptop fails in production in predictable ways: a burst of traffic hits a rate limit, a long answer outlives a timeout, the balance runs out at night, a key leaks. Each one has a fix that takes minutes to set up now and hours to set up during an outage. This page is the list, in the order that matters, with the reasons. The details of each error are in [errors](/docs/errors) and [rate limits](/docs/rate-limits). This page says what to do about them in a live system. ## The checklist | Done | Item | | ---- | --------------------------------------------------------------------------------------------- | | [ ] | The API key is in an environment variable or secret store, never in code or client bundles | | [ ] | One key per environment (development, staging, production), each with a spend cap | | [ ] | Production keys restricted to the models the service uses | | [ ] | `max_tokens` is set on every request | | [ ] | Client timeouts are set explicitly, and longer for long or reasoning requests | | [ ] | Retries with exponential backoff and jitter, only on the errors that are worth retrying | | [ ] | `Retry-After` is honored | | [ ] | 402 and `window_exhausted` are never retried in a loop | | [ ] | Parallel requests are capped below your plan's concurrency limit | | [ ] | `x-tokens-request-id` is logged for every request, successful or not | | [ ] | Usage alerts are on, and someone reads them | | [ ] | A plan for key rotation exists and has been tried once | | [ ] | The product has a defined behavior when the model is unavailable | ## Keys: one per environment, capped Create a separate key for each environment and each service. When something goes wrong, the key name in the [usage page](/dashboard/usage) and the key list tells you where the spend came from, and you can revoke one key without taking the others down. Set a **monthly spend cap** when you create the key. A cap turns a runaway loop, a retry storm or a leaked key into an error instead of a bill. It cannot be edited afterwards: to change it, create a new key. Also set the **allowed models** list, so a bug cannot call a model more expensive than the one you tested with. Keys per account are limited by plan (the default is 3 active keys). If you want development, staging and production keys plus one for a CI job, check your plan's limit on [pricing](/pricing) before you plan the layout. Local development can share the key your tool already uses, as long as it has a cap. Keep keys out of source control, out of container images and out of client-side code. If your users run in a browser or a phone, see [browser and mobile](/docs/browser-and-mobile). The full guide is in [API keys](/docs/api-keys). ## Set max_tokens on every request Before a request runs, the gateway reserves its worst-case cost against your plan or wallet. The output part of that reservation uses your `max_tokens` (or `max_completion_tokens`), and 8,192 when you leave it unset. Two things follow: - A large or missing `max_tokens` can get a request refused near a key's spend cap or when the balance is low, even though the real answer would be short. - When your balance covers only part of the reservation, the gateway lowers `max_tokens` to what the balance affords, down to a floor of 16. The answer then stops early with `finish_reason: "length"`. If you see truncated answers when money is low, this is why. Pick a limit that fits the longest answer you want, not the largest the model allows. ## Retries Retry the errors that can go away on their own, and nothing else. | Retry? | Errors | How | | --------------------------------- | ------------------------------------------------------------------------------------- | ----------------------------------------------------------------------- | | Yes, after `Retry-After` | 429 `rate_limited`, `concurrency_limit`, `rate_limit_exceeded` | Wait the number of seconds in the header, plus a little random jitter | | Yes, with exponential backoff | 500, 502, 503, 504, and connection errors or timeouts on your side | Start near 1 second, double each time, cap at 30 seconds, 4 or 5 tries | | Only after the reset | 429 `window_exhausted`, `model_limit_reached` | Not in a loop. `Retry-After` can be hours or days. Pause the job | | No, fix something first | 400, 401, 402, 403, 404, 413 | These repeat forever. Alert a person | Three details that decide whether retries help or hurt: - **A retry is a new request.** There is no idempotency key. If the first request actually completed on the gateway while your client timed out, you were billed for it, and the retry is billed again. Keep the number of attempts low, and be careful with long prompts. - **Check the SDK first.** The OpenAI and Anthropic SDKs retry some errors by themselves (`max_retries`). Stacking your own loop on top multiplies the attempts. Either turn the SDK's retries off, as the example in [rate limits](/docs/rate-limits) does, or rely on them and add nothing. - **Do not retry a stream that has already started and been used.** If an answer broke halfway, a retry sends the whole prompt again and bills it again. Decide whether the partial text is good enough. Retry only streams that failed before any content arrived. See [streaming](/docs/streaming). The gateway already fails over between providers for models that have more than one, before it returns an error to you. A 5xx you see means the gateway already tried. ### A request wrapper This wrapper calls chat completions with a timeout, a fresh request id per attempt, backoff that honors `Retry-After`, no retry on the codes that need a person, and a log line for every failure. It uses the global `fetch` of Node.js 18 or later. ```typescript title="tokens-client.ts" const BASE_URL = "https://tokens.bd/v1"; const RETRY_STATUS = new Set([429, 500, 502, 503, 504]); const WAIT_FOR_RESET = new Set(["window_exhausted", "model_limit_reached"]); export type TokensResult = | { ok: true; data: unknown; requestId: string | null } | { ok: false; status: number; code: string; requestId: string | null }; const sleep = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms)); export async function chat( body: Record, maxAttempts = 4 ): Promise { let delay = 1000; for (let attempt = 1; ; attempt++) { const clientRequestId = crypto.randomUUID(); let status = 0; // 0 means no response: timeout or connection error let code = "network_error"; let requestId: string | null = null; let wait = delay; try { const res = await fetch(`${BASE_URL}/chat/completions`, { method: "POST", headers: { Authorization: `Bearer ${process.env.TOKENS_API_KEY}`, "Content-Type": "application/json", "x-request-id": clientRequestId, }, body: JSON.stringify(body), signal: AbortSignal.timeout(180_000), }); requestId = res.headers.get("x-tokens-request-id"); if (res.ok) return { ok: true, data: await res.json(), requestId }; status = res.status; const err = (await res.json().catch(() => null)) as { error?: { code?: string } } | null; code = err?.error?.code ?? "unknown"; const retryAfter = Number(res.headers.get("retry-after")); if (retryAfter > 0) wait = retryAfter * 1000; } catch { // Timeout or connection error: handled below like a 5xx. } console.warn( JSON.stringify({ event: "tokens_call_failed", attempt, status, code, requestId, clientRequestId }) ); const retryable = status === 0 || RETRY_STATUS.has(status); if (!retryable || WAIT_FOR_RESET.has(code) || attempt >= maxAttempts) { return { ok: false, status, code, requestId }; } await sleep(Math.min(wait, 60_000) + Math.random() * 500); delay = Math.min(delay * 2, 30_000); } } ``` The 180-second timeout suits a non-streaming call with a moderate `max_tokens`. See the next section for long requests. ## Timeouts Set timeouts on purpose. Library defaults are rarely right for model calls. - **The gateway is patient.** For a model with a single provider it waits up to 600 seconds for the first byte, and up to 300 seconds of silence between chunks once a stream is flowing. A reasoning model can think for minutes before it says anything. If the provider does not start in time you get 504 `upstream_timeout`. - **Your client must be at least as patient** as the request needs. Node.js's built-in `fetch` gives up after 300 seconds without response headers and after 300 seconds between body chunks by default, according to the documentation of undici, the HTTP client behind it. A request that needs longer needs a longer limit. - **Stream anything that can run long.** A non-streaming request sends nothing until the whole answer exists, so a proxy or load balancer in front of your app can cut it as idle. A stream sends data as it goes. Use `stream: true` for long answers and long agent tasks. - **Mind your own infrastructure.** Reverse proxies, serverless platforms and mobile networks often have idle or total-time limits of 30 to 120 seconds. Check each hop between your user and Tokens. - **Cancel when the user leaves.** Pass an abort signal. The gateway stops the upstream request and bills only what was generated. - **Do not set one timeout for everything.** A short summary and an agent turn with tools deserve different limits. When a model has a backup provider, the gateway moves to it after about 30 seconds without a first token on a stream (120 seconds on a non-streaming request). You never see the switch. See [streaming](/docs/streaming). ## Retry-After `Retry-After` is the only rate-limit header Tokens sends. It is in seconds, on 429 responses. There are no `X-RateLimit-*` headers. Its meaning depends on the code: | Code | `Retry-After` | Retry? | | --------------------- | ---------------------------------------------- | --------------------------------------- | | `rate_limited` | Seconds until the next minute starts (up to 60) | Yes | | `concurrency_limit` | 2 | Yes, after reducing parallelism | | `rate_limit_exceeded` | Passed through from the provider, if present | Yes, with backoff | | `window_exhausted` | Seconds until the window resets, often hours | No. Pause or tell the user | | `model_limit_reached` | Seconds until the billing period resets | No. Use another model or wait | Wait at least that long, and add a little random jitter so many clients do not wake together. Always read the `code` first, then the header. ## Stay under the concurrency limit Requests per minute and requests in flight are limits on your whole account, shared by all keys. A streaming request holds a slot until it finishes. If your service fans out, cap parallelism with a semaphore or queue sized a little below your plan's concurrency limit, instead of sending everything and retrying the 429s. Creating more keys does not raise the limits. Abandoned streams hold a slot until they end, so close them. ## Request ids Log `x-tokens-request-id` for every call, on success and on failure. It is the one value support needs to find your request. Send your own `x-request-id` as well, so a line in your logs can be matched to Tokens' id. If a call times out on your side, you have no response and no Tokens id, so log your own id before the call. The full guide is in [request ids and debugging](/docs/request-ids-and-debugging). ## Alerts and spend - **Turn on usage alerts** under Notifications in the dashboard: warnings at 50, 75 and 90 percent of a plan limit, and a low-balance email when the wallet drops under $5. Each is sent once per threshold. See [usage and alerts](/docs/usage-and-alerts). - **Turn on the renewal reminder.** Plans do not renew by themselves. A plan that quietly expired looks exactly like a 402. - **Check before long jobs.** `GET /v1/tokens/usage` returns your windows, your wallet balance and the key's cap. It is not billed and does not count toward the per-minute limit, so a batch job can check it before it starts and between steps. - **Watch the usage page after each release.** A change in prompt size, a retry bug or a new agent shows up as a jump in tokens or cost per request. See [Usage](/dashboard/usage). ```bash curl -s https://tokens.bd/v1/tokens/usage -H "Authorization: Bearer $TOKENS_API_KEY" \ | jq '{wallet: .wallet.balanceUsd, windows: [.windows[] | {type, percentUsed, resetsAt}]}' ``` ## Handle 402 and 429 **402 (`insufficient_credits`, `no_funding`, `outstanding_debt`, `member_cap_reached`).** Your plan or wallet cannot pay for the request. Retrying cannot fix it, and each rejected request still counts against your per-minute limit. In your code: 1. Stop sending to that key. Open a circuit so the rest of your service does not keep trying. 2. Alert the person who can top up, with the request id. 3. Tell your own users something short, such as "the service is temporarily unavailable". Do not show them a billing message. 4. After a top-up in [billing](/dashboard/billing), close the circuit. **429.** Branch on `code`: - `rate_limited`, `concurrency_limit`, `rate_limit_exceeded`: wait `Retry-After`, then retry. If it happens often, you are over capacity: queue requests or cap parallelism. - `window_exhausted`: your plan's 5-hour, weekly or monthly usage is used up. Do not loop. Show the reset time, pause the job, or move to a funded wallet if you use one. - `model_limit_reached`: this one model's allowance is used up for the period. Other models still work. Switch to another model or wait. A `monthly_spend_cap_exceeded` (403) is the key's own cap. Waiting seconds does not help. Use another key, or wait for the next month. ## Graceful degradation Decide now what the product does when Tokens, or one model, is not available. - **Have a second model.** Choose one from `GET /v1/models` that is enough for your task, and fall back to it on 5xx after your retries, on `model_not_available` and on `model_limit_reached`. Test it, because models differ in tool calling and context size. See [choosing a model](/docs/choosing-a-model). - **Do not fall back on errors another model will not fix:** 401, 402, 403 and most 400s apply to the request, the key or the account. - **Fail clearly.** Show the user an honest message and a way to try again. Queue background work and run it later. - **Cut the load when constrained.** Shorter prompts, a smaller `max_tokens`, or turning off optional features such as automatic summaries. - **Check the status page.** If many things fail at once, look at [status](/status) before you look at your code. This is the wrapper above with a fallback model chosen from the environment: ```typescript title="answer.ts" import { chat } from "./tokens-client"; type Messages = { role: "system" | "user" | "assistant"; content: string }[]; export async function answer(messages: Messages) { const primary = await chat({ model: "deepseek/deepseek-v4.1-flash", messages, max_tokens: 800 }); if (primary.ok) return primary; const fallbackModel = process.env.FALLBACK_MODEL; const worthFallingBack = primary.status === 0 || primary.status >= 500 || primary.code === "model_not_available" || primary.code === "model_limit_reached"; if (!fallbackModel || !worthFallingBack) return primary; return chat({ model: fallbackModel, messages, max_tokens: 800 }); } ``` ## Rotate keys Plan rotation before you need it, and try it once while nothing is on fire. - **Rotate on a schedule and after any leak or staff change.** A leaked key is revoked or rotated immediately. - **Rotating a key replaces its secret at once.** The old secret stops working with no grace period, and every client still using it gets 401 `invalid_api_key`. The key keeps its name, cap, allowed models and history. - **For no downtime, create a second key first.** Deploy it, check that traffic uses it on the [usage page](/dashboard/usage), then revoke the old key. This needs a free key slot on your plan. - **Keep the secret in one place** (a secret manager or your platform's environment settings) so that a change is one update plus a restart, not a search through repositories. - **Test the failure.** In staging, revoke the staging key and confirm that your service alerts and recovers once you set the new one. ## Before launch Run through the table at the top. Then send a few test requests with a deliberately wrong key, a model your key is not allowed to use, and a prompt that is too large, and check that your service logs the request id, does not retry, and shows your user a sensible message. Real traffic will find these cases whether or not you do. --- # Request ids and debugging > The request id headers on every Tokens response, how to quote one to support, what to log, how to reproduce a failing call with curl, and how to read the usage page and the error body. Section: API Reference. Page: https://tokens.bd/docs/request-ids-and-debugging When a request fails, or costs something you did not expect, you need to point at that one call. On Tokens, every response carries a request id, and every request that was billed gets a row on the usage page with the same id. This page shows where the id is, what to log around it, how to turn a failure from your app into a curl command anyone can run, and how to read the two places that tell you what happened: the error body and the usage page. ## The request id headers Every response from `/v1`, success or error, carries these headers: | Header | Value | | --------------------- | ----------------------------------------------------------------------------------------------- | | `x-tokens-request-id` | The gateway's id for this request (a UUID). Always generated by Tokens. This is the one to quote. | | `x-request-id` | The `x-request-id` you sent, unchanged, or the same id as above if you sent none | | `x-trace-id` | On successful inference responses: the same value as `x-tokens-request-id` | Error bodies repeat the id as `request_id`. On `/v1/chat/completions`, `/v1/responses` and the other OpenAI-format endpoints it is inside `error`. On `/v1/messages`, which uses Anthropic's error format, it is at the top level of the body, next to `error`: ```json { "error": { "message": "Prepaid wallet balance is insufficient for this request.", "type": "insufficient_quota", "code": "insufficient_credits", "param": null, "request_id": "8f0c7a4e-2b1d-4c55-9a51-3f7e2d9b6c10" } } ``` ```json { "type": "error", "error": { "type": "billing_error", "message": "Prepaid wallet balance is insufficient for this request.", "code": "insufficient_credits" }, "request_id": "8f0c7a4e-2b1d-4c55-9a51-3f7e2d9b6c10" } ``` Some things to know: - **Tokens does not pass provider headers through.** Only the content type, `cache-control`, `retry-after`, `x-request-id` and `x-tokens-*` headers reach you. A request id from the model's own provider never appears, so quote the Tokens id. - **The id arrives with the headers.** For a streamed answer it is available before the first token. Log it at the start of the call, so you have it if the stream dies later. - **Your own `x-request-id` is echoed back as is.** Use it to match a line in your logs to Tokens' id. Send a fresh one for each attempt, not one per user action, so a retry and its first try can be told apart. - **A timeout on your side leaves you with no Tokens id,** because no response arrived. That is why you also log your own id before sending. - **SDK helpers can show your id, not ours.** Some SDKs expose a `request_id` read from `x-request-id`. If you send your own `x-request-id`, that value is yours. Read `x-tokens-request-id` from the raw headers to get the gateway's id. ## Read the id in your code :::code-tabs ```bash title="cURL" curl -sS -i https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"deepseek/deepseek-v4.1-flash","messages":[{"role":"user","content":"ping"}],"max_tokens":20}' \ | grep -i "^x-tokens-request-id" ``` ```python title="Python" import os import openai from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) try: raw = client.chat.completions.with_raw_response.create( model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": "ping"}], max_tokens=20, ) print("request id:", raw.headers.get("x-tokens-request-id")) completion = raw.parse() except openai.APIStatusError as e: print("failed:", e.status_code, e.code, e.response.headers.get("x-tokens-request-id")) raise ``` ```typescript title="Node.js" import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY, }); try { const { data, response } = await client.chat.completions .create({ model: "deepseek/deepseek-v4.1-flash", messages: [{ role: "user", content: "ping" }], max_tokens: 20, }) .withResponse(); console.log("request id:", response.headers.get("x-tokens-request-id")); console.log(data.choices[0]?.message.content); } catch (err) { if (err instanceof OpenAI.APIError) { const headers = err.headers as Headers | Record | undefined; const id = headers instanceof Headers ? headers.get("x-tokens-request-id") : headers?.["x-tokens-request-id"]; console.error("failed:", err.status, err.code, id); } throw err; } ``` ::: With the Anthropic SDKs the same headers are on the raw response; each SDK documents how to reach it. For a `fetch` or `requests` call, read the response headers directly. ## What to log One structured line per attempt is enough. Log it when the response headers arrive, and again if something fails later. | Field | Why | | --------------------------- | ---------------------------------------------------------------- | | `x-tokens-request-id` | Lets support find the request | | Your own request id | Ties the call to your own logs and to a retry sequence | | Time, with time zone (UTC) | Fallback when an id is missing | | Endpoint and model | `chat/completions`, `messages`, and the exact model id | | HTTP status and `error.code` | What happened. Branch on the code, not the message | | Latency, and time to first token for streams | Separates slow providers from slow clients | | `usage` from the response | Tokens in, cached tokens, tokens out, so you can explain a cost | | Attempt number | Shows retries | | Which key (name, never the secret) | Finds the right key when you have several | Do not log: - **The API key,** in any form. Not in headers, not in a dumped config. - **Full prompts and answers by default.** They can contain customer data. Tokens itself stores usage metadata by request id, not prompt content, so the id is what finds a request. If you log bodies to debug, do it with a short retention and redaction. ## Quote an id to support When you open a ticket in [support](/dashboard/support), include: 1. The `x-tokens-request-id`, or several if it is a pattern. 2. When it happened, with the time zone. 3. The endpoint and the model id. 4. The HTTP status and `error.code`, or, for a billing question, what you expected to be charged. 5. Whether it is reproducible, and the curl command if you have one (next section). Never include the API key. The id lets support find usage metadata for your call. [Getting help](/docs/support) explains the rest of the process, and the [status page](/status) shows whether something is broken for everyone. ## Reproduce a failing request with curl If a request fails in your app, take the app out of the picture. A curl command that fails the same way settles whether the problem is your code, your configuration, the key or the account. 1. **Capture the exact body** your app sent: the JSON for `model`, `messages`, `max_tokens`, tools and the rest. Remove customer text you do not want to share. 2. **Save it to a file** and send it with the same key and the same endpoint. 3. **Save headers and body separately** so the id and the error survive. :::code-tabs ```bash title="Chat completions" curl -sS -D headers.txt -o response.json -w "HTTP %{http_code}\n" \ https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -H "x-request-id: repro-$(date +%s)" \ -d @body.json grep -i "^x-tokens-request-id" headers.txt jq '.error // .' response.json ``` ```bash title="Messages" curl -sS -D headers.txt -o response.json -w "HTTP %{http_code}\n" \ https://tokens.bd/v1/messages \ -H "x-api-key: $TOKENS_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -H "x-request-id: repro-$(date +%s)" \ -d @body.json grep -i "^x-tokens-request-id" headers.txt jq '.error // .' response.json ``` ```powershell title="Windows PowerShell" curl.exe -sS -D headers.txt -o response.json -w "HTTP %{http_code}`n" ` https://tokens.bd/v1/chat/completions ` -H "Authorization: Bearer $env:TOKENS_API_KEY" ` -H "Content-Type: application/json" ` -d "@body.json" Select-String -Path headers.txt -Pattern "x-tokens-request-id" Get-Content response.json ``` ::: In Windows PowerShell 5.1 `curl` is an alias for another command, so use `curl.exe`. To watch a streamed answer arrive, add `-N` and `"stream": true` to the body. Then narrow it down, changing one thing at a time: | If the curl call... | It points to | | --------------------------------- | -------------------------------------------------------------------------------------------- | | Fails the same way | The request itself, the key or the account. Read `error.code` and look it up in [errors](/docs/errors) | | Works | Your app: a different key in the environment, a changed body, a proxy, a timeout, or a library default | | Works without `stream`, fails with it | A buffering proxy, or a client that does not handle SSE. See [streaming](/docs/streaming) | | Fails only with a certain model | That model's parameters or availability. Try another from `GET /v1/models` | | Fails only with tools or a large body | The tool schema, the 10 MB body limit, or the model's context window | | Fails only some of the time | A rate limit, or a provider problem. Check `Retry-After` and the [status page](/status) | A quick sanity check that the key and the base URL are right is `GET /v1/models`. It does not run a model and does not count toward your per-minute limit. ## Read the error body Branch on `error.code`. Messages are for people and can change. The shape and the codes are in [errors](/docs/errors). A short way to read one: | Part | Read it as | | ------------------- | ---------------------------------------------------------------------------------------------- | | HTTP status | The class: 4xx is about the request, key or account, 5xx is the gateway or a provider | | `error.code` | The exact reason. `insufficient_credits`, `window_exhausted`, `model_not_found` and so on | | `error.message` | Context: for example the list of models a key may use, or when a window resets | | `request_id` | What you quote to support | | `Retry-After` header | How long to wait, in seconds, on a 429 | Two behaviors that confuse people: - **Provider errors are generic.** When the model's provider rejects or fails a request, the message is replaced with a generic one so provider internals do not leak. The status and `code` still say what kind of failure it was. For a 400 from a provider, look at your parameters and at whether the model supports them. - **A stream can fail after a 200.** Once streaming has started the status is already 200. If the provider fails midway, the connection closes without `data: [DONE]` (or without `message_stop` on Messages), and there is no error body. Treat a stream with no finish reason as incomplete, and use the request id you logged at the start. ## Read the usage page [Usage](/dashboard/usage) lists your recent requests, newest first, in an activity table: | Column | What it shows | | -------- | ---------------------------------------------------------------------------------------------------------- | | Time | When the request was recorded | | Usage | The cost of the request | | Tokens | `input · N cached · output`. The cached part, shown only when there is one, is cache reads and writes together | | Timing | How long the request took | | Model | The model id you called | | Mode | `api` for requests through `/v1` | | Status | `COMPLETED` for a billed request | | Trace ID | The request id. Click it to copy the full value; the table shows only the first eight characters | The Trace ID of an API request is the same value as its `x-tokens-request-id`. To find one request, copy the id from your logs and look for it in this column. The table has pages, so for an older request use the per-page selector at the bottom to show more rows. What this page can and cannot tell you: - **Only billed requests appear.** The gateway records usage when a request has run and been billed. A request that was rejected before running (bad key, no balance, a rate limit) or that failed at the provider has no row. For those, your own log of the request id and the error body is the record. - **A request you cancelled still appears.** If your client disconnected mid-stream, you are billed for the input and the output generated so far, and the row shows those numbers. - **An Estimated badge means no usage was reported.** If the provider sent no token counts, Tokens estimates them from the size of the request and the answer and marks the row. The estimate is what you were billed. - **Cached tokens explain a low cost.** If a long prompt cost little, check the cached part of the Tokens column. See [prompt caching](/docs/prompt-caching). - **A big cost from a small prompt has usual causes:** a long conversation history sent again on every turn, a large `max_tokens` on a model that uses it, reasoning tokens counted as output, retries that were billed more than once, or several requests from a loop. The per-request tokens tell you which. The totals, the daily chart and the plan windows are explained in [usage and alerts](/docs/usage-and-alerts). For a check from code, `GET /v1/tokens/usage` returns your windows, balance and the key's cap without being billed. ## A short debugging routine 1. Read `error.code` and the status. Look the code up in [errors](/docs/errors). 2. Get the `x-tokens-request-id` from the response, or from your log. 3. Reproduce with curl and a saved body. Change one thing at a time. 4. Check [status](/status) if several requests or models fail at once. 5. Look at [Usage](/dashboard/usage) if the question is about cost or tokens. 6. If it is still unexplained, open a ticket with the id, the time, the model, the status and the code. --- # Plans, credits and wallet > How subscription plans, credits, usage windows and the pay-as-you-go wallet work on Tokens, what happens when a window or your balance runs out, and how renewal, cancellation and coupons work. Section: Account & Billing. Page: https://tokens.bd/docs/plans-and-wallet Tokens has two ways to pay for usage: a subscription plan with a credit allowance, and a prepaid wallet for pay-as-you-go. This page explains how plans, credits and the wallet fit together, and exactly what happens when something runs out. Current plans and prices are on [pricing](/pricing). ## Subscription plans vs pay-as-you-go wallet | | Plan | Wallet | | ------------ | ------------------------------------------------ | ----------------------------------------- | | What you buy | A weekly or monthly allowance of credits | A USD balance | | Good for | Steady daily use, predictable cost | Occasional use, or extra on top of a plan | | Limits | Plan's usage windows, rate limit and concurrency | Rate limit and concurrency | | Models | The models included in the plan | Every model sold pay-as-you-go | | Renewal | One payment per period, renewed by you | Top up when you want | You can have both. With a plan, usage is paid from the plan's credits first. ## How credits work Plan allowances are measured in credits. By default **100 credits = 1 USD**; the exact rate can vary by plan, and the plan page shows it. Every request costs what the model's catalog price says for the tokens it used (input, output and cache-read tokens), and that cost is deducted from your credits when the response finishes. Prices per model are on [/models](/models). Plan credits belong to the billing period they were bought for. ## Per-model allowances, free models and deals Some plans give particular models their own allowance, make some models free, or put a model on a deal. The plan's page on [pricing](/pricing) lists them, for example "$10 on DeepSeek V4.1 Flash" or "Free on Ling 3.1 Flash". - **An allowance is a cap inside your plan's usage, not extra on top.** Using a model draws from the plan's credits and from that model's allowance at the same time. With $10 of plan usage and a $10 allowance on DeepSeek V4.1 Flash, spending $5 on DeepSeek leaves $5 of allowance for DeepSeek and $5 of plan usage for every model. If you then spend the other $5 on another model, the plan is empty and DeepSeek stops too, even though its own allowance still shows $5. - **A model without an allowance** uses the plan's credits only, at the price in the [model catalog](/models). - **A free model** costs nothing. It doesn't use your credits or any allowance, and it keeps working after your plan's credits run out. - **A deal** lowers the price of that model, so the same credits go further. Wallet money is never discounted: if a request is paid from your wallet, it costs the full catalog price. - **Allowances reset with your billing period**, the same day your plan's credits do. Your dashboard's Usage page shows how much of each allowance you have used and when it resets, and you get an email when a model reaches 80% and 100% of its allowance. When one model's allowance is used up, requests to that model return `429 model_limit_reached` until the reset, and your other models keep working. On a plan with pay-as-you-go fallback, requests to that model are paid from your wallet instead, at the full price. ```json { "error": { "message": "You've used your $10.00 allowance for DeepSeek V4.1 Flash this billing period. It resets on 4 Nov 2026. Other models on your plan still work. See your plan: https://tokens.bd/dashboard/billing", "type": "rate_limit_error", "code": "model_limit_reached" } } ``` ## Usage windows: 5-hour, weekly and monthly Some plans also spread the allowance over time with usage windows, so that a single heavy day can't burn a month's credits. A plan can have any of these: | Window | How it resets | | -------------- | ------------------------------------------------------------------------------------------------- | | 5-hour session | Starts with your first request after the previous session ended, and resets exactly 5 hours later | | Weekly | Every Monday at 00:00 UTC | | Monthly | At the end of your billing period | A window is measured either in credits or in number of requests, depending on the plan. You can see each window's use and reset time in the dashboard, with `node tokens.mjs usage`, or with `GET /v1/tokens/usage`. See [usage, limits and alerts](/docs/usage-and-alerts). ## What happens when a window or your balance runs out ### A usage window is used up: 429 `window_exhausted` ```json { "error": { "message": "Your plan's session_5h limit has been reached. Resets in 5400s. Upgrade your plan or wait for the window to reset: https://tokens.bd/dashboard/billing", "type": "rate_limit_error", "code": "window_exhausted", "param": null, "request_id": "..." } } ``` The response has a `Retry-After` header with the number of seconds until the window resets. Wallet balance doesn't bypass a plan window: while a window is used up, requests on the plan wait for the reset. Don't retry in a tight loop; `Retry-After` can be hours. ### Plan credits run out Each plan is set either to stop when its credits run out, or to continue from your wallet (pay-as-you-go fallback). On a fallback plan, requests carry on as long as the wallet has money. Otherwise they fail with `402 insufficient_credits` until the next period or a new plan. If you're not sure how your plan behaves, ask [support](/docs/support). ### The wallet is low or empty - **Low balance:** if your balance can't cover the `max_tokens` you asked for, Tokens may lower `max_tokens` to what the balance covers (never below 16). The answer comes back shorter, usually with `finish_reason: "length"`. Top up before long agent sessions. - **Empty:** `402 insufficient_credits`. All 402 messages link to [/dashboard/billing](/dashboard/billing). ### Other billing errors | HTTP | Code | Meaning | | ---- | ---------------------------- | -------------------------------------------------------------------------------- | | 402 | `no_funding` | No active plan and no wallet balance | | 402 | `outstanding_debt` | A previous request left an unpaid amount; top up to clear it | | 429 | `model_limit_reached` | This model's allowance on your plan is used up; other models still work | | 403 | `tier_permission_denied` | Your plan doesn't include this model and the wallet has no balance to pay for it | | 403 | `monthly_spend_cap_exceeded` | The key's own monthly cap, not your account; see [API keys](/docs/api-keys) | The full list is in [errors](/docs/errors). ## Top up the wallet Open [/dashboard/billing](/dashboard/billing) and add funds. The minimum top-up is $5 if you pay in dollars and ৳500 if you pay in taka. The wallet is kept in USD; if you pay in taka, the amount is converted at the exchange rate locked when you start checkout. See [paying in BDT](/docs/paying-in-bdt). The Wallet page in the dashboard shows every credit and debit in a ledger. A low-balance email alert is on by default and fires when the wallet drops below $5. You can switch it off under Notifications in the dashboard. ## Renewal and cancellation - **Renewal is a one-time payment per period.** Nothing is charged automatically. Before the period ends you get a renewal reminder (if that notification is on), and you renew from [/dashboard/billing](/dashboard/billing). - **Cancelling** stops the plan at the end of the current period. You keep the plan and its remaining credits until then. - **Upgrading** is buying a different plan; compare them on [pricing](/pricing). ## Coupons and referrals If you have a coupon code, enter it at plan checkout. The discount is shown before you pay. The referral program gives you a link in the form `/r/` under Referrals in the dashboard. When people you refer pay, a commission is credited to your wallet. Current terms are shown on the Referrals page. ## Related - [Paying in BDT](/docs/paying-in-bdt) - [Usage, limits and alerts](/docs/usage-and-alerts) - [Errors](/docs/errors) --- # Paying in BDT > Pay for Tokens plans and wallet top-ups in Bangladeshi taka: available payment methods, how the exchange rate is locked at checkout, the minimum top-up, receipts and refunds. Section: Account & Billing. Page: https://tokens.bd/docs/paying-in-bdt You can pay for Tokens in Bangladeshi taka (BDT). This page covers how BDT checkout works, how the exchange rate is set, the minimum top-up, receipts and refunds. ## Why paying in BDT matters Most AI providers bill in US dollars and expect an international card. In Bangladesh that usually means a dual-currency card, an annual travel or e-commerce quota, and a foreign transaction fee on top. Tokens takes payment locally in taka and gives you access to models from many providers, so none of that is needed. Prices on Tokens are shown in USD, BDT. The default display currency is BDT. ## Payment methods The methods available to your account right now are: Manual Payment (bKash, Nagad). Only enabled methods appear at checkout, and the list can change. If the method you want isn't listed, check again later or ask [support](/docs/support). ## How BDT checkout works 1. Open [/dashboard/billing](/dashboard/billing). 2. Choose what to buy: a plan (compare them on [pricing](/pricing)) or a wallet top-up. 3. If you're buying a plan and have a coupon, enter it now. The discounted price is shown before you pay. 4. Pick a payment method and complete the payment. 5. Once the payment is confirmed, the plan activates or the wallet is credited. Every payment is confirmed with the payment provider on the server before any credit is added or any plan is activated, and each confirmation is applied once, so a retried or duplicated notification can't credit you twice. Some methods confirm within seconds. Others need a check first, such as paying an invoice or a bank transfer with proof of payment, where those are offered; for these, credit arrives once the payment has been verified. Your payment history in [/dashboard/billing](/dashboard/billing) shows the status of each one. :::tip Don't pay twice because a payment is still pending. Check its status in your payment history first, and open a [support ticket](/dashboard/support) with the transaction reference if it seems stuck. ::: ## How the exchange rate works Your wallet balance and all model prices are kept in US dollars. When you pay in taka, the amount is converted to dollars at the exchange rate **locked at the moment you start checkout**. If the published rate changes while you're paying, your payment still uses the locked rate. The rate is set by Tokens and can change over time. The rate in effect is the one checkout shows you. For BDT payments, the receipt records the locked rate (for example `৳125.00 / $1.00 USD`), so you can always work out what you got for what you paid. :::note The taka you pay is converted once, at checkout. After that, usage is deducted in dollars at each model's catalog price on [/models](/models). Rate changes after your payment don't change your wallet balance. ::: ## Minimum top-up The minimum wallet top-up is **৳500** when you pay in taka and **$5** when you pay in dollars. The two minimums are separate: ৳500 is not converted to check it against $5. If you expect to use Tokens every day, compare the plans on [pricing](/pricing) as well, since plans come with their own credit allowance and usage windows. ## Receipts and invoices Every confirmed payment gets a receipt with a number in the form `REC-XXXXXXXX`. Find them in [/dashboard/billing](/dashboard/billing). Each receipt shows: - what you bought (plan subscription or wallet deposit) - the amount paid and its currency - the credit you received in USD - the locked exchange rate, for BDT payments Receipts are printable. Use your browser's print dialog to save one as PDF for your accounts. A receipt email is also sent if billing receipts are switched on under Notifications, which they are by default. If your payment method works through an invoice, the invoice comes from the billing portal you're sent to at checkout, and the Tokens receipt is issued once it's paid. ## Refunds Refunds are handled by the Tokens team, case by case, under the [refund policy](/refund-policy). To ask for one: 1. Open a ticket in [/dashboard/support](/dashboard/support). 2. Include the receipt number (`REC-...`), the payment method and the reason. Please read the refund policy before buying a plan, especially for usage that has already been consumed. ## Common questions ### Do I need a dual-currency or international card? No. Pay with any of the methods listed above. Card payments only apply if a card method is listed for your account. ### Can I pay for a plan in BDT and top up the wallet in USD? Yes, if both currencies are available at checkout. Both end up in the same place: plan credits and wallet balance are counted in dollars. ### My payment went through but my balance didn't change. The payment may still be waiting for confirmation. Check the status in your payment history. If it says confirmed and your balance hasn't moved, open a support ticket with the receipt number or transaction reference; see [getting help](/docs/support). ## Related - [Plans, credits and wallet](/docs/plans-and-wallet) - [Usage, limits and alerts](/docs/usage-and-alerts) - [Getting help](/docs/support) --- # Usage, limits and alerts > Track spend and remaining allowance on Tokens from the dashboard, a CSV export, the GET /v1/tokens/usage endpoint or the CLI, set up usage alerts, and handle rate limits and Retry-After correctly. Section: Account & Billing. Page: https://tokens.bd/docs/usage-and-alerts Coding agents can spend a lot in one long session, so it pays to know where your usage stands before you start, not after. This page covers the four ways to check usage on Tokens, the email alerts, and the rate limits that apply to every account. ## Read the usage dashboard The Usage page in the dashboard shows: - **Totals** for spend and number of requests, plus your recent burn rate. - **Usage limits**: each usage window on your plan, how much of it is used and when it resets. - **A chart** of daily spend and requests over the last 7, 14 or 30 days, with a breakdown sorted by spend. - **Recent activity**: the latest requests with model, tokens and cost. The dashboard Overview has a shorter version of the same numbers, and the Wallet page lists every top-up and charge. ## Export usage as CSV **Export CSV** on the Usage page downloads your request records for the last 30 days, one row per request: | Column | Meaning | | ---------------------------------------------- | ---------------------------------------------- | | Record ID | Internal id of the usage record | | Request ID | The `x-tokens-request-id` of that request | | Date | When it happened (UTC, ISO 8601) | | Model | The model id you called | | Input Tokens, Output Tokens, Cache Read Tokens | Token counts for the request | | Cost (USD) | What it cost, to six decimal places | | Source | Where it came from, such as `v1` for API calls | The Request ID column is useful when you need to ask [support](/docs/support) about one specific call. ## Check usage from code with GET /v1/tokens/usage `GET /v1/tokens/usage` returns your plan, usage windows, wallet balance and the calling key's limits. It uses the same API key as inference and is not metered, so you can poll it from scripts, status bars or CI. ```bash curl -s https://tokens.bd/v1/tokens/usage \ -H "Authorization: Bearer $TOKENS_API_KEY" ``` An example response (values are illustrative): ```json { "object": "tokens.usage", "plan": { "name": "Example Plan", "tier": "monthly", "periodEnd": "2026-10-31T00:00:00.000Z" }, "windows": [ { "type": "session_5h", "label": "5-Hour Session", "unit": "usd", "limit": 5, "used": 1.85, "remaining": 3.15, "percentUsed": 37, "resetsAt": "2026-10-03T14:20:00.000Z" }, { "type": "weekly", "label": "Weekly Ceiling", "unit": "usd", "limit": 25, "used": 9.4, "remaining": 15.6, "percentUsed": 38, "resetsAt": "2026-10-05T00:00:00.000Z" } ], "wallet": { "balanceUsd": 12.5 }, "key": { "monthlySpendCapUsd": 20, "allowedModels": null } } ``` | Field | Meaning | | ------------------------ | --------------------------------------------------------------------------------- | | `plan` | Your active plan, or `null` if you're pay-as-you-go only | | `windows[].type` | `session_5h`, `weekly` or `monthly` | | `windows[].unit` | `usd` for credit-based windows (in dollars), `requests` for request-count windows | | `windows[].percentUsed` | Rounded percentage of the window used | | `windows[].resetsAt` | When the window resets (UTC) | | `wallet.balanceUsd` | Wallet balance in USD | | `key.monthlySpendCapUsd` | This key's monthly cap, or `null` if it has none | | `key.allowedModels` | This key's allowed models, or `null` if it can use all of yours | To print just the windows with `jq`: ```bash curl -s https://tokens.bd/v1/tokens/usage -H "Authorization: Bearer $TOKENS_API_KEY" \ | jq -r '.windows[] | "\(.label): \(.percentUsed)% (resets \(.resetsAt))"' ``` ## Check usage with the Tokens CLI If you set up your agents with the [Tokens CLI](/docs/tokens-cli), it reads the same endpoint: ```bash node tokens.mjs usage node tokens.mjs usage --json ``` The first prints your plan, a progress bar for each window with its reset time, your wallet balance and the key's cap. `--json` prints the raw response shown above. ## Set up usage alerts Under Notifications in the dashboard you can choose which emails you get: - **Usage warnings** at 50%, 75% and 90% of a plan limit (each threshold can be switched on or off), and when a limit is reached. Each alert is sent once per threshold, not on every request. - **Low balance**, when your wallet drops below $5. On by default. - **Renewal reminders** before your plan period ends. Plans don't renew automatically, so leave this on. - **Billing receipts** and **payment failures**. :::tip Before you leave an agent running unattended, check two things: that usage warnings are on, and that the key it uses has a monthly spend cap. A cap is set when you create a key; see [API keys](/docs/api-keys). ::: ## Rate limits and concurrency Two limits apply to every account, separately from usage windows and credits: | Limit | Default | Error | | --------------------------------------- | ----------------------------------------------------------------------------------------- | ----------------------- | | Requests per minute, per user | 60 RPM (your plan can set a different value); the dashboard playground has its own 10 RPM | `429 rate_limited` | | Requests in flight at once, per account | Set by your plan; 3 without a plan | `429 concurrency_limit` | Rate limits apply to your account, not to each key. Creating more keys doesn't raise them. Tokens doesn't send `X-RateLimit-*` headers. Use the `Retry-After` header on 429 responses instead: | Code | What `Retry-After` tells you | | --------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | | `rate_limited` | Seconds until the next minute starts | | `concurrency_limit` | 2 seconds | | `model_limit_reached` | Seconds until your billing period ends and the model's allowance resets (can be days) | | `window_exhausted` | Seconds until the usage window resets (can be hours) | | `rate_limit_exceeded` | Comes from the upstream provider, after Tokens' automatic failover had no other source left to try. Back off and retry, or switch models | ## Handle Retry-After in your code Coding agents already retry 429s. In your own code, wait for `Retry-After` on short limits and stop on `window_exhausted`, because sleeping for hours inside a request loop is rarely what you want: ```python title="retry.py" import os import time from openai import OpenAI, RateLimitError client = OpenAI( base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], max_retries=0, # we handle retries below ) def ask(messages, attempts=5): for attempt in range(attempts): try: return client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", messages=messages, ) except RateLimitError as err: if err.code == "window_exhausted": raise # resets in hours; report it instead of sleeping wait = float(err.response.headers.get("retry-after", 2 ** attempt)) time.sleep(wait) raise RuntimeError("Still rate limited after retries") print(ask([{"role": "user", "content": "One-line summary of HTTP 429."}]).choices[0].message.content) ``` To cut concurrency errors, limit how many requests your script sends in parallel to your plan's concurrency limit. More detail is in [rate limits](/docs/rate-limits) and [errors](/docs/errors). ## Related - [Plans, credits and wallet](/docs/plans-and-wallet): what happens when a window or balance runs out - [API keys](/docs/api-keys): monthly spend caps per key - [Models and usage endpoints](/docs/models-and-usage) --- # Security and data privacy > What Tokens stores and doesn't store about your requests, how your prompts reach model providers, and how to secure your account and API keys with MFA, session control and key hygiene. Section: Account & Billing. Page: https://tokens.bd/docs/security-and-privacy If you send proprietary code through an AI gateway, you should know exactly what that gateway keeps. This page sets out what Tokens stores, what it doesn't, where your prompts go, and how to lock down your account. The public summary is on the [security page](/security). ## What Tokens stores about your requests Tokens stores **usage metadata** for each request, because billing needs it: - the model you called - input, output and cache-read token counts - the cost - latency and timestamps - which API key was used - the request id (`x-tokens-request-id`) That is what you see in the usage dashboard and the CSV export, and nothing more. ## What Tokens doesn't store - **Prompt and response content is not stored.** The gateway streams your request to the model provider and meters tokens as they pass. There is no archive of prompts, code or model responses. - **Application logs redact prompt and message fields**, so content doesn't end up in logs by accident. - **Tokens doesn't train models on your data and doesn't sell it.** ## Where your prompts go: upstream providers To answer a request, Tokens forwards your prompt to an upstream provider that serves the model you chose. From that point, **that provider's own data-retention and training policies apply** to its service. Those policies differ by provider and sometimes by model, and Tokens can't change them. Practical consequences: - Choosing a model is also choosing whose policy applies. Check the provider's terms for the model you use before sending sensitive code or personal data. - With automatic failover, a request can be served by another upstream source for the same model if the first one fails. If it matters for your work, [contact support](/docs/support) and ask which providers serve a given model. - Don't send secrets (passwords, private keys, production credentials) in prompts at all. Most coding agents read files from your working directory, so keep `.env` files and credentials out of what the agent can see, or use the agent's ignore settings. ## How Tokens protects the platform - **Upstream provider credentials** are encrypted at rest with AES-256-GCM. - **Your API keys** are stored as hashes. The full secret is shown once, at creation, and can't be shown again. - **Passwords** are handled by the authentication service and stored only as salted hashes. - **Customer data** is protected by row-level security in the database. Staff see only what their role requires. - **Admin access** requires multi-factor authentication, times out after inactivity, and is restricted by role. Sensitive admin actions are written to an audit log. - **Payments** are confirmed with the payment provider on the server before credit is added, and each confirmation is applied exactly once. Tokens never sees or stores your full card number or banking password. ## Secure your account: MFA and sessions Open account security in the dashboard (`/account/security`): ### Turn on two-factor authentication Add an authenticator app (TOTP), such as any app that scans a QR code and shows 6-digit codes. After that, signing in needs your password and a current code. If someone gets your password, they still can't reach your keys or billing. ### Review active sessions The **Active Sessions** list shows where your account is signed in. Revoke any session you don't recognise, or use **sign out everywhere** if you think your account was accessed by someone else. Then change your password and check [/dashboard/keys](/dashboard/keys) for keys you didn't create. If you sign in with Google, your Google account's own security (including its two-step verification) protects that sign-in route as well. ## API key hygiene Your key is the thing most likely to leak, usually through a commit, a screenshot or a shared terminal log. - **One key per tool or machine**, so you can revoke one without breaking the rest. - **Set a monthly spend cap** on every key that runs unattended, and an allowed-models list where a tool only needs one or two models. Both are set at creation; see [API keys](/docs/api-keys). - **Keep keys in environment variables or a secrets manager**, never in committed files. Add `.env` to `.gitignore`. - **Never put a key in browser or mobile code.** Tokens doesn't support browser calls; call it from a server. - **Don't paste keys into support tickets**, chats or issues. Support never needs your key; the request id is enough. - **If a key leaks, rotate or revoke it right away.** Rotation takes effect immediately; the old secret stops working at once. Watch the usage dashboard for unfamiliar models or traffic at odd hours. It's often the first sign a key is being used by someone else. ## Delete your account Account deletion is handled by support. Open a ticket in [/dashboard/support](/dashboard/support) from the account you want deleted. Revoke your API keys first so nothing keeps running in the meantime. ## Report a vulnerability If you think you've found a security problem in Tokens, report it privately before disclosing it. The [security page](/security) has the contact details. Include what you found, how to reproduce it and how to reach you. Please don't access other customers' data or disrupt the service while testing. ## Related - [API keys](/docs/api-keys) - [Usage, limits and alerts](/docs/usage-and-alerts) - [Getting help](/docs/support) --- # Getting help > How to get help with Tokens: check the status page, find the request id, and open a support ticket in the dashboard with the details that get it solved fast. Section: Account & Billing. Page: https://tokens.bd/docs/support When something breaks, the fastest fix usually comes from three places, in this order: the status page, the troubleshooting docs, and a support ticket with the right details. This page explains each and shows how to find the request id that lets support trace your exact call. ## Check the status page first [/status](/status) shows the current state of each part of the platform: - **Inference API**: the `/v1` endpoints - **Dashboard & Auth**: signing in and the dashboard - **Billing**: checkout and payments - **Upstream providers**: the model providers behind Tokens It also shows 90 days of uptime history, measured by real probes. If a component is degraded, your problem is probably that, and there's no need to open a ticket for it. If one model is failing but others work, try another model until it's resolved. ## Try the troubleshooting docs Most problems have a known cause: - [Errors](/docs/errors) lists every error code, what it means and what to do. - [Troubleshooting](/docs/troubleshooting) covers common setup problems with agents and SDKs. - [FAQ](/docs/faq) answers billing and account questions. - The guide for your tool, such as [Claude Code](/docs/claude-code) or [Codex CLI](/docs/codex-cli), has tool-specific fixes. You can also run a live test from [/dashboard/connect](/dashboard/connect), which sends a short request with your key and shows the result. If that works but your tool doesn't, the problem is in the tool's configuration, not your key or account. ## Find the request id Every response from Tokens, successful or not, has an `x-tokens-request-id` header. Error bodies also include it as `request_id`. With it, support can find your exact request in seconds; without it, they're searching by time and model. :::code-tabs ```bash title="cURL" curl -sS -D - -o /dev/null https://tokens.bd/v1/chat/completions \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"deepseek/deepseek-v4.1-flash","messages":[{"role":"user","content":"ping"}]}' \ | grep -i x-tokens-request-id ``` ```python title="Python" import os from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) raw = client.chat.completions.with_raw_response.create( model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": "ping"}], ) print(raw.headers.get("x-tokens-request-id")) completion = raw.parse() # the normal ChatCompletion object ``` ```js title="Node.js" import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY }); const { data, response } = await client.chat.completions .create({ model: "deepseek/deepseek-v4.1-flash", messages: [{ role: "user", content: "ping" }], }) .withResponse(); console.log(response.headers.get("x-tokens-request-id")); ``` ::: If the request came from a coding agent, you won't see the headers. Instead, find the request in the Usage page's recent activity, or in the CSV export, which has a Request ID column. See [usage, limits and alerts](/docs/usage-and-alerts). ## Open a support ticket Support tickets are built into the dashboard at [/dashboard/support](/dashboard/support). A ticket has a subject, a message and a priority: | Priority | Use it for | | -------- | -------------------------------------------------------------------------------- | | High | You can't make any requests, or a payment was taken and not credited | | Medium | One model or feature is failing, or something is wrong but you have a workaround | | Low | Questions, billing clarifications, feature requests | Replies arrive in the same ticket thread, and you can reply there too. ### What to include A ticket with these details can usually be answered on the first reply: - [ ] The **`x-tokens-request-id`** of a failing request (or several) - [ ] The **model id** you called, exactly as sent - [ ] The **time** it happened, with your time zone - [ ] The **HTTP status and error code**, for example `403 model_not_allowed_on_key`, and the error message - [ ] **What you were using**: the agent or SDK and its version, and the endpoint (`/v1/chat/completions`, `/v1/messages`, `/v1/responses`) - [ ] **What you already tried** For billing questions, include the receipt number (`REC-...`) or the payment's transaction reference instead. :::danger[Never paste your API key] Support doesn't need your key, and a ticket is not a safe place for it. The key's name, or the short prefix shown next to it in [/dashboard/keys](/dashboard/keys), is enough to identify it. If you've already pasted a full key anywhere, rotate it. ::: A good example: ```text Subject: 504 upstream_timeout on long Claude Code sessions Since about 14:30 (UTC+6) on 3 Oct, Claude Code requests fail after ~10 minutes with 504 upstream_timeout. Short requests work. Model: (id copied from /models) Request ids: 3f2c..., 9a71... Claude Code version: (output of `claude --version`) Tried: a different model works; the connection test at /dashboard/connect passes. ``` ## Other things support handles - **Refunds**, under the [refund policy](/refund-policy). Include the receipt number. - **Account deletion**. Open the ticket from the account you want deleted. - **Questions about which providers serve a model**, if data handling matters for your work. See [security and data privacy](/docs/security-and-privacy). --- # Migrate from OpenAI > Move an app that calls the OpenAI API to Tokens: the three settings that change, what stays the same, the endpoints and behaviours that differ, how to test the switch with a capped key and how to roll back. Section: Guides. Page: https://tokens.bd/docs/migrate-from-openai If your app already calls the OpenAI API, moving it to Tokens takes three changes: the base URL, the API key and the model id. The request and response formats for chat completions, streaming and tool calling stay the same, so most code does not change. This page lists what does differ, so you find it in a test and not in production. ## What changes and what stays the same | Setting | OpenAI | Tokens | Where you set it | | ------------ | -------------------------------------------- | ------------------------------------------------ | ------------------------------------------------------------------------ | | Base URL | `https://api.openai.com/v1` (the SDK default) | `https://tokens.bd/v1` | `base_url` (Python) or `baseURL` (Node.js), or the `OPENAI_BASE_URL` variable | | API key | `sk-...` | `tok_live_...` from [API keys](/docs/api-keys) | `api_key` or `apiKey`, or the `OPENAI_API_KEY` variable | | Model id | OpenAI's own id | An alias in `provider/model` form from `/models` | The `model` field of every request | | Org and project | `OpenAI-Organization`, `OpenAI-Project` headers | Not used. Remove them. | Client options | The official OpenAI Python and Node.js SDKs read `OPENAI_API_KEY` and `OPENAI_BASE_URL` when you do not pass the values in code (checked in the SDK sources, October 2026). Setting both variables switches the endpoint without a code change. You still change the model id in code or config. What stays the same: - The request and response JSON of `POST /v1/chat/completions`, including `messages`, `tools`, `tool_choice`, `response_format`, `stream` and the `usage` object. See [Chat Completions](/docs/chat-completions). - Server-Sent Events streaming. The usage chunk at the end appears only when you send `stream_options: {"include_usage": true}`, as with OpenAI. - `Authorization: Bearer ` authentication. - The Responses API at `POST /v1/responses` (see below). - The OpenAI SDK classes and error types. An error status raises the same exception it would with OpenAI. ## Before and after :::code-tabs ```diff title="Python" import os from openai import OpenAI -client = OpenAI(api_key=os.environ["OPENAI_API_KEY"]) +client = OpenAI( + base_url="https://tokens.bd/v1", + api_key=os.environ["TOKENS_API_KEY"], +) resp = client.chat.completions.create( - model="your-openai-model", + model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": "What does HTTP 429 mean?"}], max_tokens=300, ) print(resp.choices[0].message.content) ``` ```diff title="Node.js" import OpenAI from "openai"; -const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY }); +const client = new OpenAI({ + baseURL: "https://tokens.bd/v1", + apiKey: process.env.TOKENS_API_KEY, +}); const resp = await client.chat.completions.create({ - model: "your-openai-model", + model: "deepseek/deepseek-v4.1-flash", messages: [{ role: "user", content: "What does HTTP 429 mean?" }], max_tokens: 300, }); console.log(resp.choices[0].message.content); ``` ```diff title="curl" -curl https://api.openai.com/v1/chat/completions \ - -H "Authorization: Bearer $OPENAI_API_KEY" \ +curl https://tokens.bd/v1/chat/completions \ + -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ - "model": "your-openai-model", + "model": "deepseek/deepseek-v4.1-flash", "messages": [{"role": "user", "content": "What does HTTP 429 mean?"}], "max_tokens": 300 }' ``` ::: Keep the `Content-Type: application/json` header when you use curl or a bare HTTP client. Without it Tokens does not read the body and answers 400 `invalid_request` ("must specify a 'model' field"). The SDKs set it for you. ### Choose the model id Do not translate OpenAI's id by hand. Tokens ids are aliases that follow `provider/model`, and they do not always match the provider's own id. List what your key can call, then copy the id: ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` The list is filtered for the key: a key with an allowed-models list, or an account without a plan or wallet balance, sees fewer models. Prices, context windows and capabilities are on [/models](/models), not in the API response. [Choosing a model](/docs/choosing-a-model) helps you pick one. A different model gives different answers, so the switch is also a model change. Test your prompts, not just the connection. ## Endpoints | OpenAI endpoint | On Tokens | | ---------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- | | `POST /v1/chat/completions` | Supported. | | `POST /v1/responses` | Supported where the provider behind the model implements it. See [Responses API](/docs/responses). | | `POST /v1/completions` (legacy) | Supported where the provider implements it. Many chat models do not. | | `POST /v1/embeddings` | Only for catalog models that are embedding models. | | `GET /v1/models` | Supported, filtered for the calling key. | | `GET /v1/models/{id}` (`models.retrieve`) | Not supported: 404 `unsupported_endpoint`. Call the list and filter it. | | Images, audio, files, uploads, batches, fine-tuning, moderations | Not supported: 404 `unsupported_endpoint`. | | Assistants, threads, runs | Not supported: 404 `unsupported_endpoint`. OpenAI retired the Assistants API on 26 August 2026 and points to the Responses API. | | Realtime API | Not supported. | Any path outside the supported list returns 404 with code `unsupported_endpoint`. If your app uses one of these endpoints, keep that part of the app on OpenAI and move only the text calls. Use two clients, each with its own base URL and key. Tokens adds two endpoints OpenAI does not have: `GET /v1/tokens/usage` (plan windows, wallet balance and key limits, see [Models and usage](/docs/models-and-usage)) and the Anthropic-style `POST /v1/messages` ([Messages](/docs/messages)). ### The Responses API If your code uses `client.responses.create`, it works against the Tokens base URL with the same change. Two cautions from the [Responses API page](/docs/responses): - Tokens does not store prompts or responses. Do not rely on `store: true` or `previous_response_id` to keep conversation state. Send the full conversation in `input` on every call. - Hosted tools such as web search and file search are provider features. Do not assume they work through the gateway. Test them first. A model that is served only by a provider that speaks Anthropic's Messages protocol answers chat completions (Tokens translates the request) but not `/v1/responses`, `/v1/completions` or `/v1/embeddings`. Those calls fail with 400 `endpoint_not_supported_for_model`. Use chat completions for such a model. ## Differences that can bite ### Rate limits are per account, and lower by default OpenAI limits requests and tokens per minute by organization and project, and returns `x-ratelimit-*` headers. Tokens limits requests per minute per account (60 by default, or your plan's value) and concurrent requests per account (10 with a plan, 3 without). The [rate limits](/docs/rate-limits) page lists no tokens-per-minute limit. Extra keys do not raise either limit, because both are per account. - A 429 carries `Retry-After` in seconds. There are no `x-ratelimit-*` headers, so code that reads them gets nothing. Poll `GET /v1/tokens/usage` to see what is left in a plan window. - A parallel job that was fine on OpenAI can hit `concurrency_limit`. Cap the parallelism on your side, for example with a semaphore. - `window_exhausted` and `model_limit_reached` can carry a `Retry-After` of hours or days. Do not retry those in a loop. The SDKs retry 429 twice by default; pass `max_retries=0` (Python) or `maxRetries: 0` (Node.js) if you handle retries yourself. The backoff example in [rate limits](/docs/rate-limits) does this. ### Errors have the same shape and different codes Gateway errors use OpenAI's JSON shape: `error.message`, `error.type`, `error.code`, `error.param` and an added `error.request_id`. Branch on `error.code`. The ones that differ from what OpenAI code usually expects: | Situation | OpenAI | Tokens | | ---------------- | ------------------------------------------- | ----------------------------------------------------------------------------------------------- | | Out of credit | 429 with a quota or spend-limit code | 402 `insufficient_credits`, `no_funding` or `outstanding_debt`. `type` is `insufficient_quota`. | | Per-minute limit | 429 with `x-ratelimit-*` headers | 429 `rate_limited` with `Retry-After` | | Provider failure | 500, or 503 `server_is_overloaded` | 502 `upstream_unreachable`, 504 `upstream_timeout`, or the provider's 5xx as `upstream_error` | Tokens also has codes that have no OpenAI counterpart: `model_not_found` (404, the id is not in the catalog), `tier_permission_denied` (403, your plan does not include the model), and the key limits `model_not_allowed_on_key` and `monthly_spend_cap_exceeded` (403). If your code treats every 429 as "retry later" and every quota problem as a 429, it will retry 402s, which never succeed. The full list is in [errors](/docs/errors). Retry 429 (except `window_exhausted` and `model_limit_reached`) and 5xx with backoff. Do not retry 400, 401, 402, 403 or 404. Error messages from the provider behind a model are replaced with a generic one, for example "The request was rejected by the upstream provider." A 400 from a model that does not accept a parameter therefore does not tell you which parameter. Check the request against that model's page in [/models](/models). ### Request ids are different OpenAI returns `x-request-id`. Tokens returns `x-tokens-request-id` on every response, and echoes your own `x-request-id` in `x-request-id` if you send one. Keep `x-tokens-request-id` in your logs. [Support](/docs/support) searches by it. Tokens passes on only a short list of provider response headers (`content-type`, `cache-control` and `retry-after`), so provider-specific headers such as rate-limit headers do not reach you. ### Parameters depend on the model Tokens does not validate the body beyond `model` and `n` (1 to 4). `tools`, `response_format`, `reasoning_effort`, `seed`, `logprobs` and similar fields go to the provider behind the model. A model that does not support one ignores it or answers 400. Features you took for granted on OpenAI's models, such as strict structured outputs or image input, depend on the model you choose, so check its page. One rewrite does happen. For OpenAI-style reasoning models (the o-series and GPT-5 and later) on chat completions, the gateway renames `max_tokens` to `max_completion_tokens` and removes `temperature` and `top_p` unless they equal 1, because those models refuse them. ### Output limits and the credit reservation Before it forwards a request, Tokens reserves the worst-case cost, using your `max_tokens` (or 8,192 output tokens if you set none). With a low balance or a key near its cap, a large `max_tokens` can be refused or, on a low balance, lowered to what you can afford (never below 16). Set `max_tokens` to what you need. You pay for the tokens used, not the reservation. See [Chat Completions](/docs/chat-completions). ### Browsers, size and privacy - No CORS headers: calls from a browser fail. Call Tokens from a server. See [authentication](/docs/authentication). - The request body can be up to 10 MB (413 `request_entity_too_large`). Large base64 images count toward it. - Tokens adds one network hop, so latency to the first token is not lower than calling the provider directly. - Prompts reach the provider that serves the model, and that provider's data policy applies. Tokens itself stores usage metadata, not prompt content. See [security and privacy](/docs/security-and-privacy). ### Billing You pay Tokens, in USD or BDT, from a plan or a wallet, at the catalog price of each model. Your OpenAI invoice stops for the traffic you move. See [plans and wallet](/docs/plans-and-wallet). ## Test the switch safely 1. **Create a second key** at [/dashboard/keys](/dashboard/keys). Name it for the test, set a **low monthly spend cap** (for example a few dollars) and an **allowed-models list** of the one or two models you will try. The cap and the list cannot be edited later, so create a new key to change them. Tokens counts the worst-case cost of each request against the cap, so a very large `max_tokens` can be refused near the cap. 2. **Switch by configuration.** Read the base URL, key and model id from environment variables or a config file, so a deploy does not need a code change to move between providers. 3. **Run both for a while.** Send the same prompts to OpenAI and Tokens (a replay of logged requests, or a mirror of live traffic whose Tokens answer you discard) and compare. A small script is enough: ```python import os import time from openai import OpenAI openai_client = OpenAI(api_key=os.environ["OPENAI_API_KEY"]) tokens_client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_TEST_KEY"]) CANDIDATES = [ ("openai", openai_client, "your-openai-model"), ("tokens", tokens_client, "deepseek/deepseek-v4.1-flash"), ] prompt = [{"role": "user", "content": "Write a Python function that parses an ISO 8601 date."}] for name, client, model in CANDIDATES: start = time.perf_counter() raw = client.chat.completions.with_raw_response.create( model=model, messages=prompt, max_tokens=400 ) elapsed = time.perf_counter() - start resp = raw.parse() print(name, resp.choices[0].finish_reason, resp.usage.total_tokens, f"{elapsed:.2f}s") print(" request id:", raw.headers.get("x-tokens-request-id") or raw.headers.get("x-request-id")) ``` 4. **Compare what matters to your app**, not only the text: | Check | How | | ------------------------ | ---------------------------------------------------------------------------------------------------- | | Answer quality | Run your own test prompts or evals on both. Models differ. | | Tool calls | Are the arguments valid JSON for your schema? Does the model call the right tool? | | `finish_reason` | More `length` than before means `max_tokens` is too low for this model. | | Token counts and cost | Token counts differ by model. Compare cost per finished task in [usage](/docs/usage-and-alerts). | | Latency | Time to first token with `stream: true`, from where your app runs. | | Errors | Count them by `error.code`. A 402 or 429 `concurrency_limit` shows a sizing problem. | 5. **Ramp up.** Move a small share of traffic (a feature flag or a percentage), watch for a day or two, then increase it. Replace the test key with a production key that has the cap you want. Create it before you cut over, because [rotating or revoking](/docs/api-keys) takes effect at once. ## Roll back Rolling back is the reverse of the switch, if you kept the way back open: 1. Keep your OpenAI key and its billing active until Tokens has carried production traffic for a full billing cycle. 2. Point the base URL, key and model id back through the same configuration, and redeploy or flip the flag. If you used `OPENAI_BASE_URL`, unset it. 3. Revoke or rotate the Tokens key you no longer use at [/dashboard/keys](/dashboard/keys). Your wallet balance and plan on Tokens stay on your account. For refunds see the [refund policy](/refund-policy). Nothing needs to be exported, because Tokens does not keep prompt content. ## Where to go next - [Chat Completions](/docs/chat-completions), [Responses API](/docs/responses) and [Streaming](/docs/streaming) for the request formats. - [Python](/docs/python) and [Node.js](/docs/nodejs) for SDK setup. - [Errors](/docs/errors) and [Rate limits](/docs/rate-limits) for the full code list and a backoff example. - [Migrate from OpenRouter](/docs/migrate-from-openrouter) and [Migrate from Anthropic](/docs/migrate-from-anthropic). Sources, checked October 2026: OpenAI [API reference overview](https://developers.openai.com/api/reference/overview), [error codes](https://developers.openai.com/api/docs/guides/error-codes), [rate limits](https://developers.openai.com/api/docs/guides/rate-limits), [Assistants migration](https://developers.openai.com/api/docs/assistants/migration), and the [openai-python](https://github.com/openai/openai-python) and [openai-node](https://github.com/openai/openai-node) sources. Tokens behaviour is from the gateway code and the pages linked above. It has not been tested against a live OpenAI account. --- # Migrate from OpenRouter > Move an app from OpenRouter to Tokens: the base URL, key and model ids to change, what happens to OpenRouter routing fields, headers and model suffixes, how errors and limits differ, and how to test and roll back. Section: Guides. Page: https://tokens.bd/docs/migrate-from-openrouter OpenRouter and Tokens both give you one OpenAI-compatible endpoint for models from many providers, so an app written for OpenRouter usually needs the same three changes as any OpenAI app: the base URL, the key and the model id. The work is in what OpenRouter does on top of the OpenAI format: provider routing, model fallbacks, model suffixes, attribution headers and cost fields. Tokens does not do these. This page says exactly what happens to each one. ## What changes and what stays the same | Setting | OpenRouter | Tokens | Where you set it | | -------- | -------------------------------------------------- | --------------------------------------------------- | -------------------------------------- | | Base URL | `https://openrouter.ai/api/v1` | `https://tokens.bd/v1` | `base_url` or `baseURL` in the SDK | | API key | An OpenRouter key | `tok_live_...` from [API keys](/docs/api-keys) | `api_key` or `apiKey` | | Model id | `provider/model`, an OpenRouter slug | `provider/model`, a Tokens alias from `/models` | The `model` field of every request | | Headers | `HTTP-Referer`, `X-Title` (attribution, optional) | Not used. Remove them. | `default_headers` or `defaultHeaders` | What stays the same: - The OpenAI request and response format of `POST /v1/chat/completions`, with `messages`, `tools`, `stream` and the `usage` object. See [Chat Completions](/docs/chat-completions). - `Authorization: Bearer ` authentication, and Server-Sent Events streaming. - Any OpenAI SDK, or the Vercel AI SDK, LangChain and similar libraries you pointed at OpenRouter. Change their base URL and key in the same place. ## Before and after :::code-tabs ```diff title="Python" import os from openai import OpenAI client = OpenAI( - base_url="https://openrouter.ai/api/v1", - api_key=os.environ["OPENROUTER_API_KEY"], - default_headers={ - "HTTP-Referer": "https://example.com", - "X-Title": "My app", - }, + base_url="https://tokens.bd/v1", + api_key=os.environ["TOKENS_API_KEY"], ) resp = client.chat.completions.create( - model="provider/openrouter-model-slug", + model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": "What does HTTP 429 mean?"}], max_tokens=300, ) print(resp.choices[0].message.content) ``` ```diff title="Node.js" import OpenAI from "openai"; const client = new OpenAI({ - baseURL: "https://openrouter.ai/api/v1", - apiKey: process.env.OPENROUTER_API_KEY, - defaultHeaders: { "HTTP-Referer": "https://example.com", "X-Title": "My app" }, + baseURL: "https://tokens.bd/v1", + apiKey: process.env.TOKENS_API_KEY, }); const resp = await client.chat.completions.create({ - model: "provider/openrouter-model-slug", + model: "deepseek/deepseek-v4.1-flash", messages: [{ role: "user", content: "What does HTTP 429 mean?" }], max_tokens: 300, }); console.log(resp.choices[0].message.content); ``` ```diff title="curl" -curl https://openrouter.ai/api/v1/chat/completions \ - -H "Authorization: Bearer $OPENROUTER_API_KEY" \ - -H "HTTP-Referer: https://example.com" \ - -H "X-Title: My app" \ +curl https://tokens.bd/v1/chat/completions \ + -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ - "model": "provider/openrouter-model-slug", + "model": "deepseek/deepseek-v4.1-flash", "messages": [{"role": "user", "content": "What does HTTP 429 mean?"}], "max_tokens": 300 }' ``` ::: Keep `Content-Type: application/json` when you use curl. Without it Tokens does not read the body and answers 400 `invalid_request`. If you pass the settings through the `OPENAI_BASE_URL` and `OPENAI_API_KEY` variables, which the official OpenAI Python and Node.js SDKs read (checked in the SDK sources, October 2026), change the variables and the model id and nothing else. ## Model ids look the same and are not OpenRouter slugs and Tokens aliases both use `provider/model`. They are separate lists. A slug that works on OpenRouter can be unknown on Tokens, or can name a different version, and Tokens aliases do not always match the provider's own id. Do not rewrite ids by pattern. 1. List what your key can call: `curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY"`. Each `data[].id` is valid for that key. 2. Prices, context windows and capabilities are on [/models](/models). The API list has ids only. See [Choosing a model](/docs/choosing-a-model). 3. Keep a table from your old ids to the new ones in configuration, so the mapping is not scattered in code. The `model` value must match an alias exactly. An unknown id returns 404 `model_not_found`. ## OpenRouter features, one by one This is what Tokens does with each OpenRouter feature. It comes from the gateway code, not from guesses. **Model suffixes.** OpenRouter documents `:free`, `:nitro`, `:floor`, `:exacto` (and the deprecated `:online`, `:thinking`, `:extended`) as part of the model id. Tokens looks up the whole string as the alias, so `some/model:nitro` returns 404 `model_not_found`. Remove the suffix. There is no equivalent: Tokens has no suffixes and no per-request provider sorting. A cheaper or faster model is a different model id. **Router models.** `openrouter/auto` and other OpenRouter router ids are not in the Tokens catalog (404 `model_not_found`). Pick a fixed model. **Provider routing (`provider`).** Tokens does not read this field. It does not reject it either: for a model whose provider speaks the OpenAI protocol (the usual case), the rest of the request body is passed to the provider as sent. Whether the provider ignores the field or answers 400 is up to the provider, and Tokens did not check any provider's behaviour. Remove it. Which provider serves a model is Tokens' choice, not yours, and the `model` field of the response is always the id you asked for. **Fallbacks (`models`, `route`).** Not implemented. The same rule as above: forwarded, not acted on, so the field does nothing. Tokens does fail over, but only between sources of the same model: on 429, 502, 503, 504 or a dropped connection, and before any byte of the answer has reached you, it retries on another source of that model. It never switches to a different model. If you want a model fallback, do it in your code: ```python import os import openai from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], max_retries=0) MODELS = ["deepseek/deepseek-v4.1-flash", "your-second-model-id"] # ids from GET /v1/models def ask(messages): last = None for model in MODELS: try: return client.chat.completions.create(model=model, messages=messages, max_tokens=512) except openai.APIStatusError as e: if e.status_code not in (429, 500, 502, 503, 504) or e.code == "window_exhausted": raise last = e raise last ``` `window_exhausted` is excluded because it is an account-wide plan window: another model will hit it too. **Prompt transforms and plugins (`transforms`, `plugins`).** OpenRouter lets you compress long prompts, parse files, search the web and repair responses through these. Tokens does none of it. The fields are forwarded to the provider as unknown body fields, like `provider`. Trim the prompt yourself: Tokens does not shorten a prompt that is over the model's context window, so the provider decides what happens to it. **`reasoning` and `usage` fields.** Forwarded like the others. Where a model supports reasoning, the provider decides what it does with them. The `usage: {"include": true}` request field is not needed: OpenRouter's documentation calls it deprecated and returns usage anyway, and Tokens' answers always carry the `usage` object the provider sent. **Attribution headers.** `HTTP-Referer`, `X-Title` and the `X-OpenRouter-*` headers only matter on OpenRouter, for app pages and rankings. Tokens forwards a short allow-list of request headers to providers (`content-type`, `accept`, `openai-beta` and `anthropic-version`, among a few) and drops the rest without an error. Sending them is harmless and has no effect. Remove them so the code does not suggest otherwise. **If you are not sure a field is safe.** Test it against the model you plan to use, on the test key described below. A request that works is the only proof that a provider accepts a field. When a model is served only by a provider that speaks Anthropic's Messages protocol, Tokens translates your chat completions request instead of forwarding it. Only these fields are carried: `messages` (text, images, tool calls and tool results), `max_tokens` or `max_completion_tokens`, `temperature`, `top_p`, `stop`, `stream`, `tools`, `tool_choice`, `parallel_tool_calls` and `user`. Everything else is dropped, including `n`, `response_format`, `seed`, `logprobs` and the penalties. `temperature` is capped at 1, and `max_tokens` defaults to 4096 when you send none. ## Response and usage differences - **`model`** in the response is always the id you asked for, whichever provider answered. OpenRouter's `model` shows the model it routed to, so code that logged it to see the routing outcome gets nothing new. - **Cost fields.** OpenRouter adds `cost`, `cost_details` and `native_finish_reason` to the response. Tokens does not add anything to the provider's body. The cost of each request is in your [usage dashboard](/docs/usage-and-alerts) and the CSV export, priced at the model's catalog price. If you read `usage.cost` to bill your own customers, switch to the export, or compute the cost from the token counts and the price on [/models](/models). - **Usage lookups.** There is no `GET /generation?id=` or `GET /key`. `GET /v1/tokens/usage` returns plan windows, wallet balance and the calling key's cap ([Models and usage](/docs/models-and-usage)). The response header `x-tokens-request-id` is the id to log and to quote to [support](/docs/support). - **Streaming.** The usage chunk at the end of a stream appears when you send `stream_options: {"include_usage": true}`. See [Streaming](/docs/streaming). ## Differences that can bite ### Errors OpenRouter errors are `{"error": {"code": 429, "message": "...", "metadata": {...}}}` with a numeric `code`. Tokens errors are `{"error": {"message", "type", "code", "param", "request_id"}}` with a **string** `code` such as `rate_limited`. Code that reads `error.code` as a number, or reads `error.metadata`, needs a change: use the HTTP status for the class of error and `error.code` for the cause. The full table is in [errors](/docs/errors). | Situation | OpenRouter | Tokens | | ----------------------- | --------------------------------------------- | ------------------------------------------------------------------------ | | Out of credit | 402 | 402 `insufficient_credits`, `no_funding` or `outstanding_debt` | | Rate limited | 429 | 429 `rate_limited`, `concurrency_limit`, `window_exhausted` or `model_limit_reached` | | Key limit you set | Key credit limit: 402 | 403 `monthly_spend_cap_exceeded` (see [API keys](/docs/api-keys)) | | Timeout | 408 | 504 `upstream_timeout` | | Model down | 502 | 502 `upstream_unreachable` or the provider's 5xx as `upstream_error` | | No provider available | 503 | 503 `no_upstream_available` | OpenRouter documents that a failure after streaming starts arrives as an error event in a 200 response. Keep any in-body error check you added for that. The provider's own error message is replaced by a generic one, and `metadata.provider_*` details do not exist. ### Rate limits OpenRouter limits free models to 20 requests per minute and puts no platform cap on paid models. Tokens limits **requests per minute per account** (60 by default, or your plan's value) and **concurrent requests per account** (10 with a plan, 3 without). An agent or batch job that ran with high parallelism on OpenRouter can hit `concurrency_limit`; cap its parallelism. More keys do not raise the limits. - A 429 carries `Retry-After` in seconds. There are no `X-RateLimit-*` headers. Poll `GET /v1/tokens/usage` to see a plan window. - `window_exhausted` and `model_limit_reached` can mean hours or days. Do not retry them in a loop. - See [rate limits](/docs/rate-limits) for the numbers and a backoff example. ### Other differences - **No `/api/v1` path.** Tokens' path is `/v1`: `https://tokens.bd/v1`. - **Unsupported endpoints.** Images, audio, files, batches, assistants, fine-tuning and moderations return 404 `unsupported_endpoint`. Embeddings work only for embedding models. Supported: `/v1/chat/completions`, `/v1/responses`, `/v1/completions` (legacy), `/v1/embeddings`, `/v1/models`, `/v1/messages` and `/v1/messages/count_tokens`. - **Parameters depend on the model.** Besides `model` and `n` (1 to 4), Tokens does not validate the body. Tool calling, `response_format`, vision and reasoning depend on the model. The gateway renames `max_tokens` to `max_completion_tokens` and removes `temperature` and `top_p` unless they equal 1 for OpenAI o-series and GPT-5 and later models on chat completions. - **Output reservation.** Tokens reserves the worst-case cost of a request, using `max_tokens` or 8,192 output tokens if you set none. A large `max_tokens` on a low balance or a nearly full key cap can be refused or lowered. Set it to what you need. - **No CORS.** Calls from a browser fail. Call Tokens from a server. - **Body size.** Up to 10 MB. - **Privacy.** Tokens stores usage metadata, not prompt content. The provider that serves a model sees the prompt and its policy applies. See [security and privacy](/docs/security-and-privacy). - **Payment.** Tokens bills in USD or BDT from a plan or a wallet; there are no OpenRouter credits. See [plans and wallet](/docs/plans-and-wallet). If you call OpenRouter's Anthropic-style Messages endpoint with an Anthropic SDK, read [Migrate from Anthropic](/docs/migrate-from-anthropic) too. Tokens serves `POST /v1/messages`, with a base URL without `/v1`. ## Test the switch safely 1. **Create a second key** at [/dashboard/keys](/dashboard/keys) with a **low monthly spend cap** and an **allowed-models list** that holds only the models you are testing. Neither can be edited afterwards. Tokens counts the worst-case cost of each request against the cap. 2. **Read the base URL, key and model id from configuration**, not from constants, so the switch is a config change. 3. **Run both for a while.** Replay logged requests, or mirror a share of live traffic to Tokens and discard its answers. Compare the same prompts on both: | Check | How | | ---------------------- | ------------------------------------------------------------------------------------------- | | Quality | Run your own prompts or evals. The two models in a pair are rarely the same model. | | Fields you stopped sending | Run once without `provider`, `models`, `transforms` and the headers. Did anything depend on them? | | Tool calls | Valid JSON arguments for your schema, on the model you chose. | | `finish_reason` | More `length` than before means `max_tokens` is too low. | | Cost per finished task | OpenRouter's `usage.cost` against the cost in [usage](/docs/usage-and-alerts). | | Errors | Count them by `error.code`. | 4. **Ramp up** with a feature flag or a percentage. Create the production key with the cap you want before you cut over. ## Roll back 1. Keep your OpenRouter key and credits until Tokens has carried production traffic for a full billing cycle. 2. Put the old base URL, key, model ids and headers back in the configuration and deploy or flip the flag. 3. Revoke the Tokens key you no longer use at [/dashboard/keys](/dashboard/keys). Your wallet balance stays on your account; see the [refund policy](/refund-policy). Because the id mapping lives in configuration, rolling back is the same change in reverse. ## Where to go next - [Chat Completions](/docs/chat-completions), [Streaming](/docs/streaming) and [Tool calling](/docs/tool-calling). - [Errors](/docs/errors) and [Rate limits](/docs/rate-limits). - [Migrate from OpenAI](/docs/migrate-from-openai) for the shared parts of an OpenAI-style switch. Sources, checked October 2026: OpenRouter [API overview](https://openrouter.ai/docs/api-reference/overview), [errors](https://openrouter.ai/docs/api-reference/errors), [limits](https://openrouter.ai/docs/api-reference/limits), [provider routing](https://openrouter.ai/docs/guides/routing/provider-selection), [model fallbacks](https://openrouter.ai/docs/guides/routing/model-fallbacks), [model variants](https://openrouter.ai/docs/guides/routing/model-variants), [usage accounting](https://openrouter.ai/docs/guides/guides/usage-accounting) and [app attribution](https://openrouter.ai/docs/app-attribution). Tokens behaviour is from the gateway code and the pages linked above. It has not been tested against a live OpenRouter account. --- # Migrate from Anthropic > Move an app that calls the Anthropic Messages API to Tokens: base URL without /v1, key header, model ids, which Anthropic-only features pass through and which do not, how errors and limits differ, testing and rollback. Section: Guides. Page: https://tokens.bd/docs/migrate-from-anthropic If your app calls the Anthropic Messages API, moving it to Tokens takes three changes: the base URL, the API key and the model id. Tokens serves `POST /v1/messages`, streaming and `POST /v1/messages/count_tokens` in Anthropic's format, so the Anthropic SDKs and your message-handling code keep working. What needs care is the set of Anthropic-only features (prompt caching, thinking, citations, the Files and Batches APIs), because they depend on which provider serves the model you pick. ## What changes and what stays the same | Setting | Anthropic | Tokens | Where you set it | | ----------------- | ------------------------------------------ | --------------------------------------------------- | ------------------------------------------------------------------- | | Base URL | `https://api.anthropic.com` (SDK default) | `https://tokens.bd`, **without `/v1`** | `base_url` (Python) or `baseURL` (Node.js), or `ANTHROPIC_BASE_URL` | | API key | An Anthropic key | `tok_live_...` from [API keys](/docs/api-keys) | `api_key` or `apiKey`, or `ANTHROPIC_API_KEY` | | Model id | An Anthropic id | An alias in `provider/model` form from `/models` | The `model` field of every request | | `anthropic-workspace-id` | Selects a workspace | Not used. Remove it. | Client options | The SDKs add `/v1/messages` themselves, so the base URL is the bare host. If you set `https://tokens.bd/v1` instead, requests go to `/v1/v1/messages` and fail with 404. Calling the endpoint yourself with curl is different: the full URL is `https://tokens.bd/v1/messages`. The Anthropic Python and TypeScript SDKs read `ANTHROPIC_API_KEY`, `ANTHROPIC_AUTH_TOKEN` and `ANTHROPIC_BASE_URL` when you pass nothing in code (checked in the SDK sources, October 2026). Be careful in a shell where you also run Claude Code or other tools that read the same variables. See the [Anthropic SDK page](/docs/anthropic-sdk) and [Claude Code](/docs/claude-code). What stays the same: - The request and response shapes of `POST /v1/messages`: `system`, `messages`, content blocks, `max_tokens` (still required), `tools`, `tool_choice`, `stop_sequences` and the `usage` object. - Server-Sent Events streaming with the usual event names (`message_start`, `content_block_delta`, `message_stop`). `client.messages.stream(...)` works. - Authentication. Tokens accepts the key in `x-api-key` (what the SDKs send) or in `Authorization: Bearer`, so the SDK needs no auth change. If both arrive, the Bearer header wins. - The `anthropic-version` header. It is forwarded to the provider, and the default `2023-06-01` is used when you send none. - Error bodies on `/v1/messages` in Anthropic's shape, so the SDKs raise the same exception classes. ## Before and after :::code-tabs ```diff title="Python" import os import anthropic -client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"]) +client = anthropic.Anthropic( + base_url="https://tokens.bd", + api_key=os.environ["TOKENS_API_KEY"], +) message = client.messages.create( - model="your-claude-model", + model="deepseek/deepseek-v4.1-flash", max_tokens=512, messages=[{"role": "user", "content": "Explain idempotency keys in two sentences."}], ) print(message.content[0].text) ``` ```diff title="Node.js" import Anthropic from "@anthropic-ai/sdk"; -const client = new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY }); +const client = new Anthropic({ + baseURL: "https://tokens.bd", + apiKey: process.env.TOKENS_API_KEY, +}); const message = await client.messages.create({ - model: "your-claude-model", + model: "deepseek/deepseek-v4.1-flash", max_tokens: 512, messages: [{ role: "user", content: "Explain idempotency keys in two sentences." }], }); console.log(message.content[0].text); ``` ```diff title="curl" -curl https://api.anthropic.com/v1/messages \ - -H "x-api-key: $ANTHROPIC_API_KEY" \ +curl https://tokens.bd/v1/messages \ + -H "x-api-key: $TOKENS_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ - "model": "your-claude-model", + "model": "deepseek/deepseek-v4.1-flash", "max_tokens": 512, "messages": [{"role": "user", "content": "Explain idempotency keys in two sentences."}] }' ``` ::: ### Choose the model id Anthropic's ids and Tokens ids are different lists. Tokens ids are aliases in `provider/model` form, and they do not always match the provider's own id. The Anthropic SDK lists them for you, because `GET /v1/models` answers in Anthropic's format when the request carries `anthropic-version`: ```python for m in client.models.list(): print(m.id, m.display_name) ``` With curl, send `anthropic-version` to get that shape. With only `Authorization: Bearer` you get the OpenAI list shape. In both, the list holds only models this key can call, so a key with an allowed-models list, or an account with no plan or wallet balance, sees fewer. The Anthropic format also works for models from other makers, not only Claude models. That is a feature, but it is also a model change: answers, tool-call behaviour and token counts differ. Prices, context windows and capabilities are on [/models](/models); [Choosing a model](/docs/choosing-a-model) helps you pick. ## Which provider serves the model decides the features Each model on Tokens is served by one or more providers. When a provider speaks the Messages API itself, your request goes to it unchanged except for the model id, so thinking blocks, prompt caching and `anthropic-beta` features come back untouched. Tokens prefers such a provider when one serves the model. When a model is only served by a provider that speaks the OpenAI format, Tokens translates your Messages request into a chat completions request and the answer back. The catalog does not say which case applies to a model, so test every feature you depend on against the exact model you choose. What the translation carries: `system` (text, joined into one system message), `messages` with text, image, `tool_use` and `tool_result` blocks, `max_tokens`, `stop_sequences`, `temperature`, `top_p`, `stream`, `tools` and `tool_choice`. Everything else is not copied. | Anthropic feature | Provider speaks Messages | Translated model | | -------------------------------------- | ------------------------------------------------------- | ----------------------------------------------------------------------------- | | Streaming, client tools, images | Passed through | Supported (translated) | | Prompt caching (`cache_control`) | Passed through. `usage` carries the cache counts the provider reports. | Markers dropped. Caching, if any, is the provider's own and not under your control. | | Extended or adaptive `thinking` | Passed through | Parameter dropped. No thinking blocks. | | `document` blocks and `citations` | Passed through | Document blocks dropped. No citations. | | Anthropic-defined server tools (web search, code execution, computer use) | Passed to the provider as sent. Tokens does not run them, so it is up to the provider. | Every `tools` entry becomes a function tool, so these do not work. | | `anthropic-beta` header | Forwarded | No effect | | `top_k`, `metadata` | Passed through | Dropped | Tokens billing reads the provider's `usage` numbers, and cached input is priced at the model's cache-read rate where the catalog has one. Prompt caching also saves money only on models whose provider supports it; see [Choosing a model](/docs/choosing-a-model) and [Messages](/docs/messages). ## Endpoints and features Tokens does not have Tokens serves these Anthropic endpoints: `POST /v1/messages`, `POST /v1/messages/count_tokens` and `GET /v1/models`. Every other path returns 404 `unsupported_endpoint`, in Anthropic's error shape (`not_found_error`). | Anthropic API | On Tokens | | ---------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- | | Message Batches (`/v1/messages/batches`) | Not supported. Send requests one by one, or run your own queue. | | Files API (`/v1/files`) | Not supported. Send images and documents inline as base64, or as an image URL. A `file_id` in a content block cannot work. | | Skills, Managed Agents, Agents, Sessions | Not supported. | | `GET /v1/models/{id}` (`models.retrieve`) | Not supported. List the models and filter. | | Models API fields `capabilities`, `lifecycle` | Not returned. The list has `type`, `id`, `display_name` and `created_at`, and it has no pagination: `has_more` is `false`. | | `POST /v1/messages/count_tokens` | Supported. See below. | ### Count tokens `POST /v1/messages/count_tokens` takes the body of a Messages request without `max_tokens` and returns `{"input_tokens": N}`. It never runs the model and is never billed. When a provider that serves the model counts tokens natively you get its exact count. Otherwise Tokens estimates (about four characters per token, plus a fixed amount per image and per turn) and sets the response header `x-tokens-estimated: true`. Use the estimate for budgeting only. ## Differences that can bite ### Rate limits Anthropic limits requests, input tokens and output tokens per minute for each model class and reports them in `anthropic-ratelimit-*` headers. Tokens limits **requests per minute per account** (60 by default, or your plan's value) and **concurrent requests per account** (10 with a plan, 3 without). The [rate limits](/docs/rate-limits) page lists no token-per-minute limit. More keys do not raise either limit. - A 429 carries `retry-after` in seconds. There are no `anthropic-ratelimit-*` headers. Poll `GET /v1/tokens/usage` to see a plan window. - Both Anthropic SDKs retry 429 and 5xx twice by default and honour `retry-after`. `window_exhausted` and `model_limit_reached` can mean hours or days, so retrying them is useless. Set `max_retries=0` (`maxRetries: 0` in TypeScript) and handle retries yourself if you want control; the backoff example in [rate limits](/docs/rate-limits) shows how. - A job that fans out many streams in parallel can hit `concurrency_limit` first. Limit its parallelism. ### Errors Errors on `/v1/messages` use Anthropic's shape: `{"type": "error", "error": {"type", "message"}}`. Errors that Tokens raises itself (key, credit, plan and limit errors) add `error.code`, and carry a `request_id`. Read `error.code` to tell causes apart. The full list is in [errors](/docs/errors). | Situation | Anthropic | Tokens | | ------------------------ | ------------------------------------------------- | ----------------------------------------------------------------------------------- | | Out of credit | 402 `billing_error` | 402 `billing_error`, code `insufficient_credits`, `no_funding` or `outstanding_debt` | | A spend limit you set | 400 `invalid_request_error`, or 429 for some workspaces | 403 `permission_error`, code `monthly_spend_cap_exceeded` for a key's cap | | Key problem | 401 `authentication_error`, 403 `permission_error`| Same types, with codes such as `invalid_api_key`, `key_inactive`, `key_expired` | | Rate limited | 429 `rate_limit_error` | 429 `rate_limit_error`, code `rate_limited`, `concurrency_limit` or `window_exhausted` | | Provider failure | 500 `api_error`, 529 `overloaded_error` | 502 `upstream_unreachable`, 503 `no_upstream_available`, 504 `upstream_timeout` | | Request too large | 413 `request_too_large` at 32 MB | 413 at **10 MB** (code `request_entity_too_large`) | A 403 is not retried by the SDKs. If your code treated the spend-limit 400 or 429 as "stop and alert", map the Tokens 403 `monthly_spend_cap_exceeded` to the same behaviour. The provider's own error text is replaced by a generic message, so a 400 for an unsupported parameter does not say which one. Check the request against the model's page on [/models](/models). ### The request id header is different Anthropic returns a `request-id` header, and the SDKs expose it as `_request_id`. Tokens does not return that header, so `_request_id` is `None`. Read `x-tokens-request-id` instead: ```python raw = client.messages.with_raw_response.create( model="deepseek/deepseek-v4.1-flash", max_tokens=64, messages=[{"role": "user", "content": "ping"}], ) print(raw.headers.get("x-tokens-request-id")) message = raw.parse() ``` Tokens passes on only a short list of provider response headers (`content-type`, `cache-control` and `retry-after`), so `anthropic-organization-id`, `anthropic-workspace-id` and the rate-limit headers do not reach you. [Support](/docs/support) searches by `x-tokens-request-id`. ### Output limits and the credit reservation Anthropic does not count `max_tokens` against your output-token rate limit, so many apps set it high. Tokens reserves the worst-case cost of a request before it forwards it, using your `max_tokens`. On a low balance or a key close to its cap, a large `max_tokens` can be refused. On a low balance Tokens can lower it to what you can afford (never below 16), and the answer then ends with `stop_reason: "max_tokens"`. You pay for the tokens used, not the reservation. Set `max_tokens` to what you need. ### More differences - **No CORS.** Calls from a browser fail. Call Tokens from a server. - **Body size.** Up to 10 MB, against 32 MB at Anthropic. Large base64 images or PDFs count. - **Latency.** Tokens adds one network hop. - **Privacy.** Tokens stores usage metadata, not prompt content, and the provider that serves the model sees the prompt. See [security and privacy](/docs/security-and-privacy). - **Billing.** Tokens bills in USD or BDT from a plan or a wallet, at the catalog price of each model. See [plans and wallet](/docs/plans-and-wallet). - **Claude Code and other agents.** Agents that use the Anthropic protocol have their own setup. See [Claude Code](/docs/claude-code). ## Test the switch safely 1. **Create a second key** at [/dashboard/keys](/dashboard/keys) with a **low monthly spend cap** and an **allowed-models list** with only the models you are testing. Neither can be edited later. Tokens counts the worst-case cost of each request against the cap. 2. **Read the base URL, key and model id from configuration.** Both Anthropic SDKs accept the three values from environment variables, so a deploy can switch without a code change. 3. **Run both for a while.** Replay logged requests, or mirror a share of live traffic and discard the Tokens answers. Compare: | Check | How | | ------------------------------------ | ---------------------------------------------------------------------------------------------------- | | Quality | Run your own prompts or evals. A different model gives different answers. | | Prompt caching | Read `usage.cache_read_input_tokens` on the second call with the same prefix. Zero means no cache hit on this model. | | Thinking, citations, server tools | Send one request that uses each, to the exact model. Check the response has the blocks you expect. | | Tool calls | `tool_use` inputs match your schema. | | `stop_reason` | More `max_tokens` than before means the limit is too low for this model. | | Cost per finished task | Compare Anthropic's invoice with the cost in [usage](/docs/usage-and-alerts). | | Errors | Count them by `error.code`. | 4. **Ramp up** with a feature flag or a percentage. Create the production key with the cap you want before you cut over, because a rotation or a revoke takes effect at once. ## Roll back 1. Keep your Anthropic key and billing active until Tokens has carried production traffic for a full billing cycle. 2. Put the old base URL (or unset `ANTHROPIC_BASE_URL`), key and model ids back in the configuration, and deploy or flip the flag. 3. Revoke the Tokens key you no longer use at [/dashboard/keys](/dashboard/keys). Your wallet balance stays on your account; see the [refund policy](/refund-policy). If you use the Files or Batches APIs, those parts never left Anthropic, so they need no rollback. ## Where to go next - [Messages](/docs/messages) for headers, streaming events and the translation rules. - [Anthropic SDK](/docs/anthropic-sdk) for Python and TypeScript setup. - [Errors](/docs/errors) and [Rate limits](/docs/rate-limits). - [Migrate from OpenAI](/docs/migrate-from-openai) and [Migrate from OpenRouter](/docs/migrate-from-openrouter). Sources, checked October 2026: Anthropic [API overview](https://platform.claude.com/docs/en/api/overview), [errors](https://platform.claude.com/docs/en/api/errors), [rate limits](https://platform.claude.com/docs/en/api/rate-limits), [prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching), [citations](https://platform.claude.com/docs/en/build-with-claude/citations), [List Models](https://platform.claude.com/docs/en/api/models/list) and the [anthropic-sdk-python](https://github.com/anthropics/anthropic-sdk-python) and [anthropic-sdk-typescript](https://github.com/anthropics/anthropic-sdk-typescript) sources. Tokens behaviour is from the gateway code and the pages linked above. It has not been tested against a live Anthropic account. --- # One key per customer or environment > Building a product on Tokens: what a key can be limited by, how keys are created and revoked (dashboard only, no key-management API), how many you can have, how to track spend per customer, and what your own app must do. Section: Guides. Page: https://tokens.bd/docs/one-key-per-customer If you build a product on Tokens, you will want to know which part of your spend belongs to which customer, and to stop one customer or one broken service from using up everything. Separate API keys are the tool Tokens gives you for this. This page says what a key can and cannot do, so you can decide where to use keys and where to do the work in your own app. Read [API keys](/docs/api-keys) first for the basics. This page builds on it. ## What Tokens can and cannot do here | You might expect | What exists | | ---------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------- | | Create a key with a spend cap | Yes, in the dashboard at [/dashboard/keys](/dashboard/keys) | | Restrict a key to some models | Yes, an allowed-models list set at creation | | Create, rotate or revoke a key by API | **No.** There is no key-management endpoint you can call with a `tok_live_` key | | Set an expiry date on a key | **No.** The creation form has no expiry field. `key_expired` applies only to keys an administrator provisioned with an expiry | | Edit a key's cap or models later | **No.** Both are fixed at creation | | A separate rate limit for each key | **No.** Requests per minute and concurrency are per account, shared by all keys | | Spend per key in the dashboard or an API | **No.** See [Track spend per customer](#track-spend-per-customer) | | Hundreds of keys on one account | Only if your plan allows it. The default is 3 active keys | The dashboard talks to routes under `/api/keys`. They are authenticated by your signed-in browser session (and your multi-factor check, if you use one), not by a Tokens API key, and they are not a supported public API. Do not build a provisioning service on them. The [Tokens CLI](/docs/tokens-cli) signs in through the browser and receives one new key per login; it is for setting up your own coding agents, not for issuing keys to customers. ## What a key can be limited by When you create a key in the dashboard you can set: | Setting | Effect | | ----------------------- | ----------------------------------------------------------------------------------------------------------------------------- | | Name | Up to 64 characters. Use it to record whose key it is: `acme-prod`, `staging`, `batch-worker`. | | Monthly spend cap (USD) | The most the key may spend in a calendar month. Left blank, the key takes your plan's default cap if the plan defines one. | | Allowed models | The only model ids the key may call. Empty means every model your account can use. `GET /v1/models` lists only allowed models. | A key also stops working when it is revoked or its account is suspended. Rate limits, concurrency and usage windows are not key settings; they belong to your account, see [rate limits](/docs/rate-limits). To read a key's own settings from code, call `GET /v1/tokens/usage` with that key. The `key` object in the response shows its `monthlySpendCapUsd` and `allowedModels`. The `plan`, `windows` and `wallet` parts of the same response describe your whole account, not the key. ```bash curl -s https://tokens.bd/v1/tokens/usage \ -H "Authorization: Bearer $CUSTOMER_KEY" | jq .key ``` ## How many keys you can have Each plan sets a maximum number of active keys, and the default is 3 (pay-as-you-go accounts also default to 3). Creating one more returns `key_limit_reached`. Revoked keys do not count. This decides your design: - **A handful of services or environments** (production, staging, a batch worker, an internal tool): one key each fits inside the default. - **Many customers**: a key per customer works only if your plan's limit is high enough. Check your plan on [pricing](/pricing), or ask [support](/docs/support) what limit is possible for you. Do not plan around a number you have not confirmed. - **Customers beyond your key limit**: use one key (or a few) for the whole product and do the per-customer work in your own app, as described below. :::note More keys do not mean more throughput. The per-minute limit and the concurrency limit are counted per account, so ten keys share the same limits as one. If one customer can send a burst, put your own limit in front of Tokens, for example a queue or a semaphore per customer. ::: ## Create keys for environments and services A good minimum set: | Key | Cap | Allowed models | | ---------------- | -------------------------------- | ------------------------------------ | | `prod` | The most you accept to lose to a bug in one month | The models your product uses | | `staging` | A small amount | One cheap model | | `ci` or `batch` | The cost of the job, with margin | The one model the job needs | Set the cap and the model list when you create the key, because you cannot change them afterwards. If a cap turns out to be wrong, create a new key with the right values, switch your app to it, then revoke the old one. Keep the old key's secret out of any place you have already moved on from. ## What happens at the cap When a request would take the key over its monthly cap, the gateway refuses it before it reaches a model: ```json { "error": { "message": "Monthly spend cap of $20.00 for this API key has been reached. Update the key cap or use an alternate key: https://tokens.bd/dashboard/billing", "type": "permission_denied_error", "code": "monthly_spend_cap_exceeded", "param": null, "request_id": "..." } } ``` The status is `403`, not `429`, and there is no `Retry-After` header: waiting a few seconds does not help, and the key stays blocked until the next calendar month. Details to design around: - The check adds up what the key has already spent this month and the worst-case cost of the new request, which depends on its `max_tokens` (8,192 output tokens if you do not set it). A request with a large `max_tokens` can be refused slightly before the cap is actually reached, while a smaller one still passes. - Spend is counted when a request finishes. Requests in flight at the moment you reach the cap are not counted yet, so a key can end a little over its cap. - Only that key stops. Your other keys, your plan and your wallet are not affected, unless the account itself is out of credit (see [plans, credits and wallet](/docs/plans-and-wallet)). - Rotating a key does not reset its spend. The key keeps the same identity, so the month's usage still counts against the cap. In your app, treat `monthly_spend_cap_exceeded` as "this customer or environment has used its allowance", not as an outage. Show your own message, and do not retry. See [errors](/docs/errors) for the full list of codes. ## Track spend per customer Tokens does not show spend per key. The Usage page, the CSV export and `GET /v1/tokens/usage` report your account as a whole, and the CSV has no key column. So the record of which customer used what has to be kept by your app. For every call you make on a customer's behalf, store: 1. your customer id, 2. the `x-tokens-request-id` response header, 3. the model you called, and 4. the `usage` object from the response (`prompt_tokens`, `completion_tokens` for Chat Completions). You can then multiply the tokens by the model's prices from [the catalog](/models) to get a cost per customer. Your totals will not match the bill to the last cent if a plan discount or allowance applies, so use the **Request ID** column in the exported CSV ([usage and alerts](/docs/usage-and-alerts)) to look up what a given request actually cost. For streaming responses, add `"stream_options": {"include_usage": true}` to the request so the last chunk carries the usage. See [streaming](/docs/streaming). ```python title="record_usage.py" import os from openai import OpenAI client = OpenAI( base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"], ) def ask_for_customer(customer_id: str, prompt: str) -> str: raw = client.chat.completions.with_raw_response.create( model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": prompt}], max_tokens=500, ) completion = raw.parse() save_usage_row( # your own database write customer_id=customer_id, request_id=raw.headers.get("x-tokens-request-id"), model=completion.model, prompt_tokens=completion.usage.prompt_tokens, completion_tokens=completion.usage.completion_tokens, ) return completion.choices[0].message.content ``` If you use a key per customer, the key's cap protects you; the table above is still how you bill them. ## Team accounts Tokens has team accounts (organizations) with roles, and they are switched on per deployment, so you may not see them yet in your dashboard. Where they are on, keys belong to the organization rather than to one person. Based on the code, the roles are owner, admin, developer, billing and viewer; the billing and viewer roles cannot create keys, a developer sees the keys they created, and owners and admins can rotate or revoke any key in the organization. An organization can also give each member a monthly spend cap, which returns `402 member_cap_reached` when it is reached. Read [teams and roles](/docs/teams-and-roles) for what your account has. ## Rotate and revoke - **Rotate** replaces the secret and keeps the name, cap, allowed models and usage history. The old secret stops working immediately; there is no grace period. Requests that use it get `401 invalid_api_key`. - **Revoke** disables the key permanently. Requests with it get `403 key_inactive`. To rotate without downtime, create a second key with the same settings first, deploy it to your app, then revoke the old one. Because key settings are fixed, "same settings" means you enter them again. Rotate or revoke at once if a key leaks, then look at [usage](/dashboard/usage) for requests you do not recognise. ## What your own app must do - **Map your user to a key.** Keep a table from customer or environment to the key. The secret is shown once, so store it in a secrets manager or an encrypted column, never in plain text and never in logs. - **Keep keys on your server.** Tokens does not answer browser calls (no CORS headers), and anything in a browser or mobile app can be extracted by users. Your clients call your backend, and your backend calls Tokens. See [browser and mobile apps](/docs/browser-and-mobile). - **Limit abuse yourself.** Per-customer rate limits, request-size limits and a maximum `max_tokens` belong in your code, because Tokens' rate limits are per account. - **Handle the key errors.** Map `monthly_spend_cap_exceeded`, `model_not_allowed_on_key`, `key_inactive` and `invalid_api_key` to a clear state in your product, and alert yourself on `key_inactive` and `invalid_api_key`, which mean your own configuration is wrong. - **Watch the account.** Every key spends from the same plan and wallet. Turn on the low-balance and usage alerts in the dashboard ([usage and alerts](/docs/usage-and-alerts)), and check [production checklist](/docs/production-checklist) before launch. - **Read the terms.** What you may build and resell on top of your account is set by the [terms of service](/terms), not by this page. ## Checklist 1. Count the keys you need and compare with your plan's active-key limit before you design around them. 2. Create each key with a cap and an allowed-models list, because you cannot add them later. 3. Store every secret on the server, and record your customer id with each request. 4. Test the cap path: create a key with a tiny cap and confirm your app handles `403 monthly_spend_cap_exceeded`. 5. Write down how you rotate a key, and try it once before you need it. --- # Cutting your spend > Practical ways to spend less on Tokens, from the choices that change the bill most to the ones that prevent surprises: model choice, max_tokens, context size, caching, plan allowances, background models, cancelling streams, and spend caps. Section: Guides. Page: https://tokens.bd/docs/cutting-costs A request costs the tokens you send plus the tokens the model writes, each at the model's price per million tokens. Everything on this page changes one of three things: which price you pay, how many tokens go in, or how many come out. The sections run roughly from the choices that usually change a bill the most to the ones that mainly prevent surprises. Your own usage will differ, so measure before and after each change instead of trusting any ordering, including this one. ## Measure first You cannot cut what you cannot see. Three places show what a request cost: - **Every response** carries a `usage` object with the input and output token counts. Log it. - **The Usage page** in the dashboard lists recent requests with model, tokens and cost, and charts daily spend with a breakdown sorted by spend. Start with whichever model is on top. - **The CSV export** has one row per request with input, output and cache-read tokens and the cost in USD. See [usage, limits and alerts](/docs/usage-and-alerts). Look for two patterns: one model that accounts for most of the spend, and requests whose input is far larger than the output. The first points to model choice, the second to context size. ## 1. Pick the right model for the job The biggest difference is usually the price per token of the model itself. Model prices on Tokens span a wide range, and many everyday jobs (commit messages, classification, test scaffolding, small edits) do not need the most expensive model. - Use a strong model for hard, long-running work, and a cheap one for bulk or simple work. - Run the same real task on two models and compare **cost per finished task**, not price per token. A pricier model that finishes in fewer turns can cost less. - Models that always think before answering add output tokens to every call. [Choosing a model](/docs/choosing-a-model) lists representative models with their list prices, and [/models](/models) has your actual price for each one. Models that your plan includes can also be cheaper for you than pay-as-you-go, see section 5. ## 2. Set `max_tokens` `max_tokens` is the most the model may write. Output tokens cost more than input tokens on most models, and a runaway answer is the easiest way to overspend. - Set it to what a good answer needs, not to the largest value the model allows. A classification label needs tens of tokens, a code review a few thousand. - If you leave it out, Tokens assumes 8,192 output tokens when it checks your balance, key cap and plan limits. Setting a smaller value lets requests pass near a cap that a larger one would not. - On the Responses API the field is `max_output_tokens`; some newer OpenAI models use `max_completion_tokens` on Chat Completions. The gateway reads all three. - When `finish_reason` is `length`, the answer was cut off by your limit. If that happens often, the limit is too low for the task; do not just double it everywhere. - If your wallet cannot cover the `max_tokens` you asked for, Tokens may lower it to what the balance covers, never below 16. The answer then comes back shorter. See [plans, credits and wallet](/docs/plans-and-wallet). ```python title="max_tokens.py" import os from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) resp = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": "Label this commit message as fix, feat or chore: 'handle empty cart'"}], max_tokens=10, ) print(resp.choices[0].message.content, resp.usage.completion_tokens) ``` ## 3. Trim the context Chat APIs are stateless. Each request carries the whole conversation, so a long session pays for the same early messages again on every turn. Coding agents add file contents and tool output on top, which is why input tokens dominate long agent sessions. What helps: - **Send less per turn.** Drop tool results the model no longer needs, send the relevant file or function instead of the whole tree, and keep the system prompt short. - **Start new conversations** for new tasks instead of continuing an old one. - **Summarize old turns.** Replace the first part of a long conversation with a short summary, and keep the last few messages as they are. - **Do not paste what a tool can fetch.** Agents that search and read selectively usually send less than one that is given everything up front. - **Remember the price of a large window.** A 1M-token context is available on many models, but every token you fill is billed on every request. Some providers charge a higher rate above a size threshold, noted in [choosing a model](/docs/choosing-a-model). To see how large your inputs really are, read `usage.prompt_tokens` from a response, or use [token counting](/docs/token-counting) before you send. ```python title="trim_history.py" def trim_history(messages, keep_last=6): """Keep the system prompt and the most recent messages.""" system = [m for m in messages if m["role"] == "system"] rest = [m for m in messages if m["role"] != "system"] return system + rest[-keep_last:] ``` This is a blunt tool: cutting turns can remove facts the model needs. Summarize instead of dropping when the early turns matter. ## 4. Use prompt caching When the start of your prompt is the same from request to request (the system prompt, tool definitions, a big file), the provider can read it from a cache at a lower price. Cached input is shown in the usage record as cache-read tokens and priced at the model's cache-read rate where one exists. How to structure prompts for it is in [prompt caching](/docs/prompt-caching); the saving depends on the model and your traffic, so check the cache-read tokens in your usage rows. ## 5. Use your plan's allowances and free models How a plan treats each model changes what a request costs you. From [plans, credits and wallet](/docs/plans-and-wallet): - A plan can give a model **its own allowance**. It is a cap inside your plan's usage, not extra on top: using the model draws from the plan's credits and from the model's allowance at the same time. - A **free model** costs nothing, does not use your credits or any allowance, and keeps working after your plan's credits run out. If a free model is good enough for a job (see the plan's page on [pricing](/pricing)), route that job to it. - A **deal** lowers a model's price for plan usage, so the same credits go further. - When a model's allowance is used up you get `429 model_limit_reached` until the reset. On a plan with pay-as-you-go fallback, requests to that model are instead paid from your wallet at the **full catalog price**, with no deal applied. That is a quiet way to spend more than you planned, so watch it. ## 6. Give background calls a cheap model Several tools make small calls in the background (titles, summaries, quick edits) in addition to the main conversation. If those use the same expensive model as your main work, they add up. Where a tool lets you set a second model, set a cheap one: - **Claude Code**: `ANTHROPIC_DEFAULT_HAIKU_MODEL` backs the `haiku` alias and also runs background tasks. If you leave it unset, background tasks use the main model. `CLAUDE_CODE_SUBAGENT_MODEL` sets the model for subagents. See [Claude Code](/docs/claude-code). - **OpenCode**: `small_model` takes a cheaper model for lightweight tasks. See [OpenCode](/docs/opencode). - **Crush**: you choose a large and a small model. See [Crush](/docs/crush). Check the Usage page after a session. If you see a model you did not choose to use, a background setting is the usual cause. ## 7. Cancel streams you no longer need Closing the connection ends the request on Tokens' side, and you are billed for what had been produced up to that point. This is how Tokens meters a request (checked in the gateway code, October 2026): - If you close the connection, the gateway passes the cancellation on to the provider and settles the request with the tokens counted so far. - When the provider has already reported usage, that report is used. When it has not yet (many providers send usage only at the end of a stream), Tokens estimates from the length of the text already sent. Such a record is marked as estimated. - Input tokens are not refunded. The prompt was already read. - An abandoned stream that is never closed keeps one of your concurrency slots until it ends, so close streams you do not read to the end. See [rate limits](/docs/rate-limits). In practice: if a user stops a response, abort the HTTP request instead of ignoring the remaining chunks. ```python title="cancel_stream.py" import os from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) stream = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": "Explain how HTTP caching works."}], max_tokens=800, stream=True, ) text = "" for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: text += chunk.choices[0].delta.content if len(text) > 400: # enough, stop paying for more stream.close() break print(text) ``` ## 8. Understand windows, plan and wallet Plans and the wallet limit spend in different ways, and knowing which one is running tells you what to watch. - **A plan** gives a credit allowance, and some plans spread it over a 5-hour session, a week, or a month. When a window is used up, requests return `429 window_exhausted` until it resets, whatever your wallet holds. That is a built-in brake on a heavy day. - **The wallet** is prepaid pay-as-you-go. No usage window limits it, only its balance. A plan with fallback turns to the wallet when credits run out, at full price. - A key's **monthly spend cap** limits one key only, see section 9. Read where you stand with `GET /v1/tokens/usage`. It is free, not metered, and does not count towards your per-minute limit. ```bash curl -s https://tokens.bd/v1/tokens/usage -H "Authorization: Bearer $TOKENS_API_KEY" \ | jq -r '.windows[] | "\(.label): \(.percentUsed)% used, resets \(.resetsAt)"' ``` A window that has not been used yet in its current period is left out of the response, so an empty list after a reset is normal. Use the response to stop a batch job before it hits a window: ```python title="gate.py" import os import sys import requests resp = requests.get( "https://tokens.bd/v1/tokens/usage", headers={"Authorization": f"Bearer {os.environ['TOKENS_API_KEY']}"}, timeout=10, ) resp.raise_for_status() usage = resp.json() for window in usage["windows"]: if window["percentUsed"] >= 80: sys.exit(f"{window['label']} is {window['percentUsed']}% used, resets {window['resetsAt']}") balance = (usage.get("wallet") or {}).get("balanceUsd") if balance is not None and balance < 5: sys.exit(f"Wallet balance is ${balance:.2f}") print("OK to start the job") ``` The 80 and 5 are example thresholds; pick your own. ## 9. Caps and alerts Limits turn a mistake into an error instead of a bill. - **Monthly spend cap per key.** Set it when you create the key; it cannot be edited later. A key used by CI, a script or an unattended agent should always have one. A request that would take the key over the cap is refused with `403 monthly_spend_cap_exceeded`. See [API keys](/docs/api-keys) and, for several keys, [one key per customer](/docs/one-key-per-customer). - **Allowed models per key.** A key restricted to one cheap model cannot call an expensive one by mistake. - **Email alerts.** Under Notifications you can get warnings at 50%, 75% and 90% of a plan limit, and a low-balance alert when the wallet drops below $5. Check that they are on. ## Checklist 1. Find the model that takes most of your spend on the Usage page. 2. Ask whether a cheaper or free model can do that job, and test it on real tasks. 3. Set `max_tokens` on every request. 4. Check input size with `usage.prompt_tokens`, then trim, summarize or cache what repeats. 5. Set a cheap model for background and small tasks in your tools. 6. Cancel streams you no longer need. 7. Put a monthly cap on every key that runs unattended, and keep the alerts on. --- # Working with Bengali text > How Bengali and other non-Latin text behaves through the API: why it takes more tokens, how to measure it from the usage field, what that means for max_tokens, context and cost, prompting tips, UTF-8 in streams, and a script to compare models on your own text. Section: Guides. Page: https://tokens.bd/docs/bengali-text Bengali works through Tokens like any other text: you send it as UTF-8 in the `messages` and read the answer back. Three things differ from English. The same content usually uses more tokens, so it costs more and fills the context window faster. Some bugs only show up with multi-byte characters, especially in streams. And you should check output quality on your own text before you pick a model. This page covers each, and ends with a script that compares models on a Bengali prompt. The same points apply, to different degrees, to Hindi, Arabic, Thai, Chinese and other non-Latin scripts. ## Why Bengali takes more tokens Models do not read characters. A tokenizer first splits the text into pieces called tokens, and you pay per token. A tokenizer is built from a large body of text, and a script that is rare in that text tends to be split into more, smaller pieces. How much depends on the model's tokenizer, and the maker can change it between versions. Anthropic, for example, documents that newer Claude models use a different tokenizer from older ones and that the same text produces a different number of tokens, and tells you to recount prompts against the model you plan to use ([Anthropic: token counting](https://platform.claude.com/docs/en/build-with-claude/token-counting), checked October 2026). Bengali text is also built from many small Unicode parts. You can see this in Python: ```python text = "বাংলা" print(len(text)) # 5 Unicode code points print(len(text.encode("utf-8"))) # 15 bytes: 3 bytes for each Bengali character print(len("ক্ষ")) # 3: ক + hasanta (virama) + ষ make one conjunct ``` Every Bengali letter takes three bytes in UTF-8, and a conjunct is several code points. A tokenizer that works on bytes starts from those pieces. How far it merges them back into larger tokens depends on whether Bengali was well represented in its training text, and that differs from model to model. Do not rely on a ratio you read somewhere. Measure it on your own text and your own model, as shown next. ## Measure it with a real response Every response reports its token counts in `usage`. Send an English text and a Bengali text with the same meaning, and compare `prompt_tokens`. Set `max_tokens` low so the call costs almost nothing: you are billed for the input you send, and for at most a few output tokens. ```python title="measure_tokens.py" import os from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) ENGLISH = ( "Our team launched a new payment page last week. In the first two days, " "some users got an error when paying by card. The cause was an old library, " "and we have updated it." ) BENGALI = ( "আমাদের দল গত সপ্তাহে নতুন পেমেন্ট পেজ চালু করেছে। প্রথম দুই দিনে কয়েকজন ব্যবহারকারী " "কার্ড দিয়ে টাকা দিতে গিয়ে ত্রুটি পেয়েছেন। সমস্যাটি একটি পুরনো লাইব্রেরির কারণে হয়েছিল, " "এবং আমরা সেটি আপডেট করেছি।" ) def input_tokens(model: str, text: str) -> int: resp = client.chat.completions.create( model=model, messages=[{"role": "user", "content": text}], max_tokens=16, # the input is what we measure; keep the output tiny ) return resp.usage.prompt_tokens for model in ["deepseek/deepseek-v4.1-flash"]: # add more ids from /models en = input_tokens(model, ENGLISH) bn = input_tokens(model, BENGALI) print(f"{model}: English {en} tokens, Bengali {bn} tokens, ratio {bn / en:.2f}") ``` Notes on reading the result: - `prompt_tokens` includes a few tokens the model's chat format adds around your message. For very short texts that overhead distorts the ratio, so measure paragraphs of realistic length. - Run it for **each model you might use**. Different models can give very different counts for the same text. - Use your real prompts. A system prompt, tool definitions and pasted code are mostly English or code, and only part of your input is Bengali. - Some models that think before answering may reject a very small `max_tokens`. If you get an error, raise it a little. You can also count tokens before you send; see [token counting](/docs/token-counting). ## What it means for `max_tokens`, context and cost - **`max_tokens`** counts output tokens. A Bengali answer of the same length needs more of them, so a limit that is comfortable in English can cut a Bengali answer short. Check `finish_reason`: `length` means your limit stopped it. Raise `max_tokens` for Bengali answers, then confirm the result by measuring. - **Context window.** The window is measured in tokens, not characters. A conversation in Bengali reaches the limit sooner than the same conversation in English, so trim or summarize history earlier. Limits per model are on [/models](/models). - **Cost.** You pay per token at the model's price, so the cost of a Bengali request is its token count times the price. Compare models by the cost of your Bengali task, not by price per token alone: a model with a lower price and a worse tokenizer for Bengali can cost more per request. See [cutting your spend](/docs/cutting-costs). - **Speed.** More tokens to write means longer to finish, even at the same tokens per second. ## Prompting for Bengali These are habits that make output predictable. Test each on your own task. - **Say the output language explicitly.** A model answers in the language it considers likely, which may follow your instructions, the user message, or the system prompt. Put it in the system prompt, for example `Always answer in Bengali (Bangla script).` If you want English, say so just as plainly. - **Name the script.** "Bengali" can come back in Bangla script or in Latin letters (romanized Bengali). Ask for Bangla script if that is what you need. - **Keep code and identifiers in English.** Variable names, function names, file paths, JSON keys and error messages are best left in English. Write instructions like `Explain in Bengali, but keep code, identifiers and JSON keys unchanged.` - **Machine-readable output.** If your code parses the answer, ask for a fixed format and say which digits to use. Bengali digits (০১২৩) and ASCII digits (0123) are different characters, and `int("১২")` works in Python while a regular expression for `[0-9]` will not match it. Say `Use ASCII digits` if you need them. - **Give an example in the format you want.** One short Bengali example in the prompt shows tone and script better than a description does. - **Mind mixed-script text.** Bengali with English technical terms is normal, but similar-looking characters can cause bugs. The Bengali danda `।` (U+0964) is not a pipe `|` or a Latin `I`, and the Bengali digit zero `০` (U+09E6) is not the Latin `0`. Search, comparison and parsing treat them as different characters. - **Normalize Unicode before you compare or store.** The same Bengali text can be encoded in more than one way. For example `য়` can be one code point (U+09DF) or two (U+09AF followed by U+09BC), and the vowel sign in `কো` can be one code point (U+09CB) or two (U+09C7 followed by U+09BE). Normalize with NFC before you compare, search or deduplicate: ```python import unicodedata def clean(text: str) -> str: return unicodedata.normalize("NFC", text) assert clean("য়") == "য়" # য় composed form becomes two code points assert clean("কো") == "কো" # কো, two vowel parts become one sign ``` Normalizing the text you send also keeps your prompts consistent, which helps if you test or cache them. It does not change how the model writes its answer. ## Streaming and UTF-8 in clients With `stream: true` the answer arrives in chunks. If you use an official SDK, it decodes the stream for you and you can skip this section. If you read the response bytes yourself, remember that one Bengali character is three bytes, and a network chunk can end in the middle of one. Decoding each chunk on its own then gives a replacement character (`�`) or an error. Use an incremental decoder that keeps the unfinished bytes until the rest arrives. :::code-tabs ```python title="Python" import codecs decoder = codecs.getincrementaldecoder("utf-8")() def feed(chunk: bytes) -> str: return decoder.decode(chunk) # returns only complete characters data = "বা".encode("utf-8") # 6 bytes print(repr(data[:2].decode("utf-8", errors="replace"))) # '�' when split naively print(repr(feed(data[:2]) + feed(data[2:]))) # 'বা' with the incremental decoder ``` ```javascript title="Node.js / browser" const decoder = new TextDecoder("utf-8"); function feed(chunk) { // { stream: true } keeps an unfinished character until the next chunk return decoder.decode(chunk, { stream: true }); } ``` ::: Related points: - Decode first, then split into lines. Splitting raw bytes on newlines is safe for UTF-8 (the newline byte never appears inside a multi-byte character), but splitting a decoded string at a byte offset is not. - When you cut text to a length, cut on characters, not on bytes. A byte-based cut can end in the middle of a character. - Send requests with `Content-Type: application/json`, and make sure your HTTP library encodes the body as UTF-8. Do not write the JSON by escaping each character by hand. - Read the server-sent events format in [streaming](/docs/streaming). To get token counts at the end of a stream, add `"stream_options": {"include_usage": true}` to the request. - If you display text in a terminal or on Windows, the display font and console code page can show boxes or `?` for text that is correct in memory. Check by writing the string to a UTF-8 file. ## Check which models handle Bengali well There is no reliable shortcut. A maker's claim that a model supports a language does not tell you how it does on your task, and a score on a public test does not tell you how it does on your text. Do not choose from a ranking. Test: 1. **Collect 20 to 50 real prompts.** Use the kind of text your product handles: support messages, summaries, code explanations, form data. 2. **Run the same prompts on several models** with the script below, and keep the answers. 3. **Read the answers yourself, or have a Bengali speaker read them.** Look at correctness first, then fluency, whether the script and language are what you asked for, and whether names, numbers and technical terms survived. 4. **Compare cost and speed** from `usage` and the run time, alongside quality. 5. **Re-run when you change models or prompts.** Models change often, so keep your prompt set. Candidates come from [choosing a model](/docs/choosing-a-model) and [/models](/models). Use the exact ids from there, or from `GET /v1/models`, which lists what your key can call. ## A script to compare models on the same prompt This script sends one Bengali prompt to each model and prints the token usage, the finish reason, the time taken and the answer. Add the model ids you want to compare to `MODELS`, or set them in the `TOKENS_MODELS` environment variable, separated by commas. ```python title="compare_models.py" import os import time import unicodedata from openai import OpenAI client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]) MODELS = os.environ.get("TOKENS_MODELS", "deepseek/deepseek-v4.1-flash").split(",") SYSTEM = ( "তুমি একজন সহায়ক সহকারী। সবসময় বাংলা অক্ষরে বাংলায় উত্তর দেবে। " "কোড, ভেরিয়েবলের নাম এবং JSON কী ইংরেজিতেই রাখবে।" ) PROMPT = unicodedata.normalize( "NFC", "নিচের বার্তাটি দুই বাক্যে সংক্ষেপে লেখো এবং গ্রাহকের জন্য একটি বিনীত উত্তরের খসড়া দাও।\n\n" "বার্তা: আমি গতকাল কার্ড দিয়ে টাকা দিয়েছি, কিন্তু আমার অ্যাকাউন্টে ক্রেডিট আসেনি। " "পেমেন্টের রসিদ আমার কাছে আছে। দয়া করে দ্রুত দেখুন।", ) print(f"Prompt: {len(PROMPT)} characters\n") for model in (m.strip() for m in MODELS if m.strip()): started = time.time() try: resp = client.chat.completions.create( model=model, messages=[ {"role": "system", "content": SYSTEM}, {"role": "user", "content": PROMPT}, ], max_tokens=1000, ) except Exception as err: # a model your key cannot call, a rate limit, and so on print(f"=== {model}\nerror: {err}\n") continue elapsed = time.time() - started choice = resp.choices[0] usage = resp.usage print(f"=== {model}") print( f"input {usage.prompt_tokens} tokens, output {usage.completion_tokens} tokens, " f"finish_reason={choice.finish_reason}, {elapsed:.1f}s" ) print(choice.message.content) print() ``` Run it with your key set: :::code-tabs ```bash title="macOS / Linux" export TOKENS_API_KEY="tok_live_your_key" export TOKENS_MODELS="deepseek/deepseek-v4.1-flash" python compare_models.py ``` ```powershell title="Windows PowerShell" $env:TOKENS_API_KEY = "tok_live_your_key" $env:TOKENS_MODELS = "deepseek/deepseek-v4.1-flash" $env:PYTHONUTF8 = "1" python compare_models.py ``` ::: On Windows, `PYTHONUTF8=1` makes Python print Bengali correctly in the console; without it you may see an encoding error even though the request worked. When you read the output, look at `output tokens` next to `finish_reason`. A `length` means the answer was cut off at `max_tokens`, and the models you are comparing may need different limits to finish. The calls are billed like any others; each run of the script above is small, and a cap on the key you test with keeps a mistake cheap. See [API keys](/docs/api-keys). ## Troubleshooting **The answer is in English although the prompt was Bengali.** Add an explicit instruction to the system prompt, in English and in Bengali, such as `Always answer in Bengali (Bangla script).` Check that no later instruction or example in the conversation is in English. **The answer is Bengali written in Latin letters.** Ask for Bangla script by name, and show a short example. **Boxes, `?` or `�` appear in the output.** If it is only on screen, it is a font or console problem. If it is in the saved file or the data, check your stream decoding and any step that cuts bytes. **The answer stops mid-sentence.** Check `finish_reason`. If it is `length`, raise `max_tokens`. If your wallet is low, Tokens may have lowered it; see [plans, credits and wallet](/docs/plans-and-wallet). **The bill is higher than the English equivalent.** That is expected from the token counts. Measure it with the first script, choose the model by cost on your own text, and see [cutting your spend](/docs/cutting-costs). **Comparison or search misses text that looks identical.** Normalize both sides to NFC, and check for look-alike characters such as `।` and `|`. --- # Troubleshooting > Fix common Tokens API errors by symptom: 401 invalid key, 403 model not allowed, 402 insufficient credits, 429 limits, 404 model not found, 5xx, wrong base URL, stuck streams and Windows env vars. Section: Troubleshooting. Page: https://tokens.bd/docs/troubleshooting Find your symptom below and apply the fix. Most Tokens API errors come with a specific `error.code` in the response body, and the code tells you more than the HTTP status, so start by reading it: ```bash curl -i https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" ``` `-i` also prints `x-tokens-request-id`. Keep it, because support will ask for it. If many things are failing at once, check [/status](/status) before you debug your own setup. ## 401: missing_api_key or invalid_api_key **Symptom:** every request fails immediately, including `GET /v1/models`. **Fix:** - `missing_api_key`: no key reached the gateway. Send it as `Authorization: Bearer ` or `x-api-key: `. In a shell, run `echo $TOKENS_API_KEY` (or `echo $env:TOKENS_API_KEY` in PowerShell) to check the variable isn't empty in _this_ session. - `invalid_api_key`: the key is wrong, revoked, or was rotated. Keys start with `tok_live_` followed by 48 hex characters. Look for a truncated paste, stray quotes or a trailing newline. After a rotation the old secret stops working immediately, so update every client. Lost a key? Keys are shown once, so create a new one in [/dashboard/keys](/dashboard/keys). ## 403: model_not_allowed_on_key or tier_permission_denied **Symptom:** some models work and others return 403. **Fix:** - `model_not_allowed_on_key`: the key was created with an allowed-models list that doesn't include this model. Allow-lists can't be edited after creation, so create a new key with the models you need. - `tier_permission_denied`: your plan doesn't include this model and you have no wallet balance to pay for it. Pick a model your plan includes, add funds to the wallet, or change plans in [/dashboard/billing](/dashboard/billing). See [Plans and wallet](/docs/plans-and-wallet). - `monthly_spend_cap_exceeded`: the key reached its monthly spend cap. Wait for the next month or use a key with a higher cap. - `key_inactive`, `key_expired`, `account_suspended`: the key or account has been disabled. Contact [support](/docs/support). `GET /v1/models` lists exactly the models this key can use right now, which settles most 403 questions quickly. ## 402: insufficient_credits **Symptom:** requests that used to work start failing with 402. Related codes are `no_funding` (no active plan and no funded wallet), `outstanding_debt` (settle the balance shown in the message) and `member_cap_reached`. **Fix:** your plan credits and wallet can't cover the request. Top up or renew in [/dashboard/billing](/dashboard/billing). The minimum top-up is $5 in dollars or ৳500 in taka. Subscriptions are one-time payments per period and don't renew on their own, so a plan that has quietly expired looks exactly like this. Replies cut short while your balance is low are the same problem in milder form: the gateway lowers `max_tokens` to what the balance covers (floor 16). Top up, and set a low-balance alert under Notifications. ## 429: window_exhausted vs rate_limited vs concurrency_limit Three different 429s with three different fixes. Read `error.code` and the `Retry-After` header. Tokens doesn't send `X-RateLimit-*` headers. | Code | Meaning | What to do | | ------------------- | ------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- | | `window_exhausted` | Your plan's usage window (rolling 5-hour session, weekly or monthly) is used up | Wait for reset. `Retry-After` is the seconds until then, often hours. Check with `GET /v1/tokens/usage`. Otherwise, upgrade the plan. | | `rate_limited` | Too many requests per minute on your account (default 60 RPM) | Back off for the `Retry-After` time. Spread out batch jobs. | | `concurrency_limit` | Too many requests in flight at once for your account (`Retry-After: 2`) | Lower parallelism. Agents with parallel sub-agents, or several agents on one account, hit this first. | Limits apply per account, not per key, so creating more keys won't raise them. Details are in [Rate limits](/docs/rate-limits). A 429 with code `rate_limit_exceeded` comes from the upstream provider. The gateway already tried to fail over, so retry after a short pause or try another model. ## 404: model_not_found **Symptom:** `model_not_found` (or `400 model_not_available`) for a model you're sure exists. **Fix:** model IDs are exact and include the provider prefix, for example `deepseek/deepseek-v4.1-flash`, not `deepseek-v4.1-flash`. List the IDs your key can actually use: ```bash curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY" | jq -r '.data[].id' ``` `model_not_available` means the ID is valid but the model can't be served right now. Pick another one from [/models](/models). A `404 unsupported_endpoint` is different. It means the path isn't one Tokens serves (images, audio, files, batches, assistants, fine-tuning and moderations aren't supported), or the base URL is wrong. See the next section. ## Agent says "model not found" or 404: wrong base URL This is the most common setup mistake. Tools fall into two groups: | Tool type | Base URL | Why | | -------------------------------------------------------------------------- | ---------------------- | ------------------------------- | | OpenAI-compatible (Cursor, Cline, Aider, OpenCode, OpenAI SDKs, LangChain) | `https://tokens.bd/v1` | They append `/chat/completions` | | Anthropic-compatible (Claude Code, Anthropic SDKs) | `https://tokens.bd` | They append `/v1/messages` | Using `/v1` with Claude Code produces requests to `/v1/v1/messages`. Leaving `/v1` off in an OpenAI tool sends requests to `/chat/completions` on the site root. Both end in a 404 that many agents report as "model not found". Check the config file for the agent you use ([Claude Code](/docs/claude-code)), or let the [Tokens CLI](/docs/tokens-cli) write it for you. ## 502, 503, 504: upstream errors Codes: `502 upstream_unreachable`, `503 no_upstream_available`, `504 upstream_timeout`. The gateway fails over to another upstream source on server errors, 429, timeouts, connection errors and provider-side key or model problems before replying. If you still get a 5xx, every source it tried failed. Retry with backoff (two or three attempts), try a different model, and check [/status](/status) for an incident affecting upstream providers. If the error persists for a single model, open a ticket with the `x-tokens-request-id`. Long requests are fine. The gateway waits up to 600 seconds for response headers. If your own client times out first, raise its timeout or switch to streaming. ## Streaming hangs or arrives all at once **Symptom:** a streamed response shows nothing for a long time and then appears in one block, or never finishes. **Fix:** something between your client and Tokens is buffering the server-sent events. - Test with `curl -N` and `"stream": true` first ([cURL](/docs/curl)). If that streams, the problem is in your stack. - Behind your own nginx: set `proxy_buffering off;` for the route, or send `X-Accel-Buffering: no` from your app. - Don't gzip `text/event-stream` responses in your own middleware. - Corporate proxies and some antivirus HTTPS scanners buffer whole responses. Try another network to confirm. - In your own route handlers, forward chunks as they arrive instead of collecting the full body. See [Node.js](/docs/nodejs) and [Streaming](/docs/streaming). ## Tool calls fail or are ignored **Symptom:** a 400 error mentioning tools, or the model answers in plain text and never calls your function. **Fix:** not every model supports tool calling. Check the model's page in [/models](/models) and switch to one that lists tool support. Then check your schema: `parameters` must be a valid JSON Schema object, and each tool result has to be sent back with the matching `tool_call_id`. It's also why an agent can chat but never edit files. See [Tool calling](/docs/tool-calling). ## Windows: environment variable isn't picked up - `$env:TOKENS_API_KEY = "tok_live_your_key"` sets it for the **current PowerShell window only**. - `setx TOKENS_API_KEY "tok_live_your_key"` saves it for **new** windows only. The window you ran it in doesn't see it, so open a new terminal (and restart VS Code or your agent). - In `cmd.exe`, `set TOKENS_API_KEY=tok_live_your_key` has no quotes around the value. Any quotes become part of the key. - In Windows PowerShell 5.1, `curl` is an alias for `Invoke-WebRequest`. Use `curl.exe` for the examples in these docs. ## Still stuck Open a ticket from [/dashboard/support](/dashboard/support) with the `x-tokens-request-id`, the time, the model ID and the `error.code`. Don't include your API key. [Support](/docs/support) explains what to send, and [Errors](/docs/errors) has the full code reference. --- # FAQ > Short answers to common questions about Tokens: OpenAI and Anthropic compatibility, models, privacy, paying in BDT, balances, keys, receipts, account deletion and uptime. Section: Troubleshooting. Page: https://tokens.bd/docs/faq Short answers to the questions people ask most about Tokens, with links to the full docs. If your question is about a specific error code, [Troubleshooting](/docs/troubleshooting) is faster. ## API and compatibility ### Is Tokens OpenAI-compatible? Yes. Set the base URL to `https://tokens.bd/v1` and use your Tokens key as the API key. Chat Completions, the legacy Completions endpoint, the Responses API, Models and Embeddings (for embedding models only) are supported, streaming included. Image generation, audio, files, batches, assistants, fine-tuning and moderations are not. Sending an image to a model in a chat request works: see [Image input](/docs/vision). See [Overview](/docs/overview) and [Chat Completions](/docs/chat-completions). ### Is Tokens Anthropic-compatible? Yes. The Messages API is at `https://tokens.bd/v1/messages`. Tools and SDKs built for Anthropic take the base URL `https://tokens.bd` (no `/v1`) because they add the path themselves. You can call any model in the catalog this way, not only Claude models, because the gateway translates the request. See [Messages](/docs/messages) and [Claude Code](/docs/claude-code). ### Can I call Tokens from a browser app? No. Tokens' responses carry no CORS headers, so browser `fetch` calls fail, and a key in front-end code is visible to anyone who opens dev tools. Call Tokens from your server or a CLI, and have the browser call your server. [Node.js](/docs/nodejs) has a Next.js route handler example. ## Models ### Which models can I use? The catalog is curated by our team and spans several providers, including Claude, GPT, Gemini, Grok, DeepSeek, Kimi, GLM, Qwen and MiniMax model families, and it changes as models are added or retired. [/models](/models) has the full list with prices and capabilities. Your key can see the models available to it right now with `GET https://tokens.bd/v1/models`. Model IDs look like `deepseek/deepseek-v4.1-flash`. ### Do all models support tool calling? No. Tool calling, vision and reasoning depend on the model. Check the model page in [/models](/models) before you rely on a feature. See [Tool calling](/docs/tool-calling). ## Privacy ### Do you store my prompts? No. Tokens stores usage metadata only: model, token counts, cost, latency, timestamps, the key used and the request ID. Prompt and response content isn't stored, and logs redact prompts. We don't train on your data or sell it. Your prompts are sent to the upstream model provider that serves the request, and that provider's own data policy applies. Details are on [/security](/security). ## Billing and payments ### Is there a free tier? Current plans and what each includes are listed on [/pricing](/pricing). To create an API key you need a verified email and either an active plan or a positive wallet balance. ### How do I pay in BDT? Prices are shown in USD, BDT, with BDT as the default. Payment methods currently available: Manual Payment (bKash, Nagad). Your wallet is kept in USD. When you pay in BDT, the amount is converted at the exchange rate locked at checkout. The minimum top-up is $5 in dollars or ৳500 in taka. See [Paying in BDT](/docs/paying-in-bdt). ### What happens when my balance runs out? Requests fail with `402 insufficient_credits` and a link to [/dashboard/billing](/dashboard/billing). Before that point, when your balance can't cover a request's full `max_tokens`, the gateway lowers `max_tokens` to what the balance covers, so replies can come back shorter than you expect. A low-balance alert (default $5) and usage alerts at 50/75/90/100% are under Notifications in the dashboard. Subscriptions are paid per period and don't renew on their own, so renew before the period ends. See [Plans and wallet](/docs/plans-and-wallet). ### How do I get a receipt? Every payment creates a receipt (numbered `REC-XXXXXXXX`) under [/dashboard/billing](/dashboard/billing). Open it and print it or save it as a PDF from your browser. Refunds are handled by our team under the policy at [/refund-policy](/refund-policy). ## API keys ### Can I limit what a key can do? Yes, when you create it. Each key can have a monthly spend cap in USD and a list of allowed models. These can't be edited after creation, so to change them, create a new key and revoke the old one. Rate limits (requests per minute and concurrent requests) apply to your whole account, not per key. See [API keys](/docs/api-keys) and [Rate limits](/docs/rate-limits). ### How do I rotate a key? In [/dashboard/keys](/dashboard/keys), rotate the key to get a new secret. The old secret stops working **immediately**. There's no grace period. To avoid downtime, create a second key first, move your apps and agents to it, then revoke the old key. ## Account and reliability ### How do I delete my account? Contact [support](/docs/support) and ask for deletion. It's done manually by our team. Revoke your API keys first if you want them disabled right away. ### Is there an SLA? There's no published SLA. [/status](/status) shows current status and 90-day uptime for the Inference API, Dashboard & Auth, Billing and upstream providers, measured by real probes. If a provider fails, the gateway automatically fails over to another source for the same model where one is available. ### How do I get help? Open a ticket from [/dashboard/support](/dashboard/support). Include the `x-tokens-request-id` header from the failing response, the model ID and the time. Never include your API key. See [Support](/docs/support). --- # Dashboard tour > A page-by-page tour of the Tokens dashboard: what each page is for, the one thing to do there, and where the details are documented. Section: Dashboard. Page: https://tokens.bd/docs/dashboard-tour The dashboard at [/dashboard](/dashboard) is where you pay, create API keys, connect your tools and check what you spent. This page walks through the sidebar page by page. For each page it says what the page is for and the one thing to do there. Pages with their own guide link to it. Everything here needs a signed-in account. The sidebar has three groups: Workspace, Account and Resources. On a phone the sidebar is behind the menu button in the top bar. ## First steps The Overview page shows a checklist until you finish it. Do these in order: 1. **Verify your email.** You cannot create keys before this. 2. **Choose a plan or top up wallet.** A key needs an active plan or a positive wallet balance. 3. **Create an API key.** See [API keys](/docs/api-keys). 4. **Connect your agent.** Set the base URL and key in your tool. The Connect your agent page, described below, does this for you. 5. **Execute your first request.** Send one from your tool or from the [playground](/docs/playground). ## The top bar The page title and the **Docs** and **Website** buttons are always in the top bar, next to the theme toggle. If teams are enabled for your account, an organization switcher sits next to the title. It shows which organization you are working in and lets you change it. See [Teams and roles](/docs/teams-and-roles). ## Workspace ### Overview [/dashboard](/dashboard) is your summary. Its title is "Developer Dashboard". It has: - three cards: **Prepaid Wallet** (balance, with a "Manage & Top Up" link), your plan (credits remaining this period, with an "Upgrade" link) and **API Credentials** (number of active keys and your plan's key limit); - **Usage limits**, one meter per usage window on your plan, with its reset time; - **Usage by model this period**; - **Active API Keys**, with a **Revoke** button per key; - a chart of spend and requests, and **Recent Completion Requests**. The one thing to do here: press **Create API Key** if you have none, otherwise check the plan card before a long session. **Refresh** reloads all the numbers. ### Usage [/dashboard/usage](/dashboard/usage) is for finding out where your spend went: totals for the current billing period, per-model and per-day breakdowns, and your recent requests. The one thing to do here: press **Export CSV** to download the last 30 days. Details are in [Usage, limits and alerts](/docs/usage-and-alerts). ### Playground [/dashboard/playground](/dashboard/playground) sends a single prompt to any model you can use, from the browser. The one thing to do here: pick a key and a model and press **Send Request**. See [Playground](/docs/playground). It costs money like any other request. ### Model catalog [/dashboard/models](/dashboard/models), titled "Model Catalog & Pricing Matrix", lists every model with its alias, context window, price and the plan tiers that include it. The one thing to do here: copy the model id you want to use in your tool. See [Model catalog](/docs/model-catalog) and [Choosing a model](/docs/choosing-a-model). ### Wallet & ledger [/dashboard/wallet](/dashboard/wallet), titled "Prepaid Wallet & Accounting Ledger", shows your wallet balance and **Transaction history**: every top-up, charge and credit, with the request or payment reference and the balance after each entry. The one thing to do here: use **Deposit USD** or **Deposit BDT** to add funds (a button appears for each currency that is enabled). See [Plans, credits and wallet](/docs/plans-and-wallet) and [Paying in BDT](/docs/paying-in-bdt). ### API keys & limits [/dashboard/keys](/dashboard/keys), titled "API Keys & Gateway Protection", is where you manage credentials. The one thing to do here: press **Generate Key**. Each key row has actions to rotate and revoke it. Below the keys, the **Security Audit Log** lists security events such as key rotations, with a filter for severity. See [API keys](/docs/api-keys). ### Connect your agent [/dashboard/connect](/dashboard/connect) gets a coding agent talking to Tokens. It has three parts: - **Fastest: one command.** A command to run in a terminal that signs in through the website and writes the settings for the supported agents. It needs Node.js 18 or newer. See [Tokens CLI](/docs/tokens-cli). - **Or set it up by hand.** Pick your agent, paste a key (or press **Create a new key for** your agent) and choose a model. The page shows the exact settings with your values filled in. - **Test the connection.** **Run test** sends one tiny request with the key and model you chose and tells you whether it worked. The one thing to do here: pick your agent and copy the settings. Existing keys cannot be shown again, so paste one you saved or create a new one. ## Account ### Team & members [/dashboard/team](/dashboard/team) is for sharing an account with other people: members, roles, per-member spending caps and a per-member usage table. The item only appears in the sidebar when teams are enabled for your account. The one thing to do here: press **Invite Member**. See [Teams and roles](/docs/teams-and-roles). ### Billing & plans [/dashboard/billing](/dashboard/billing), titled "Subscription & Billing", has three tabs: **Subscription Plans**, **Prepaid Wallet & Ledger** and **Payments History**. The one thing to do here: pick a plan under **Subscription Plans** and pay. The checkout has fields for a coupon and for a referral code. See [Plans, credits and wallet](/docs/plans-and-wallet). ### Notifications [/dashboard/notifications](/dashboard/notifications), titled "Email Alerts & Notifications", controls which emails you get: usage warnings, low wallet balance, renewal reminders and receipts. The one thing to do here: check that usage warnings are on before you leave an agent running unattended. See [Notifications](/docs/notifications). ### Referrals & rewards [/dashboard/referrals](/dashboard/referrals) holds your referral link and shows what you have earned. The one thing to do here: copy your link with **Copy Link**. You need an active paid plan to get one. See [Referrals](/docs/referrals). ### Settings [/dashboard/settings](/dashboard/settings) holds your profile, your billing currency, notification preferences, sign-in security and account deletion. The one thing to do here: set the billing currency you want to pay in. See [Account settings](/docs/account-settings). ### Account security `/account/security`, titled "Account Security", has **Two-Factor Authentication** (an authenticator app) and **Active Sessions** (the devices signed in to your account). The one thing to do here: turn on two-factor authentication. See [Account settings](/docs/account-settings). ## Resources ### Docs & agent setup Opens this documentation, which has a setup guide for each tool. Start with [Claude Code](/docs/claude-code) or [other tools](/docs/other-tools) if yours is not listed. ### Support tickets [/dashboard/support](/dashboard/support), titled "Support", lists your tickets. The one thing to do here: press **New ticket** when something is wrong. Tokens emails you when it replies. See [Support](/docs/support). ### Help & support Opens a support dialog without leaving the page you are on. ### Website Takes you back to the public site. ## Related - [Quickstart](/docs/quickstart) - [API keys](/docs/api-keys) - [Usage, limits and alerts](/docs/usage-and-alerts) - [Teams and roles](/docs/teams-and-roles) --- # Teams and roles > Create a team organization on Tokens, invite members, give them a role and a monthly spending cap, switch between organizations, and remove a member. Includes what each role can do and what a team shares. Section: Dashboard. Page: https://tokens.bd/docs/teams-and-roles An organization lets several people work under one Tokens account: shared keys and usage reporting, a role for each person, and a monthly spending cap for each member. This page covers creating a team, inviting people, what each role can do, spending caps, switching between organizations and removing a member. :::note[Teams may not be enabled for your account] Teams are a feature that Tokens switches on per deployment. When it is off, the **Team & members** item is missing from the dashboard sidebar and the organization switcher is hidden. If you open [/dashboard/team](/dashboard/team) anyway, it says "Teams and collaborative workspaces are coming soon." Ask [support](/docs/support) if you need it. ::: ## Personal and team organizations Every account has a **Personal** organization, created automatically. Your keys, plan and wallet live there until you create a team. A personal organization has no members to invite: **Invite Member** is only shown in team organizations. A **Team** organization is one you create. You become its **owner**. Everything below happens in the organization that is active in the switcher, so check the name and the Team or Personal tag in the top bar before you create a key or invite someone. ## Create a team 1. Open the organization switcher in the dashboard top bar. 2. Press **Create Team Organization**. 3. Enter the **Team / Organization Name** and press **Create Team**. The dashboard reloads with the new team active and you as its owner. Open [/dashboard/team](/dashboard/team) to manage it. Owners and admins can also open **Org Settings** there to change the **Organization Name**, **Billing Email (Receipts)**, **Company / Legal Name** and **BIN / VAT / Tax ID**. ## Invite members Invitations are links. Tokens does not email them for you. 1. In the team organization, open [/dashboard/team](/dashboard/team) and press **Invite Member**. 2. Enter the **Email Address**, choose a **Role** and optionally a **Monthly Spend Cap (Optional USD)**. 3. Press **Generate Invitation**. 4. Copy the **Invitation Link (valid for 7 days)** and send it to the person yourself. The invited person opens the link while signed in to a Tokens account that uses the same email address, and presses **Accept and join organization**. A link issued to one address does not work for another one: the error says which address the invitation was issued to. After joining, they open the organization switcher and choose the team to start working in it. Things that can stop an invitation: - **Seat limit.** The number of members comes from the plan active on the organization. Without a plan, the limit is 1, which is the owner. Past the limit, you get "Organization has reached its plan seat limit of N members. Upgrade your plan to invite more team members." Pending invitations count towards the limit until they are accepted or expire. - **Already a member.** "This user is already a member of the organization". - **Expired or used link.** Generate a new invitation. :::tip Set the member's spending cap in the members table after they have joined, and check that it shows there. See the section on per-member spending caps below. ::: ## Roles and what each can do There are five roles. The first four can be chosen when you invite someone or change a role; the owner is the person who created the team. | Action | Owner | Admin | Billing | Developer | Viewer | | ---------------------------------------- | :---: | :---: | :-----: | :-------: | :----: | | See the member list | Yes | Yes | Yes | Yes | Yes | | See the usage table for all members | Yes | Yes | Yes | Own row only | Yes | | Create API keys | Yes | Yes | No | Yes | No | | See keys in the organization | All | All | None | Own keys | None | | Rotate or revoke keys | Any | Any | No | Own keys | No | | Invite members, change roles, set caps | Yes | Yes | No | No | No | | Remove members | Yes | Yes | No | No | No | | Edit Org Settings | Yes | Yes | No | No | No | The invitation form describes the roles like this: - **Developer**: "Create & manage own keys, view own usage". - **Admin**: "Manage members, keys, view all usage". - **Billing**: "Manage subscription & top-ups, invoices". - **Viewer**: "View org usage & models, read-only". Limits on admins: - An admin cannot change another admin's role, and cannot remove another admin. - An admin cannot set a spending cap for themselves or for the owner. - Nobody can demote or remove the owner from the dashboard. The role list for a member has Admin, Billing, Developer and Viewer only. Billing and viewer members cannot create keys and see no keys, so they cannot run requests in the team. In the [playground](/docs/playground) they see "No active API keys found" for the team. ## Per-member spending caps A cap limits how much one member's requests can spend in a calendar month. Set it in the **Monthly Spend Cap** column of the members table (in dollars), or leave it blank for no cap. The table saves when you leave the field. How it is enforced: - Tokens adds up what that member has spent in the team since the start of the month, plus requests still running, plus the most the new request could cost. If the total would pass the cap, the request is refused before it reaches the model. - The refusal is `402 member_cap_reached`. The message tells the member to ask the organization administrator to raise the limit. It is listed in [Errors](/docs/errors). - The cap is per person, not per key. Individual keys can also have their own monthly cap; both are checked. See [API keys](/docs/api-keys). - A cap of 0 is treated as no cap. Use a small positive number to block a member. The **Usage by Member (Current Month)** table on the same page shows each member's requests, total tokens, month spend in USD and cap. ## Switch between organizations Use the organization switcher in the top bar. It lists every organization you belong to, with a check mark on the active one. Choosing another one reloads the dashboard. Keys you create, usage you see and the team page all follow the active organization. You can belong to several teams at once and keep your Personal organization as well. ## What is shared and what is not Per organization, shared by its members: - **API keys.** A key belongs to the organization it was created in. Which keys you can see depends on your role (see the table above). - **Usage records.** Every request is recorded against the organization and the member who made it. The usage-by-member table on the team page reads from them. - **Funding.** Requests made with a team key are charged against the wallet and plan of the organization the key belongs to. The invitation page says members issue keys "funded by the organization's credit pool". Not shared: - **Your Personal organization.** Its keys, plan and wallet stay yours and are not visible to the team. - **Your profile, email and sign-in.** Teammates see your name and email in the member list and nothing else. - **Referral earnings.** They are paid to your own wallet. See [Referrals](/docs/referrals). :::warning[Check the funding before you roll out] The seat limit and the team's funding come from the plan and wallet attached to the team organization, not from your personal ones. Before inviting a whole team, make a test request with a team key and check that it is charged as you expect, or ask [support](/docs/support) how your team is funded. ::: ## Remove a member Owners and admins can press **Remove** on a member row. A confirmation dialog titled "Remove Team Member" warns: "All API keys created by this member will be permanently revoked immediately." Press **Confirm & Revoke Keys** to continue. When you remove someone: - Every active key they created in this organization is revoked at once. Tools using those keys get `401 invalid_api_key` on the next request. - They lose access to the team's keys, usage and settings. - Their own account and Personal organization are untouched. If the removed team was their active organization, they are moved back to their Personal one. If you want to keep a key a member made, create a new key yourself first and move your tools to it. Revoked keys cannot be restored. ## Related - [API keys](/docs/api-keys) - [Usage, limits and alerts](/docs/usage-and-alerts) - [Plans, credits and wallet](/docs/plans-and-wallet) - [Dashboard tour](/docs/dashboard-tour) --- # Referrals > How the Tokens referral program works: get your link, what the person you invite gets, what you earn and when it is paid, the rules and limits, and where to track your referrals in the dashboard. Section: Dashboard. Page: https://tokens.bd/docs/referrals The referral program pays you in wallet credit when someone you invite buys a paid plan, and gives them a discount on that purchase. This page explains how to get your link, what each side gets, when the reward is granted, the rules that apply and where to follow your referrals. :::note[The numbers are set by Tokens and can change] Discount and commission rates, the hold period and the limits below are program settings. This page tells you where to read the current values in the dashboard instead of repeating them. Where it gives an example figure, it says so. ::: ## Get your referral link Open [/dashboard/referrals](/dashboard/referrals), titled "Developer Referrals & Rewards". You need an active **Weekly or Monthly** paid plan to have a referral code. Without one, the page shows "Active Paid Membership Required" and an **Upgrade Subscription** button. With a qualifying plan, the page shows "Your Personal Referral Link & Share Kit": - your link, in the form `/r/` on the Tokens site, with a **Copy Link** button; - your **Referral Code**, with **Copy Code Only**; - two ready-made share messages, **Copy Bangla Message** and **Copy English Message**; - a **Personal QR Code** for your link. The badge at the top of the card shows the commission rate on your code ("Cash Credit Rate"). The discount the person you invite gets is the one shown on your link's landing page, so open your own link to see exactly what they see. ## What the invited person gets When someone opens your link, Tokens remembers your code in their browser for 30 days. The first referral link they open wins; a later link does not replace it. The link's landing page shows the discount for your code and a button to sign up. If they sign up and later buy a plan, the discount is taken off the plan price at checkout. On the Billing page the **Referral Partner Attribution** field is filled in for them. Someone who did not arrive through a link can type a code into **Referral Code (Optional)** at checkout instead. Rules for the discount: - It applies to a plan purchase, not to a wallet top-up. - It does not stack with a coupon. If the person enters a coupon too, the bigger of the two discounts is used, and if the coupon wins, the referral does not count for you. - A code only works for a limited time after the person signs up (30 days by default). After that, checkout does not accept it. ## What you get When the person you invited pays, you earn a commission. It is a percentage of the amount they actually paid, after their discount. A purchase that costs nothing after discounts earns nothing. Example, with made-up rates of 10% discount and 10% commission: a plan that costs 20 is bought for 18, and you earn 1.80. The commission is paid into your wallet, in the currency of the payment, as a wallet credit. It is not a cash withdrawal. It appears in your wallet's transaction history as a bonus credit ("Referral commission matured"). You can earn more through two extras, both shown on the Referrals page when they apply: - **Milestone Volume Ladder & Bonus Rewards.** Tiers you reach after a number of distinct approved referrals, each with a wallet bonus and sometimes a higher commission rate. A progress bar shows how many more you need for the next tier. - **Flash Event** boosters. Time-limited promotions that add a bonus to each referral during the event. The banner shows what you get and when it ends. A commission rate, with any milestone boost, never goes above a program maximum (30% by default). ## When rewards are granted A new referral does not pay at once. It goes through these states, shown in the **Status** column of the conversions table: | Status | What it means | | --------------------- | ----------------------------------------------------------------------------------------------- | | **In 14 Day Escrow** | The person has paid. The commission is held for a hold period (14 days by default). | | **Approved & Credited** | The hold ended and the commission was added to your wallet ("Settled to Wallet"). | | **Under Review** | The referral looked unusual and is being checked ("Routine fraud check in progress"). | | **Revoked** | The commission was cancelled. | When the hold ends, Tokens checks two things before it pays: the payment must still be completed (not refunded or reversed), and you must still hold an active Weekly or Monthly paid plan. If either is no longer true, the referral is revoked and nothing is paid. A scheduled job does the payout, so the credit can arrive some time after the date shown in the **Escrow Maturation** column. That column shows the real date for each referral, even if some text on the page mentions a fixed number of days. ## Rules and limits These are the program rules in force today. Rates and limits are program settings and can change. - **Paid plan required.** Only subscribers on an active Weekly or Monthly paid plan get a code and earn commission, and they must stay on one until the hold ends. - **No self-referrals.** You cannot refer your own account or an account with your email address. - **Fraud checks.** Referrals that look related to you, for example the same network or device, or many sign-ups from one place in a day, are held for review instead of paid straight away. - **Daily limit.** There is a cap on how many referrals one referrer can get in 24 hours (20 by default). - **Time limit.** The invited person must buy within a set time after signing up (30 days by default). - **Payment must hold.** If the payment is refunded or reversed before the hold ends, the commission is cancelled. ## Track your referrals [/dashboard/referrals](/dashboard/referrals) shows: - four cards: **Referral Link Clicks** (unique visits on your link), **Attributed Sign-ups**, **Paying Customers** and **Total Earnings** (split into paid and pending); - **Program Rules & Automatic Escrow Maturation**, a short summary of the steps above; - **Referral Conversions & Ledger**, one row per referral with the date, the referee (shown as a short id, not a name), the commission, the status and the maturity date. Earned credit shows in your wallet at [/dashboard/wallet](/dashboard/wallet), in **Transaction history**. ## Related - [Plans, credits and wallet](/docs/plans-and-wallet) - [Dashboard tour](/docs/dashboard-tour) - [Support](/docs/support) --- # Playground > Send a prompt to any model you can use from the Tokens dashboard, see the answer and what the request cost, and copy the same request as cURL, Python or a coding-agent config. Covers billing, settings and limits. Section: Dashboard. Page: https://tokens.bd/docs/playground The playground is a page in the dashboard where you send one prompt to a model and read the answer, without writing any code. Use it to try a model, check that a key works, or draft the request you will later make from your own code. This page covers how to use it, what it costs, what each setting does, how to copy the request as code and the limits. Open [/dashboard/playground](/dashboard/playground), titled "Developer Inference Playground". ## Run a request 1. Under **Inference Parameters**, choose an **API Key Context**. Only your active keys are listed. If you have none, the page links to [API keys & limits](/dashboard/keys) so you can generate one. 2. Choose a **Model Target**. The model's context window is shown next to the label. Models that your plan does not include are shown with a lock and "(Upgrade Required)" and cannot be chosen. 3. Optionally edit the **System Prompt (Optional)**. The default is "You are a helpful and concise coding assistant." 4. Write your **User Prompt**. A counter under the box shows characters and an estimate of tokens. 5. Press **Send Request**, or press Ctrl+Enter (Cmd+Enter on macOS) in the prompt box. The answer appears under **Model Output** when the model has finished. It is not streamed. A **Copy** button copies the answer. Below it, the stats bar shows what the request cost and how long it took. Each run is one single-turn request: your system prompt and your user prompt. The playground has no conversation history and no tool calling. For those, call the API from [curl](/docs/curl) or an SDK. ## What it costs The playground is not free and not separate from your balance. A playground request goes through the same metering as a call to `/v1`, so it is paid the same way: from your plan's credits when your plan covers the model, and from your wallet when it does not or when your plan is set to fall back to the wallet. See [Plans, credits and wallet](/docs/plans-and-wallet). - After each run, the stats bar shows "This request cost" and the amount in dollars. That figure is what was charged. - The run also appears in your usage, like any other request. - Your key's own limits apply: its allowed-model list and its monthly spend cap. See [API keys](/docs/api-keys). - Before the request starts, the same checks as the API apply: a used-up usage window, an empty wallet or an outstanding debt stop it with an error. - Some plans are set not to count playground use. On those plans the playground does not use your credits or wallet. If the amount shown is zero and your model is not a free one, ask [support](/docs/support) whether your plan counts playground use. Lowering **Max Tokens** lowers the most a single run can cost, because that is the longest answer you allow. ## What the settings do | Setting | Range in the page | What it does | | --------------- | ----------------------------------- | -------------------------------------------------------------------------------------------------------------- | | **Temperature** | 0 to 1.5, steps of 0.05, default 0.7 | Lower values make answers more predictable, higher values more varied. 0 is the most repeatable. | | **Max Tokens** | 256 to 4096, steps of 128, default 2048 | The longest answer the model may write. If the answer is cut off, raise it. | | **System Prompt** | Free text | Instructions that apply to the whole request, such as a role or output format. Leave it empty to send none. | The playground caps Max Tokens at 4096. If you need longer answers, call the API from your own code. ## Copy the request as code The **Integration Snippets** box under the settings builds the request you just configured. Your **Model Target**, **System Prompt**, **User Prompt** and **Temperature** are filled in, and the address is the Tokens address. Choose a tab and press **Copy**: - **cURL** - **Claude Code** - **Cursor** - **Cline** - **OpenCode** - **Aider** - **Python** The key in every snippet is the placeholder `YOUR_API_KEY`. The page never shows your secret, and existing keys cannot be shown again, so put your own key in before you run it. Keep it in an environment variable and out of files you commit (see [API keys](/docs/api-keys)). The cURL tab produces a request of this shape: ```bash curl -X POST "https://tokens.bd/v1/chat/completions" \ -H "Authorization: Bearer $TOKENS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek/deepseek-v4.1-flash", "messages": [ {"role": "system", "content": "You are a helpful and concise coding assistant."}, {"role": "user", "content": "Say hello in one short sentence."} ], "temperature": 0.7 }' ``` The snippet does not include Max Tokens. Add `"max_tokens": 2048` to the body if you want the same limit as in the page. The request format is described in [Chat Completions](/docs/chat-completions). For the agent tabs, the full setup pages are better: see [Claude Code](/docs/claude-code), [OpenCode](/docs/opencode) and the other [coding agents](/docs/other-tools). ## Limits - **10 requests per minute.** The playground has its own limit, separate from the one on your API calls. Over it you get a rate-limit error; wait for the next minute. - **One prompt per run.** No chat history, no streaming, no tools, no images. - **Max Tokens up to 4096.** - **Needs an active key.** Revoked or suspended keys are not listed. - **In a team,** only owners, admins and developers have keys, so billing and viewer members have nothing to run with. See [Teams and roles](/docs/teams-and-roles). - **Usage windows.** If your plan's usage window is already full, the playground refuses the request with an `insufficient_quota` error and says when it resets. See [Usage, limits and alerts](/docs/usage-and-alerts). ## If a run fails The red "Execution Failed" box shows the message and, when there is one, a **Request ID**. Quote that ID to [support](/docs/support). | Message or symptom | Cause and fix | | ----------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ | | "Please select an API key to run completions." | No key is selected. Generate one on the keys page. | | "The selected API key is either invalid, expired, or does not belong to you." | The key was revoked or is not yours. Pick another key. | | Wallet or credit error (402) | Your plan or wallet cannot cover the request. Top up or change plan. See [Errors](/docs/errors). | | Rate limit error (429) | More than 10 runs in a minute, or a usage window is full. Wait and retry. | | A model is greyed out with a lock | Your plan does not include it. Pick another model, or see the [model catalog](/docs/model-catalog). | ## Related - [Quickstart](/docs/quickstart) - [Choosing a model](/docs/choosing-a-model) - [Dashboard tour](/docs/dashboard-tour) --- # Notifications and email alerts > Every email Tokens sends about your account: which ones you can switch on or off on the Notifications page, what triggers each, the thresholds, the defaults, and how to read the delivery history. Section: Dashboard. Page: https://tokens.bd/docs/notifications Tokens tells you about money and limits by email, so a long agent session does not end with a surprise. This page lists every email the platform sends to your account, which of them you can switch off, what triggers each one, and how to check that an email was sent. For the usage numbers behind the alerts, see [usage, limits and alerts](/docs/usage-and-alerts). ## Where to find the settings Open [Notifications](/dashboard/notifications) in the dashboard sidebar, under Account. The page is titled **Email Alerts & Notifications** and has three parts: - **Billing & Lifecycle Alerts**: receipts, renewal reminders, payment failures and the low wallet balance alert. - **Token Quota Milestones**: warnings as you use up your plan's credits. - **Delivery History & Audit Trail**: a log of the emails sent to you. Each switch is a checkbox that saves as soon as you click it. The page shows "Notification preferences successfully updated and saved." when it worked. If the settings cannot be loaded, the switches stay disabled and the badge next to the title reads Unavailable; press **Refresh** to try again. The **Notifications** tab in [Settings](/dashboard/settings) shows the same topics as a read-only summary, with a link to this page. Its "Enabled" badges are fixed text and do not follow your switches, so use the Notifications page as the source of truth. :::note All alerts are emails. They go to the email address on your account. Tokens has no SMS, push or in-dashboard notification for these alerts. ::: ## Alerts you can switch on or off Everything below is **on by default** for every account. Your choice is stored per account; it is not per API key. | Switch on the page | What triggers it | Sent | | ----------------------------- | ------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- | | Checkout & Wallet Receipts | A subscription checkout or a wallet top-up is confirmed by the payment provider | Right after the payment is confirmed | | Subscription Renewal Reminders | Your paid plan's period ends within 3 days | Once per plan period | | Payment Failure Notices | A payment is reported as failed or cancelled by the payment provider | Right after the failure is reported | | Low Wallet Balance Alerts | Your USD wallet balance is below $5.00 | At most once per day while the balance stays below the threshold | | Token Quota Milestones | The share of your plan's credits you have used reaches 50%, 75%, 90% or 100% | Once per threshold per plan period | ### Checkout & Wallet Receipts The label says "Receive verified receipts for subscription checkouts and prepaid credit deposits." The email is sent when the payment provider confirms the payment, for plan purchases and for wallet top-ups. It carries the invoice number, amount, currency and payment method. A printable receipt is also in [billing](/dashboard/billing). Manual payments (a bank or mobile-wallet transfer you submit with a transaction id) have their own emails, listed in the next section. ### Subscription Renewal Reminders Sent about 3 days before your plan period ends, with the plan name, the date and the price. A few rules from the code: - Only paid plans get a reminder. Free plans and plans with a price of zero do not. - If the next period is already paid for, no reminder is sent. - You get one reminder per period. Plans do not renew by themselves, so this email is your cue to pay for the next period. See [plans, credits and wallet](/docs/plans-and-wallet). ### Payment Failure Notices Sent when a payment you started is reported as failed or cancelled. The email says the payment did not go through and gives the amount. Start the payment again from [billing](/dashboard/billing). ### Low Wallet Balance Alerts The page says "Receive an email when your prepaid wallet drops below $5.00 USD to prevent API request dropouts." Details: - The check covers your **USD wallet**. A BDT wallet balance does not trigger this email. - The threshold is $5.00. The page has no field to change it. - While the balance stays below the threshold you get at most one email per day. Top up and the alerts stop. The [Wallet](/dashboard/wallet) page shows its own "Low Balance" badge, which uses a different limit ($1 for the USD wallet). The badge and the email are independent. ### Token Quota Milestones These emails warn you before your plan's credits run out. The percentage is the share of your active plan's credit pool for the current period that has been used. The four rows are: | Row | Label on the page | | ----- | ---------------------------------- | | 50% | Halfway threshold reached | | 75% | Three quarters threshold warning | | 90% | Critical quota warning | | 100% | Quota exhausted notification | How it behaves: - **Enable All** (top right of the card) is a master switch for the whole group. Turn it off and the four row switches become disabled, and none of the milestone emails are sent. Turning it back on restores whatever the four rows were set to. - Each row can be switched on its own. - Each milestone is sent once per plan period. If your usage jumps past several milestones between two checks, you get one email for each. - If you switch a milestone on after you have already passed it, the email is sent on the next check. - These emails are for accounts with an active plan. They look at the plan's credit pool, so they do not apply to a wallet-only account. The alerts are produced by a scheduled check that runs about every five minutes, so an email can arrive a few minutes after you cross a threshold, not at the exact request. A failed send is retried up to 3 times in the same period. ## Emails with no switch on the Notifications page ### Usage window and model allowance warnings If your plan has usage windows (for example a 5-hour session or a weekly ceiling) you also get a warning at **80%** and a notice at **100%** of a window. When a plan sets a spending limit on one model, you get the same two emails for that model ("... 80% of its allowance used", "... allowance used up"). These are on by default. The Notifications page does not show a switch for them, so you cannot turn them off from the dashboard. Each window or model sends each of the two emails once per window or period. To see the windows themselves, use the Usage page or `GET /v1/tokens/usage` as described in [usage, limits and alerts](/docs/usage-and-alerts). ### Payment and account emails These are always sent when the event happens: | Email | When | | --------------------------------------- | -------------------------------------------------------------------- | | We received your payment details | You submit a manual payment (a transaction id) for a plan or top-up | | Payment approved | Staff confirm your manual payment | | Payment rejected | Staff cannot verify your manual payment; the email gives the reason | | Confirm your email address | You create an account | | Reset your password, sign-in link | You ask for a password reset or a sign-in link | | Confirm your email change | An email change is requested for your account | | Support ticket reply, ticket closed | Staff answer or close a ticket you opened | Support emails link back to [your tickets](/dashboard/support); the full conversation stays in the dashboard. ## Delivery history The **Delivery History & Audit Trail** table lists the emails recorded for your account, newest first. The badge shows how many records were loaded. The page loads the latest 20. Columns: | Column | Meaning | | ---------- | ---------------------------------------------------------------- | | Event Type | The internal name of the email, for example `checkout_receipt` | | Subject | The subject line | | Recipient | The address it was sent to | | Status | `sent`, `failed`, `suppressed`, or `sent (test mode)` | | Timestamp | When it was recorded, in your browser's time zone | What the statuses mean: - **sent**: handed to the email provider. - **failed**: the provider or the mail server refused it. The platform retries milestone, renewal and window alerts up to 3 times per period. - **suppressed**: it was due, but you had that alert switched off (or your address could not be resolved), so nothing was sent. Receipts, payment failures and similar alerts leave such a row. Quota milestones, renewal reminders and low balance alerts are checked before sending, so a switched-off alert simply leaves no row. - **sent (test mode)**: the platform's email provider is set to test mode, so the message was recorded but not delivered. An empty table shows "No notification events recorded yet. Complete a checkout to create one." Account emails such as the confirmation link and support ticket emails are not tied to your user record in the log, so they do not appear in this table. ## If an email does not arrive 1. Check the history table. If the row says `suppressed`, switch the alert back on. 2. If there is no row, the event may not have happened yet. Milestones and reminders are produced by the scheduled check, so allow a few minutes. 3. Look in your spam folder, and make sure the address in [Settings](/dashboard/settings) is the one you read. 4. If the row says `failed`, or you expected an email and have no record, send the event type and time to [support](/docs/support). ## Related - [Usage, limits and alerts](/docs/usage-and-alerts) - [Plans, credits and wallet](/docs/plans-and-wallet) - [Account settings](/docs/account-settings) - [Getting help](/docs/support) --- # Reading the model catalog > How to read the public models list, a model's own page and the dashboard Model catalog: every column, badge and price field, what Available on means, and why the list your key sees can be shorter. Section: Dashboard. Page: https://tokens.bd/docs/model-catalog Tokens shows its models in three places: the public list at [/models](/models), one page per model at `/models/`, and the Model catalog inside your dashboard. They read from the same catalog but show different things. This page explains each field, what the catalog does not show, and how to find out which models a given key can call. ## Which page shows what | Page | Who can open it | Prices | Your access | | --------------------------------------------- | --------------- | ------ | ------------------------------------------------ | | [/models](/models) | Anyone | Yes | No | | `/models/` | Anyone | Yes | No | | [Model catalog](/dashboard/models) (dashboard) | Signed in | No | "Access Granted" or "Upgrade Needed" on each model | | `GET /v1/models` | A Tokens key | No | Exact list that key can call | Despite its heading, "Model Catalog & Pricing Matrix", the dashboard page has no price columns. Use the public pages for prices. Only active models are listed. Changes made to the catalog can take a few minutes to appear on the public pages. ## The public list at /models The page header reads "Models and prices" and counts the models and providers. The table has these columns: | Column | What it shows | | --------------- | ------------------------------------------------------------------------------------------------------------------- | | Model | The display name, linked to the model's page. The model id is printed below it. | | Provider | The model's provider label (see the note below). | | Context | The context window, written short: `256K`, `1M`. | | Input / 1M | Price for one million input tokens, with a bar that compares it to the most expensive price on the page. | | Output / 1M | Price for one million output tokens, with the same kind of bar. | | Available on | The plan tiers that include the model: Free, Weekly, Monthly, Pay as you go. | | Copy id | A button that copies the model id to your clipboard. | The id is what you put in the `model` field of a request. Copy it from here rather than typing it. :::note[About the Provider label] The label is worked out from the model's id and name, not stored as the model's maker. A model whose name is not recognised is grouped under "OmniRouter". Treat it as a rough filter, not as a statement of who trained the model. ::: ### What the price fields mean - **Input / 1M** is the price per one million tokens you send: your prompt, the conversation so far, file contents and tool results. - **Output / 1M** is the price per one million tokens the model writes back. - Prices are shown in one currency at a time. Use the currency switch to flip between USD and BDT. The switch offers only the currencies enabled on the platform. A BDT price is the BDT price set for the model; if none is set, the page shows the USD price multiplied by the platform's exchange rate. - Amounts show two decimals (for example `$0.30`). A price under one cent shows four. Prices change, so this page quotes none. Read the current figures on the pages above. ### Search, filter and sort - **Search models** matches the display name, the id and the provider label. - **Context window**: `Any context`, `256K+` or `1M`. The `1M` option keeps models with a context window of at least one million tokens. - **Provider chips** under the search box filter to one provider label. "All providers" clears the filter. - Click **Model**, **Context**, **Input / 1M** or **Output / 1M** to sort. A first click sorts names A to Z, context from largest, and prices from lowest. A second click on the same column reverses it. - The footer shows "Showing N of M models". If nothing matches, the table says "No models match these filters." ### Compare up to three models Tick the checkbox on up to three rows. A bar appears at the bottom of the page. The **Compare** button reads "Pick one more" until two models are ticked, then "Compare 2" or "Compare 3". The comparison window lists, side by side: | Row | Meaning | | -------------- | -------------------------------------------------------------------------------------------------------------------------------- | | Provider | The provider label | | Model id | The exact id | | Context window | Short form, for example `1M` | | Input / 1M | Input price | | Output / 1M | Output price | | Sample session | The cost of 25,000 input tokens and 1,500 output tokens, "a typical coding-agent turn". The cheapest is marked "lowest". | | Available on | The plan tiers | The address in your browser updates to `/models?compare=,` while you tick models, so you can share the comparison by copying the link. ## A model's own page Open a model from the list, or go to `/models/`. A model that is not active gives a not-found page. The page has: - The provider, the display name and a short description. If the catalog has no description, the page writes one sentence from the name, provider and context window. - The model id with a copy button, a **Get an API key** button and **Compare with others**. - Four tiles: | Tile | Meaning | | --------------------- | -------------------------------------------------------------------------------------------------------- | | Input, per 1M tokens | Input price in the main currency, with the other currency in small type below | | Output, per 1M tokens | Output price, in the same way | | Context window | Short form, with the exact token count below | | Available on | How many tiers include the model, with their names | - **What it costs in practice**: the cost of three workloads at the model's current price, in each enabled currency. These are examples, not forecasts; your own mix of input and output depends on the agent and the task. | Workload | Tokens | | ---------------------- | ----------------------- | | One agent turn | 25K input, 1.5K output | | An hour of agent work | 1M input, 60K output | | A month of daily use | 20M input, 1.2M output | - **Use this model**: ready-made setups for Claude Code, Codex, OpenCode, Cursor, the OpenAI Python SDK and curl, each with this model's id filled in. Replace the key with one of yours. See the [quickstart](/docs/quickstart). - **Similar models**: up to four, models from the same provider label first, then the closest output price, with a link that compares them. ## What the catalog does not show Check these before you plan around the catalog: - **No output limit.** The catalog has one size per model, the context window. It does not list a maximum output length. - **No cache price.** Only input and output prices are shown. The public pages and the dashboard Model catalog do not list a cache-read or cache-write price. Your usage export counts cache-read tokens in a separate column; see [usage, limits and alerts](/docs/usage-and-alerts). - **No capability badges.** There is no tag or filter for tool calling, image input, reasoning or streaming. The search box on the public page matches name, id and provider only. For what each model is good at, and which accept images, read [choosing a model](/docs/choosing-a-model). To see how a model behaves on your own task, open it in the [Playground](/dashboard/playground). ### Finding a model for a feature 1. Decide the constraint that matters: a large context, a low price, or an input type such as images. 2. For context, use the **Context window** filter (`256K+` or `1M`) on [/models](/models). For price, sort by **Input / 1M** or **Output / 1M**. 3. For features the catalog does not list (tool calling, vision, reasoning), start from [choosing a model](/docs/choosing-a-model), then send one test request to the model before you rely on it. Tool use in particular is covered in [tool calling](/docs/tool-calling). 4. Compare the shortlist with the comparison window, using the sample session cost as a rough guide. ## The dashboard Model catalog Open [Model catalog](/dashboard/models) in the sidebar. It tells you what your account can use. Three tiles at the top: | Tile | Meaning | | ------------------- | ---------------------------------------------------------------------------------------------------------------------------------- | | Active Coding Models | How many active models the catalog lists | | Catalog Source | `Live Database` when it loaded, `Loading` while it loads, `Unavailable` if the request failed | | Your Current Tier | The tier of your active plan in capitals (for example `MONTHLY`), or `PUBLIC` if you have no active plan. A `Wallet Active` badge appears when your wallet balance is above zero | Controls: - A text box ("Filter models by alias or capability...") that matches a model's id, name or description. It does not search a capability field, because the catalog has none; words in the description can match. - Tabs: **All**, **Free Starter** (models available on the Free tier), **Paid Plans** (Weekly or Monthly) and **Wallet Allowed** (Pay as you go). - **Grid** and **Table** views, and a refresh button. Each model shows its name, the context window as thousands of tokens (`1000k context` for a 1M window), its id with a copy button, and the raw tier names (`free`, `weekly`, `monthly`, `pay_as_you_go`). **Try in Playground** opens the [Playground](/dashboard/playground). Choose the model there; the Playground page does not read it from the link. The table view has the columns Model Name & Alias, Context Window, Allowed Tiers, Status and Actions. ### Access Granted and Upgrade Needed The status badge is **Access Granted** (**Available** in the table) when your wallet balance is above zero or your plan's tier is on the model's list. Otherwise it is **Upgrade Needed** (**Upgrade**). :::warning[Treat the badge as a hint] The dashboard works out the badge only when you have an active plan. With no plan, every model shows Upgrade Needed, even if you could call it with your wallet. The badge also does not check whether the model is enabled for Pay as you go. For the exact answer, call `GET /v1/models` with the key you plan to use (next section). ::: ## How the list differs by key The public pages are the same for everyone. The list of models a key can call is not. `GET /v1/models` returns the models for the key you send: ```bash curl -s https://tokens.bd/v1/models \ -H "Authorization: Bearer $TOKENS_API_KEY" | jq -r '.data[].id' ``` A model appears in that list when all of these hold: 1. It is active in the catalog. 2. If the key has an allowed-models list, the model is on it. See [API keys](/docs/api-keys). 3. Your active plan's tier is on the model's **Available on** list, or your wallet balance is above zero and **Pay as you go** is on the model's list. So two keys on one account can see different lists, and a model on the public page can be missing from your key. The reasons are covered in [models and usage endpoints](/docs/models-and-usage). Calling a model the key cannot use returns `model_not_allowed_on_key` or a plan error; see [errors](/docs/errors). ## Related - [Choosing a model](/docs/choosing-a-model) - [Models and usage endpoints](/docs/models-and-usage) - [API keys](/docs/api-keys) - [Plans, credits and wallet](/docs/plans-and-wallet) --- # Account settings > What the Settings page and Account security page let you do: view your profile, pick a default currency, set up two-factor authentication, review sessions, and request account deletion, with what happens to your keys, balance and records. Section: Dashboard. Page: https://tokens.bd/docs/account-settings Your account controls are in two places: [Settings](/dashboard/settings) and Account security (`/account/security`), both under Account in the dashboard sidebar. This page walks through every control, says what is not there, and explains what happens to your API keys, wallet balance and records when you ask for your account to be deleted. ## What is on the Settings page The page is titled **Settings**, and it has five tabs. | Tab | What it does | | ---------------- | ------------------------------------------------------------------------------------ | | Profile | Shows your name, email, user ID and role. Nothing on it can be edited. | | Default Currency | Saves your preferred currency: BDT or USD. | | Notifications | A summary of your email alerts, with a link to the Notifications page. | | Security & MFA | The same two-factor and session controls as the Account security page. | | Danger Zone | Where you request account deletion. | ## Profile The Profile tab, headed **Profile Information**, shows four read-only items: - **Full Name**: the name on your account, or "Not provided" if none was set. - **Email Address**: the address you sign in with and where alerts are sent. - **Account Identifier (User ID)**: your account's ID, with a button to copy it. Quote it when you write to [support](/docs/support). - **Role & Status**: a badge with your role (`user` for customers) and the date you joined. There is no form to change the name or email. To correct either, ask [support](/docs/support) and say which one you want changed. Email also matters for verification: you need a verified email address before you can create an API key (see [API keys](/docs/api-keys)). ## Password The Settings page has no change-password form. To set a new password: 1. On the sign-in page, choose **Forgot your password?**. 2. Enter your email address. You get an email with a **Reset password** button. 3. Open the link and enter a new password on the **Set a new password** screen. The email says to ignore it if you did not ask for it, and your password stays unchanged. If you sign in with Google or GitHub, your account password is managed by that provider. After changing a password, use **Sign out everywhere** (below) to end sessions you do not recognise. ## Connected accounts There is no screen to link or unlink accounts. The sign-in and sign-up pages show **Continue with Google** and **Continue with GitHub** buttons when the platform has turned those providers on. If you only see the email and password form, they are off. ## Default Currency The **Default Currency** tab ("Default Billing & Display Currency") offers a card for each currency enabled on the platform: - **৳ BDT**: Bangladeshi taka, for local checkout, mobile wallets and regional settlement. - **$ USD**: US dollars, for card and cryptocurrency billing. Click a card to choose it. The choice is saved at once ("Currency preference saved to your profile.") and any currency switch open on the page follows it. If an administrator has turned a currency off, its card is not shown, and saving it is refused with "Currency ... is currently disabled." The setting is your preferred pricing and checkout currency for subscriptions and wallet top-ups. It does not move money between wallets. Your USD and BDT balances stay separate; see [plans, credits and wallet](/docs/plans-and-wallet). ## Notifications The **Notifications** tab lists receipts, quota warnings, renewal reminders and low wallet balance alerts, each with an "Enabled" badge. The badges are fixed text, not live switches. The real switches, the thresholds and the delivery log are on the [Notifications page](/dashboard/notifications); see [notifications](/docs/notifications). ## Two-factor authentication Open the **Security & MFA** tab, or Account security (`/account/security`) directly. The **Two-Factor Authentication** card protects sign-in with a time-based code from an authenticator app (TOTP), such as Google Authenticator, 1Password or Authy. The badge on the card says: - **Active**: you have a verified authenticator. - **Optional**: you have none, and two-factor is not enforced on this platform. - **Required for Admin**: you have none, and staff accounts must have one. To turn it on: 1. Press **Set Up Authenticator App**. 2. Scan the QR code with your app. If you cannot scan, copy the secret key shown below the code and type it into the app. 3. Enter the 6-digit code from the app and press **Verify and Enable 2FA**. From then on sign-in asks for your password and a current code. **Remove** next to the authenticator turns it off. Tokens supports authenticator apps only; the page offers no SMS codes or recovery codes. Keep the app on a phone you will not lose. The [security and privacy](/docs/security-and-privacy) page explains why this is worth doing: your key and billing sit behind your sign-in. ## Active sessions The **Active Sessions** card, "Devices and sessions currently signed in to your account", shows the number of sessions and, for each: - the browser or device description, - the IP address, and - when it was last used. It lists three sessions and shows the rest under **View all N sessions...**. Use: - **Revoke** to end one session. That browser has to sign in again. - **Sign out everywhere** to end all your sessions, including the one you are using. You will need to sign in again. If you see a session you do not recognise, revoke it, change your password with the reset steps above, then check [your API keys](/dashboard/keys) for any you did not create. ## Theme and language - **Theme.** The dashboard header has a button that switches between light and dark. On a first visit the site follows your operating system's setting. When you press the button, your choice is stored in that browser, so it is per browser and device, not part of your account. - **Language.** There is no language setting. The dashboard is in English. These docs also exist in Bengali at `/bn/docs`. ## Delete your account You cannot delete the account with one click. Deletion is a request that the Tokens team carries out. 1. Open the **Danger Zone** tab and press **Delete Account**. 2. Read the **Request Account Deletion** dialog. It states that all active API keys will be immediately revoked, that refunds of any remaining prepaid balance must be requested before closure, and that this cannot be reversed. 3. Press **Confirm & Email Support**. This opens an email to the support address, with a subject like "Account Deletion Request - User ID ..." and your email and user ID already in the message. Nothing is deleted yet. If no email program opens, open a ticket in [support](/dashboard/support) from the account you want deleted and include the same details. The team then checks the account before closing it. No time frame is stated in the product. ### What has to be true first Deletion is refused until both of these hold: - **Every wallet balance is zero.** Wallets are per currency, so both USD and BDT must be at zero. If you have a balance, ask for a refund or an adjustment first; see the [refund policy](/refund-policy). Deleting the account does not refund anything. - **No wallet reservations are pending.** A reservation is money held for a request that is still being settled. They clear on their own, so stop your tools and try again shortly. Stop any agent that uses your keys before you ask, and save any data you want to keep, such as a [usage export](/docs/usage-and-alerts). ### What happens to your data When the team deletes an account, in this order: 1. The account's upstream model key, if one exists, is revoked. 2. The sign-in record is removed, so the account can no longer sign in. 3. The account record is anonymized in place: the email becomes a random address of the form `deleted-@tombstone.invalid`, the name becomes "Deleted user", the role is reset to `user`, and the verified-email date is cleared. 4. Every API key on the account is set to revoked ("Account deleted") and stops working at once. 5. An entry is written to the audit log with the reason. What stays: payment records, subscriptions, the wallet ledger, usage records and audit logs are not deleted, because they are financial and audit records. They remain attached to the same internal ID, but without your name or email. The account's email address is freed, so it can be used to register a new account later. The new account starts empty; it does not inherit the old keys, balance or history. What is not stored in the first place: Tokens does not keep the content of your prompts and responses. See [security and privacy](/docs/security-and-privacy). The action cannot be undone. A tombstoned account cannot be restored. ## Related - [Notifications and email alerts](/docs/notifications) - [Security and data privacy](/docs/security-and-privacy) - [API keys](/docs/api-keys) - [Getting help](/docs/support)