cURL is the fastest way to check that a key works and see exactly what the Tokens API returns, with no SDK in the way. Every example below reads your key from TOKENS_API_KEY, so set that first.
export TOKENS_API_KEY="tok_live_your_key"On Windows PowerShell, use $env:TOKENS_API_KEY = "tok_live_your_key" and call curl.exe rather than curl. In Windows PowerShell 5.1, curl is an alias for Invoke-WebRequest, which takes different flags. The \ line continuations below are for bash and zsh; in PowerShell put the command on one line or use a backtick.
Send a chat completion with cURL#
The OpenAI-compatible base URL is https://tokens.bd/v1. Authenticate with Authorization: Bearer.
curl https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"messages": [
{"role": "system", "content": "You are a terse code reviewer."},
{"role": "user", "content": "What does `set -euo pipefail` do?"}
],
"max_tokens": 300
}'The response is a standard chat completion object: the reply is in choices[0].message.content and token counts are in usage. Field-by-field details are in Chat Completions.
model is required. Request bodies can be up to 10 MB; anything larger returns 413.
Stream tokens with -N#
Add "stream": true and pass -N (--no-buffer) so cURL prints each server-sent event as it arrives instead of waiting for the whole response.
curl -N https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"messages": [{"role": "user", "content": "Count from 1 to 10, one number per line."}],
"stream": true,
"stream_options": {"include_usage": true}
}'You get data: {...} lines with choices[0].delta.content, then data: [DONE]. The usage chunk (with an empty choices array) only appears because the request sets stream_options.include_usage. Leave it out and the stream carries no token counts. More on the event format in Streaming.
Call the Anthropic Messages API#
The same key works on the Anthropic-compatible endpoint. Send it as x-api-key (or Authorization: Bearer, both are accepted).
curl https://tokens.bd/v1/messages \
-H "x-api-key: $TOKENS_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"max_tokens": 300,
"messages": [{"role": "user", "content": "Explain a git rebase in two sentences."}]
}'The model doesn't have to be an Anthropic model. The gateway translates the request for whichever provider serves the model you name. Add "stream": true and -N to stream Anthropic-style events. See Messages.
Note
SDKs and tools built for Anthropic usually want the base URL without /v1 (https://tokens.bd), because they append /v1/messages themselves. With raw cURL you type the full path.
List the models your key can use#
curl https://tokens.bd/v1/models \
-H "Authorization: Bearer $TOKENS_API_KEY"This returns the models available to this key right now, taking into account your plan, your wallet balance and the key's allowed-models list. Copy model IDs from here rather than typing them from memory. The full catalog with prices is at /models.
To pull out just the IDs with jq:
curl -s https://tokens.bd/v1/models \
-H "Authorization: Bearer $TOKENS_API_KEY" | jq -r '.data[].id'Check plan usage and wallet balance#
GET /v1/tokens/usage is specific to Tokens. It reports your plan, active usage windows, wallet balance and this key's limits. It is read-only and is never billed.
curl -s https://tokens.bd/v1/tokens/usage \
-H "Authorization: Bearer $TOKENS_API_KEY" | jqThe response has these top-level fields:
| Field | What it holds |
|---|---|
plan | Plan name, tier and periodEnd, or null if you have no active plan |
windows | Active usage windows (session_5h, weekly, monthly) with unit, limit, used, remaining, percentUsed, resetsAt |
wallet | balanceUsd, or null |
key | monthlySpendCapUsd and allowedModels for the key you used |
This is the call to make when a request fails with 429 window_exhausted and you want to know when the window resets.
Debug errors with -i and x-tokens-request-id#
Pass -i to print response headers along with the body. Every response carries x-tokens-request-id, which is the first thing support will ask for.
curl -i https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "not-a-real/model", "messages": [{"role": "user", "content": "hi"}]}'HTTP/2 404
content-type: application/json
x-tokens-request-id: 7f3c...
x-request-id: ...
{"error":{"message":"...","type":"...","code":"model_not_found","param":null,"request_id":"7f3c..."}}Read error.code, not just the HTTP status. A 403 can mean an inactive key, a model that isn't on the key's allow-list, or a spend cap you've hit, and each has a different fix. On 429, check the Retry-After header. Tokens doesn't send X-RateLimit-* headers.
For one-line status checks in scripts, -w prints just the code:
curl -s -o /dev/null -w "%{http_code}\n" https://tokens.bd/v1/models \
-H "Authorization: Bearer $TOKENS_API_KEY"200 means the key is valid. 401 means it's missing or wrong. The full list of codes and what to do about each is in Errors and Troubleshooting.
What isn't supported#
The gateway serves /v1/chat/completions, /v1/completions (legacy), /v1/messages, /v1/responses, /v1/embeddings (embedding models only), /v1/models and /v1/tokens/usage. Images, audio, files, batches, assistants, fine-tuning and moderations endpoints return 404 unsupported_endpoint.