If you build a product on Tokens, you will want to know which part of your spend belongs to which customer, and to stop one customer or one broken service from using up everything. Separate API keys are the tool Tokens gives you for this. This page says what a key can and cannot do, so you can decide where to use keys and where to do the work in your own app.
Read API keys first for the basics. This page builds on it.
What Tokens can and cannot do here#
| You might expect | What exists |
|---|---|
| Create a key with a spend cap | Yes, in the dashboard at /dashboard/keys |
| Restrict a key to some models | Yes, an allowed-models list set at creation |
| Create, rotate or revoke a key by API | No. There is no key-management endpoint you can call with a tok_live_ key |
| Set an expiry date on a key | No. The creation form has no expiry field. key_expired applies only to keys an administrator provisioned with an expiry |
| Edit a key's cap or models later | No. Both are fixed at creation |
| A separate rate limit for each key | No. Requests per minute and concurrency are per account, shared by all keys |
| Spend per key in the dashboard or an API | No. See Track spend per customer |
| Hundreds of keys on one account | Only if your plan allows it. The default is 3 active keys |
The dashboard talks to routes under /api/keys. They are authenticated by your signed-in browser session (and your multi-factor check, if you use one), not by a Tokens API key, and they are not a supported public API. Do not build a provisioning service on them. The Tokens CLI signs in through the browser and receives one new key per login; it is for setting up your own coding agents, not for issuing keys to customers.
What a key can be limited by#
When you create a key in the dashboard you can set:
| Setting | Effect |
|---|---|
| Name | Up to 64 characters. Use it to record whose key it is: acme-prod, staging, batch-worker. |
| Monthly spend cap (USD) | The most the key may spend in a calendar month. Left blank, the key takes your plan's default cap if the plan defines one. |
| Allowed models | The only model ids the key may call. Empty means every model your account can use. GET /v1/models lists only allowed models. |
A key also stops working when it is revoked or its account is suspended. Rate limits, concurrency and usage windows are not key settings; they belong to your account, see rate limits.
To read a key's own settings from code, call GET /v1/tokens/usage with that key. The key object in the response shows its monthlySpendCapUsd and allowedModels. The plan, windows and wallet parts of the same response describe your whole account, not the key.
curl -s https://tokens.bd/v1/tokens/usage \
-H "Authorization: Bearer $CUSTOMER_KEY" | jq .keyHow many keys you can have#
Each plan sets a maximum number of active keys, and the default is 3 (pay-as-you-go accounts also default to 3). Creating one more returns key_limit_reached. Revoked keys do not count.
This decides your design:
- A handful of services or environments (production, staging, a batch worker, an internal tool): one key each fits inside the default.
- Many customers: a key per customer works only if your plan's limit is high enough. Check your plan on pricing, or ask support what limit is possible for you. Do not plan around a number you have not confirmed.
- Customers beyond your key limit: use one key (or a few) for the whole product and do the per-customer work in your own app, as described below.
Note
More keys do not mean more throughput. The per-minute limit and the concurrency limit are counted per account, so ten keys share the same limits as one. If one customer can send a burst, put your own limit in front of Tokens, for example a queue or a semaphore per customer.
Create keys for environments and services#
A good minimum set:
| Key | Cap | Allowed models |
|---|---|---|
prod | The most you accept to lose to a bug in one month | The models your product uses |
staging | A small amount | One cheap model |
ci or batch | The cost of the job, with margin | The one model the job needs |
Set the cap and the model list when you create the key, because you cannot change them afterwards. If a cap turns out to be wrong, create a new key with the right values, switch your app to it, then revoke the old one. Keep the old key's secret out of any place you have already moved on from.
What happens at the cap#
When a request would take the key over its monthly cap, the gateway refuses it before it reaches a model:
{
"error": {
"message": "Monthly spend cap of $20.00 for this API key has been reached. Update the key cap or use an alternate key: https://tokens.bd/dashboard/billing",
"type": "permission_denied_error",
"code": "monthly_spend_cap_exceeded",
"param": null,
"request_id": "..."
}
}The status is 403, not 429, and there is no Retry-After header: waiting a few seconds does not help, and the key stays blocked until the next calendar month. Details to design around:
- The check adds up what the key has already spent this month and the worst-case cost of the new request, which depends on its
max_tokens(8,192 output tokens if you do not set it). A request with a largemax_tokenscan be refused slightly before the cap is actually reached, while a smaller one still passes. - Spend is counted when a request finishes. Requests in flight at the moment you reach the cap are not counted yet, so a key can end a little over its cap.
- Only that key stops. Your other keys, your plan and your wallet are not affected, unless the account itself is out of credit (see plans, credits and wallet).
- Rotating a key does not reset its spend. The key keeps the same identity, so the month's usage still counts against the cap.
In your app, treat monthly_spend_cap_exceeded as "this customer or environment has used its allowance", not as an outage. Show your own message, and do not retry. See errors for the full list of codes.
Track spend per customer#
Tokens does not show spend per key. The Usage page, the CSV export and GET /v1/tokens/usage report your account as a whole, and the CSV has no key column. So the record of which customer used what has to be kept by your app.
For every call you make on a customer's behalf, store:
- your customer id,
- the
x-tokens-request-idresponse header, - the model you called, and
- the
usageobject from the response (prompt_tokens,completion_tokensfor Chat Completions).
You can then multiply the tokens by the model's prices from the catalog to get a cost per customer. Your totals will not match the bill to the last cent if a plan discount or allowance applies, so use the Request ID column in the exported CSV (usage and alerts) to look up what a given request actually cost.
For streaming responses, add "stream_options": {"include_usage": true} to the request so the last chunk carries the usage. See streaming.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://tokens.bd/v1",
api_key=os.environ["TOKENS_API_KEY"],
)
def ask_for_customer(customer_id: str, prompt: str) -> str:
raw = client.chat.completions.with_raw_response.create(
model="deepseek/deepseek-v4.1-flash",
messages=[{"role": "user", "content": prompt}],
max_tokens=500,
)
completion = raw.parse()
save_usage_row( # your own database write
customer_id=customer_id,
request_id=raw.headers.get("x-tokens-request-id"),
model=completion.model,
prompt_tokens=completion.usage.prompt_tokens,
completion_tokens=completion.usage.completion_tokens,
)
return completion.choices[0].message.contentIf you use a key per customer, the key's cap protects you; the table above is still how you bill them.
Team accounts#
Tokens has team accounts (organizations) with roles, and they are switched on per deployment, so you may not see them yet in your dashboard. Where they are on, keys belong to the organization rather than to one person. Based on the code, the roles are owner, admin, developer, billing and viewer; the billing and viewer roles cannot create keys, a developer sees the keys they created, and owners and admins can rotate or revoke any key in the organization. An organization can also give each member a monthly spend cap, which returns 402 member_cap_reached when it is reached. Read teams and roles for what your account has.
Rotate and revoke#
- Rotate replaces the secret and keeps the name, cap, allowed models and usage history. The old secret stops working immediately; there is no grace period. Requests that use it get
401 invalid_api_key. - Revoke disables the key permanently. Requests with it get
403 key_inactive.
To rotate without downtime, create a second key with the same settings first, deploy it to your app, then revoke the old one. Because key settings are fixed, "same settings" means you enter them again. Rotate or revoke at once if a key leaks, then look at usage for requests you do not recognise.
What your own app must do#
- Map your user to a key. Keep a table from customer or environment to the key. The secret is shown once, so store it in a secrets manager or an encrypted column, never in plain text and never in logs.
- Keep keys on your server. Tokens does not answer browser calls (no CORS headers), and anything in a browser or mobile app can be extracted by users. Your clients call your backend, and your backend calls Tokens. See browser and mobile apps.
- Limit abuse yourself. Per-customer rate limits, request-size limits and a maximum
max_tokensbelong in your code, because Tokens' rate limits are per account. - Handle the key errors. Map
monthly_spend_cap_exceeded,model_not_allowed_on_key,key_inactiveandinvalid_api_keyto a clear state in your product, and alert yourself onkey_inactiveandinvalid_api_key, which mean your own configuration is wrong. - Watch the account. Every key spends from the same plan and wallet. Turn on the low-balance and usage alerts in the dashboard (usage and alerts), and check production checklist before launch.
- Read the terms. What you may build and resell on top of your account is set by the terms of service, not by this page.
Checklist#
- Count the keys you need and compare with your plan's active-key limit before you design around them.
- Create each key with a cap and an allowed-models list, because you cannot add them later.
- Store every secret on the server, and record your customer id with each request.
- Test the cap path: create a key with a tiny cap and confirm your app handles
403 monthly_spend_cap_exceeded. - Write down how you rotate a key, and try it once before you need it.