# One key per customer or environment

> Building a product on Tokens: what a key can be limited by, how keys are created and revoked (dashboard only, no key-management API), how many you can have, how to track spend per customer, and what your own app must do.

If you build a product on Tokens, you will want to know which part of your spend belongs to which customer, and to stop one customer or one broken service from using up everything. Separate API keys are the tool Tokens gives you for this. This page says what a key can and cannot do, so you can decide where to use keys and where to do the work in your own app.

Read [API keys](/docs/api-keys) first for the basics. This page builds on it.

## What Tokens can and cannot do here

| You might expect                         | What exists                                                                                                                   |
| ---------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------- |
| Create a key with a spend cap            | Yes, in the dashboard at [/dashboard/keys](/dashboard/keys)                                                                    |
| Restrict a key to some models            | Yes, an allowed-models list set at creation                                                                                   |
| Create, rotate or revoke a key by API    | **No.** There is no key-management endpoint you can call with a `tok_live_` key                                               |
| Set an expiry date on a key              | **No.** The creation form has no expiry field. `key_expired` applies only to keys an administrator provisioned with an expiry |
| Edit a key's cap or models later         | **No.** Both are fixed at creation                                                                                            |
| A separate rate limit for each key       | **No.** Requests per minute and concurrency are per account, shared by all keys                                               |
| Spend per key in the dashboard or an API | **No.** See [Track spend per customer](#track-spend-per-customer)                                                             |
| Hundreds of keys on one account          | Only if your plan allows it. The default is 3 active keys                                                                     |

The dashboard talks to routes under `/api/keys`. They are authenticated by your signed-in browser session (and your multi-factor check, if you use one), not by a Tokens API key, and they are not a supported public API. Do not build a provisioning service on them. The [Tokens CLI](/docs/tokens-cli) signs in through the browser and receives one new key per login; it is for setting up your own coding agents, not for issuing keys to customers.

## What a key can be limited by

When you create a key in the dashboard you can set:

| Setting                 | Effect                                                                                                                        |
| ----------------------- | ----------------------------------------------------------------------------------------------------------------------------- |
| Name                    | Up to 64 characters. Use it to record whose key it is: `acme-prod`, `staging`, `batch-worker`.                                |
| Monthly spend cap (USD) | The most the key may spend in a calendar month. Left blank, the key takes your plan's default cap if the plan defines one.    |
| Allowed models          | The only model ids the key may call. Empty means every model your account can use. `GET /v1/models` lists only allowed models. |

A key also stops working when it is revoked or its account is suspended. Rate limits, concurrency and usage windows are not key settings; they belong to your account, see [rate limits](/docs/rate-limits).

To read a key's own settings from code, call `GET /v1/tokens/usage` with that key. The `key` object in the response shows its `monthlySpendCapUsd` and `allowedModels`. The `plan`, `windows` and `wallet` parts of the same response describe your whole account, not the key.

```bash
curl -s https://tokens.bd/v1/tokens/usage \
  -H "Authorization: Bearer $CUSTOMER_KEY" | jq .key
```

## How many keys you can have

Each plan sets a maximum number of active keys, and the default is 3 (pay-as-you-go accounts also default to 3). Creating one more returns `key_limit_reached`. Revoked keys do not count.

This decides your design:

- **A handful of services or environments** (production, staging, a batch worker, an internal tool): one key each fits inside the default.
- **Many customers**: a key per customer works only if your plan's limit is high enough. Check your plan on [pricing](/pricing), or ask [support](/docs/support) what limit is possible for you. Do not plan around a number you have not confirmed.
- **Customers beyond your key limit**: use one key (or a few) for the whole product and do the per-customer work in your own app, as described below.

:::note
More keys do not mean more throughput. The per-minute limit and the concurrency limit are counted per account, so ten keys share the same limits as one. If one customer can send a burst, put your own limit in front of Tokens, for example a queue or a semaphore per customer.
:::

## Create keys for environments and services

A good minimum set:

| Key              | Cap                              | Allowed models                       |
| ---------------- | -------------------------------- | ------------------------------------ |
| `prod`           | The most you accept to lose to a bug in one month | The models your product uses |
| `staging`        | A small amount                   | One cheap model                      |
| `ci` or `batch`  | The cost of the job, with margin | The one model the job needs          |

Set the cap and the model list when you create the key, because you cannot change them afterwards. If a cap turns out to be wrong, create a new key with the right values, switch your app to it, then revoke the old one. Keep the old key's secret out of any place you have already moved on from.

## What happens at the cap

When a request would take the key over its monthly cap, the gateway refuses it before it reaches a model:

```json
{
  "error": {
    "message": "Monthly spend cap of $20.00 for this API key has been reached. Update the key cap or use an alternate key: https://tokens.bd/dashboard/billing",
    "type": "permission_denied_error",
    "code": "monthly_spend_cap_exceeded",
    "param": null,
    "request_id": "..."
  }
}
```

The status is `403`, not `429`, and there is no `Retry-After` header: waiting a few seconds does not help, and the key stays blocked until the next calendar month. Details to design around:

- The check adds up what the key has already spent this month and the worst-case cost of the new request, which depends on its `max_tokens` (8,192 output tokens if you do not set it). A request with a large `max_tokens` can be refused slightly before the cap is actually reached, while a smaller one still passes.
- Spend is counted when a request finishes. Requests in flight at the moment you reach the cap are not counted yet, so a key can end a little over its cap.
- Only that key stops. Your other keys, your plan and your wallet are not affected, unless the account itself is out of credit (see [plans, credits and wallet](/docs/plans-and-wallet)).
- Rotating a key does not reset its spend. The key keeps the same identity, so the month's usage still counts against the cap.

In your app, treat `monthly_spend_cap_exceeded` as "this customer or environment has used its allowance", not as an outage. Show your own message, and do not retry. See [errors](/docs/errors) for the full list of codes.

## Track spend per customer

Tokens does not show spend per key. The Usage page, the CSV export and `GET /v1/tokens/usage` report your account as a whole, and the CSV has no key column. So the record of which customer used what has to be kept by your app.

For every call you make on a customer's behalf, store:

1. your customer id,
2. the `x-tokens-request-id` response header,
3. the model you called, and
4. the `usage` object from the response (`prompt_tokens`, `completion_tokens` for Chat Completions).

You can then multiply the tokens by the model's prices from [the catalog](/models) to get a cost per customer. Your totals will not match the bill to the last cent if a plan discount or allowance applies, so use the **Request ID** column in the exported CSV ([usage and alerts](/docs/usage-and-alerts)) to look up what a given request actually cost.

For streaming responses, add `"stream_options": {"include_usage": true}` to the request so the last chunk carries the usage. See [streaming](/docs/streaming).

```python title="record_usage.py"
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
)

def ask_for_customer(customer_id: str, prompt: str) -> str:
    raw = client.chat.completions.with_raw_response.create(
        model="deepseek/deepseek-v4.1-flash",
        messages=[{"role": "user", "content": prompt}],
        max_tokens=500,
    )
    completion = raw.parse()
    save_usage_row(  # your own database write
        customer_id=customer_id,
        request_id=raw.headers.get("x-tokens-request-id"),
        model=completion.model,
        prompt_tokens=completion.usage.prompt_tokens,
        completion_tokens=completion.usage.completion_tokens,
    )
    return completion.choices[0].message.content
```

If you use a key per customer, the key's cap protects you; the table above is still how you bill them.

## Team accounts

Tokens has team accounts (organizations) with roles, and they are switched on per deployment, so you may not see them yet in your dashboard. Where they are on, keys belong to the organization rather than to one person. Based on the code, the roles are owner, admin, developer, billing and viewer; the billing and viewer roles cannot create keys, a developer sees the keys they created, and owners and admins can rotate or revoke any key in the organization. An organization can also give each member a monthly spend cap, which returns `402 member_cap_reached` when it is reached. Read [teams and roles](/docs/teams-and-roles) for what your account has.

## Rotate and revoke

- **Rotate** replaces the secret and keeps the name, cap, allowed models and usage history. The old secret stops working immediately; there is no grace period. Requests that use it get `401 invalid_api_key`.
- **Revoke** disables the key permanently. Requests with it get `403 key_inactive`.

To rotate without downtime, create a second key with the same settings first, deploy it to your app, then revoke the old one. Because key settings are fixed, "same settings" means you enter them again. Rotate or revoke at once if a key leaks, then look at [usage](/dashboard/usage) for requests you do not recognise.

## What your own app must do

- **Map your user to a key.** Keep a table from customer or environment to the key. The secret is shown once, so store it in a secrets manager or an encrypted column, never in plain text and never in logs.
- **Keep keys on your server.** Tokens does not answer browser calls (no CORS headers), and anything in a browser or mobile app can be extracted by users. Your clients call your backend, and your backend calls Tokens. See [browser and mobile apps](/docs/browser-and-mobile).
- **Limit abuse yourself.** Per-customer rate limits, request-size limits and a maximum `max_tokens` belong in your code, because Tokens' rate limits are per account.
- **Handle the key errors.** Map `monthly_spend_cap_exceeded`, `model_not_allowed_on_key`, `key_inactive` and `invalid_api_key` to a clear state in your product, and alert yourself on `key_inactive` and `invalid_api_key`, which mean your own configuration is wrong.
- **Watch the account.** Every key spends from the same plan and wallet. Turn on the low-balance and usage alerts in the dashboard ([usage and alerts](/docs/usage-and-alerts)), and check [production checklist](/docs/production-checklist) before launch.
- **Read the terms.** What you may build and resell on top of your account is set by the [terms of service](/terms), not by this page.

## Checklist

1. Count the keys you need and compare with your plan's active-key limit before you design around them.
2. Create each key with a cap and an allowed-models list, because you cannot add them later.
3. Store every secret on the server, and record your customer id with each request.
4. Test the cap path: create a key with a tiny cap and confirm your app handles `403 monthly_spend_cap_exceeded`.
5. Write down how you rotate a key, and try it once before you need it.

---
Page: https://tokens.bd/docs/one-key-per-customer
