# Production checklist

> What to set up before real users depend on the Tokens API: retries and backoff, timeouts, Retry-After, request ids, a key per environment, spend caps, alerts, key rotation, and what to do on 402 and 429.

A script that works on your laptop fails in production in predictable ways: a burst of traffic hits a rate limit, a long answer outlives a timeout, the balance runs out at night, a key leaks. Each one has a fix that takes minutes to set up now and hours to set up during an outage. This page is the list, in the order that matters, with the reasons.

The details of each error are in [errors](/docs/errors) and [rate limits](/docs/rate-limits). This page says what to do about them in a live system.

## The checklist

| Done | Item                                                                                          |
| ---- | --------------------------------------------------------------------------------------------- |
| [ ]  | The API key is in an environment variable or secret store, never in code or client bundles    |
| [ ]  | One key per environment (development, staging, production), each with a spend cap            |
| [ ]  | Production keys restricted to the models the service uses                                     |
| [ ]  | `max_tokens` is set on every request                                                          |
| [ ]  | Client timeouts are set explicitly, and longer for long or reasoning requests                 |
| [ ]  | Retries with exponential backoff and jitter, only on the errors that are worth retrying       |
| [ ]  | `Retry-After` is honored                                                                      |
| [ ]  | 402 and `window_exhausted` are never retried in a loop                                        |
| [ ]  | Parallel requests are capped below your plan's concurrency limit                              |
| [ ]  | `x-tokens-request-id` is logged for every request, successful or not                          |
| [ ]  | Usage alerts are on, and someone reads them                                                   |
| [ ]  | A plan for key rotation exists and has been tried once                                        |
| [ ]  | The product has a defined behavior when the model is unavailable                              |

## Keys: one per environment, capped

Create a separate key for each environment and each service. When something goes wrong, the key name in the [usage page](/dashboard/usage) and the key list tells you where the spend came from, and you can revoke one key without taking the others down.

Set a **monthly spend cap** when you create the key. A cap turns a runaway loop, a retry storm or a leaked key into an error instead of a bill. It cannot be edited afterwards: to change it, create a new key. Also set the **allowed models** list, so a bug cannot call a model more expensive than the one you tested with.

Keys per account are limited by plan (the default is 3 active keys). If you want development, staging and production keys plus one for a CI job, check your plan's limit on [pricing](/pricing) before you plan the layout. Local development can share the key your tool already uses, as long as it has a cap.

Keep keys out of source control, out of container images and out of client-side code. If your users run in a browser or a phone, see [browser and mobile](/docs/browser-and-mobile). The full guide is in [API keys](/docs/api-keys).

## Set max_tokens on every request

Before a request runs, the gateway reserves its worst-case cost against your plan or wallet. The output part of that reservation uses your `max_tokens` (or `max_completion_tokens`), and 8,192 when you leave it unset. Two things follow:

- A large or missing `max_tokens` can get a request refused near a key's spend cap or when the balance is low, even though the real answer would be short.
- When your balance covers only part of the reservation, the gateway lowers `max_tokens` to what the balance affords, down to a floor of 16. The answer then stops early with `finish_reason: "length"`. If you see truncated answers when money is low, this is why.

Pick a limit that fits the longest answer you want, not the largest the model allows.

## Retries

Retry the errors that can go away on their own, and nothing else.

| Retry?                            | Errors                                                                                | How                                                                     |
| --------------------------------- | ------------------------------------------------------------------------------------- | ----------------------------------------------------------------------- |
| Yes, after `Retry-After`          | 429 `rate_limited`, `concurrency_limit`, `rate_limit_exceeded`                        | Wait the number of seconds in the header, plus a little random jitter   |
| Yes, with exponential backoff     | 500, 502, 503, 504, and connection errors or timeouts on your side                    | Start near 1 second, double each time, cap at 30 seconds, 4 or 5 tries  |
| Only after the reset              | 429 `window_exhausted`, `model_limit_reached`                                          | Not in a loop. `Retry-After` can be hours or days. Pause the job        |
| No, fix something first           | 400, 401, 402, 403, 404, 413                                                          | These repeat forever. Alert a person                                    |

Three details that decide whether retries help or hurt:

- **A retry is a new request.** There is no idempotency key. If the first request actually completed on the gateway while your client timed out, you were billed for it, and the retry is billed again. Keep the number of attempts low, and be careful with long prompts.
- **Check the SDK first.** The OpenAI and Anthropic SDKs retry some errors by themselves (`max_retries`). Stacking your own loop on top multiplies the attempts. Either turn the SDK's retries off, as the example in [rate limits](/docs/rate-limits) does, or rely on them and add nothing.
- **Do not retry a stream that has already started and been used.** If an answer broke halfway, a retry sends the whole prompt again and bills it again. Decide whether the partial text is good enough. Retry only streams that failed before any content arrived. See [streaming](/docs/streaming).

The gateway already fails over between providers for models that have more than one, before it returns an error to you. A 5xx you see means the gateway already tried.

### A request wrapper

This wrapper calls chat completions with a timeout, a fresh request id per attempt, backoff that honors `Retry-After`, no retry on the codes that need a person, and a log line for every failure. It uses the global `fetch` of Node.js 18 or later.

```typescript title="tokens-client.ts"
const BASE_URL = "https://tokens.bd/v1";
const RETRY_STATUS = new Set([429, 500, 502, 503, 504]);
const WAIT_FOR_RESET = new Set(["window_exhausted", "model_limit_reached"]);

export type TokensResult =
  | { ok: true; data: unknown; requestId: string | null }
  | { ok: false; status: number; code: string; requestId: string | null };

const sleep = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms));

export async function chat(
  body: Record<string, unknown>,
  maxAttempts = 4
): Promise<TokensResult> {
  let delay = 1000;
  for (let attempt = 1; ; attempt++) {
    const clientRequestId = crypto.randomUUID();
    let status = 0; // 0 means no response: timeout or connection error
    let code = "network_error";
    let requestId: string | null = null;
    let wait = delay;

    try {
      const res = await fetch(`${BASE_URL}/chat/completions`, {
        method: "POST",
        headers: {
          Authorization: `Bearer ${process.env.TOKENS_API_KEY}`,
          "Content-Type": "application/json",
          "x-request-id": clientRequestId,
        },
        body: JSON.stringify(body),
        signal: AbortSignal.timeout(180_000),
      });
      requestId = res.headers.get("x-tokens-request-id");
      if (res.ok) return { ok: true, data: await res.json(), requestId };

      status = res.status;
      const err = (await res.json().catch(() => null)) as { error?: { code?: string } } | null;
      code = err?.error?.code ?? "unknown";
      const retryAfter = Number(res.headers.get("retry-after"));
      if (retryAfter > 0) wait = retryAfter * 1000;
    } catch {
      // Timeout or connection error: handled below like a 5xx.
    }

    console.warn(
      JSON.stringify({ event: "tokens_call_failed", attempt, status, code, requestId, clientRequestId })
    );
    const retryable = status === 0 || RETRY_STATUS.has(status);
    if (!retryable || WAIT_FOR_RESET.has(code) || attempt >= maxAttempts) {
      return { ok: false, status, code, requestId };
    }
    await sleep(Math.min(wait, 60_000) + Math.random() * 500);
    delay = Math.min(delay * 2, 30_000);
  }
}
```

The 180-second timeout suits a non-streaming call with a moderate `max_tokens`. See the next section for long requests.

## Timeouts

Set timeouts on purpose. Library defaults are rarely right for model calls.

- **The gateway is patient.** For a model with a single provider it waits up to 600 seconds for the first byte, and up to 300 seconds of silence between chunks once a stream is flowing. A reasoning model can think for minutes before it says anything. If the provider does not start in time you get 504 `upstream_timeout`.
- **Your client must be at least as patient** as the request needs. Node.js's built-in `fetch` gives up after 300 seconds without response headers and after 300 seconds between body chunks by default, according to the documentation of undici, the HTTP client behind it. A request that needs longer needs a longer limit.
- **Stream anything that can run long.** A non-streaming request sends nothing until the whole answer exists, so a proxy or load balancer in front of your app can cut it as idle. A stream sends data as it goes. Use `stream: true` for long answers and long agent tasks.
- **Mind your own infrastructure.** Reverse proxies, serverless platforms and mobile networks often have idle or total-time limits of 30 to 120 seconds. Check each hop between your user and Tokens.
- **Cancel when the user leaves.** Pass an abort signal. The gateway stops the upstream request and bills only what was generated.
- **Do not set one timeout for everything.** A short summary and an agent turn with tools deserve different limits.

When a model has a backup provider, the gateway moves to it after about 30 seconds without a first token on a stream (120 seconds on a non-streaming request). You never see the switch. See [streaming](/docs/streaming).

## Retry-After

`Retry-After` is the only rate-limit header Tokens sends. It is in seconds, on 429 responses. There are no `X-RateLimit-*` headers. Its meaning depends on the code:

| Code                  | `Retry-After`                                  | Retry?                                  |
| --------------------- | ---------------------------------------------- | --------------------------------------- |
| `rate_limited`        | Seconds until the next minute starts (up to 60) | Yes                                     |
| `concurrency_limit`   | 2                                              | Yes, after reducing parallelism         |
| `rate_limit_exceeded` | Passed through from the provider, if present   | Yes, with backoff                       |
| `window_exhausted`    | Seconds until the window resets, often hours   | No. Pause or tell the user              |
| `model_limit_reached` | Seconds until the billing period resets        | No. Use another model or wait           |

Wait at least that long, and add a little random jitter so many clients do not wake together. Always read the `code` first, then the header.

## Stay under the concurrency limit

Requests per minute and requests in flight are limits on your whole account, shared by all keys. A streaming request holds a slot until it finishes. If your service fans out, cap parallelism with a semaphore or queue sized a little below your plan's concurrency limit, instead of sending everything and retrying the 429s. Creating more keys does not raise the limits. Abandoned streams hold a slot until they end, so close them.

## Request ids

Log `x-tokens-request-id` for every call, on success and on failure. It is the one value support needs to find your request. Send your own `x-request-id` as well, so a line in your logs can be matched to Tokens' id. If a call times out on your side, you have no response and no Tokens id, so log your own id before the call. The full guide is in [request ids and debugging](/docs/request-ids-and-debugging).

## Alerts and spend

- **Turn on usage alerts** under Notifications in the dashboard: warnings at 50, 75 and 90 percent of a plan limit, and a low-balance email when the wallet drops under $5. Each is sent once per threshold. See [usage and alerts](/docs/usage-and-alerts).
- **Turn on the renewal reminder.** Plans do not renew by themselves. A plan that quietly expired looks exactly like a 402.
- **Check before long jobs.** `GET /v1/tokens/usage` returns your windows, your wallet balance and the key's cap. It is not billed and does not count toward the per-minute limit, so a batch job can check it before it starts and between steps.
- **Watch the usage page after each release.** A change in prompt size, a retry bug or a new agent shows up as a jump in tokens or cost per request. See [Usage](/dashboard/usage).

```bash
curl -s https://tokens.bd/v1/tokens/usage -H "Authorization: Bearer $TOKENS_API_KEY" \
  | jq '{wallet: .wallet.balanceUsd, windows: [.windows[] | {type, percentUsed, resetsAt}]}'
```

## Handle 402 and 429

**402 (`insufficient_credits`, `no_funding`, `outstanding_debt`, `member_cap_reached`).** Your plan or wallet cannot pay for the request. Retrying cannot fix it, and each rejected request still counts against your per-minute limit. In your code:

1. Stop sending to that key. Open a circuit so the rest of your service does not keep trying.
2. Alert the person who can top up, with the request id.
3. Tell your own users something short, such as "the service is temporarily unavailable". Do not show them a billing message.
4. After a top-up in [billing](/dashboard/billing), close the circuit.

**429.** Branch on `code`:

- `rate_limited`, `concurrency_limit`, `rate_limit_exceeded`: wait `Retry-After`, then retry. If it happens often, you are over capacity: queue requests or cap parallelism.
- `window_exhausted`: your plan's 5-hour, weekly or monthly usage is used up. Do not loop. Show the reset time, pause the job, or move to a funded wallet if you use one.
- `model_limit_reached`: this one model's allowance is used up for the period. Other models still work. Switch to another model or wait.

A `monthly_spend_cap_exceeded` (403) is the key's own cap. Waiting seconds does not help. Use another key, or wait for the next month.

## Graceful degradation

Decide now what the product does when Tokens, or one model, is not available.

- **Have a second model.** Choose one from `GET /v1/models` that is enough for your task, and fall back to it on 5xx after your retries, on `model_not_available` and on `model_limit_reached`. Test it, because models differ in tool calling and context size. See [choosing a model](/docs/choosing-a-model).
- **Do not fall back on errors another model will not fix:** 401, 402, 403 and most 400s apply to the request, the key or the account.
- **Fail clearly.** Show the user an honest message and a way to try again. Queue background work and run it later.
- **Cut the load when constrained.** Shorter prompts, a smaller `max_tokens`, or turning off optional features such as automatic summaries.
- **Check the status page.** If many things fail at once, look at [status](/status) before you look at your code.

This is the wrapper above with a fallback model chosen from the environment:

```typescript title="answer.ts"
import { chat } from "./tokens-client";

type Messages = { role: "system" | "user" | "assistant"; content: string }[];

export async function answer(messages: Messages) {
  const primary = await chat({ model: "deepseek/deepseek-v4.1-flash", messages, max_tokens: 800 });
  if (primary.ok) return primary;

  const fallbackModel = process.env.FALLBACK_MODEL;
  const worthFallingBack =
    primary.status === 0 ||
    primary.status >= 500 ||
    primary.code === "model_not_available" ||
    primary.code === "model_limit_reached";
  if (!fallbackModel || !worthFallingBack) return primary;

  return chat({ model: fallbackModel, messages, max_tokens: 800 });
}
```

## Rotate keys

Plan rotation before you need it, and try it once while nothing is on fire.

- **Rotate on a schedule and after any leak or staff change.** A leaked key is revoked or rotated immediately.
- **Rotating a key replaces its secret at once.** The old secret stops working with no grace period, and every client still using it gets 401 `invalid_api_key`. The key keeps its name, cap, allowed models and history.
- **For no downtime, create a second key first.** Deploy it, check that traffic uses it on the [usage page](/dashboard/usage), then revoke the old key. This needs a free key slot on your plan.
- **Keep the secret in one place** (a secret manager or your platform's environment settings) so that a change is one update plus a restart, not a search through repositories.
- **Test the failure.** In staging, revoke the staging key and confirm that your service alerts and recovers once you set the new one.

## Before launch

Run through the table at the top. Then send a few test requests with a deliberately wrong key, a model your key is not allowed to use, and a prompt that is too large, and check that your service logs the request id, does not retry, and shows your user a sensible message. Real traffic will find these cases whether or not you do.

---
Page: https://tokens.bd/docs/production-checklist
