An error from an AI gateway can come from three places: your request, your account, or a model provider somewhere behind the gateway. The HTTP status narrows it down, and the code field pins it. This guide goes through the AI gateway errors you'll see from Tokens, status by status, with what each one means and whether retrying makes any sense.
401 your key fix the key, don't retry
402 your balance top up, don't retry
403 your key's limits new key or different model, don't retry
404 your URL or model fix the request, don't retry
429 slow down wait for Retry-After, then retry
5xx upstream trouble back off and retry, a few timesThe shape of an error#
Every error the gateway itself produces has the same JSON body, in OpenAI's style:
{
"error": {
"message": "Model 'deepseek/deepseek-v4.1-flash' is not permitted on this API key. Permitted models: ...",
"type": "permission_denied_error",
"code": "model_not_allowed_on_key",
"param": null,
"request_id": "6f1c2a9e-..."
}
}Branch on code, not on message. Messages are written for people and can change; codes are the contract. Every response, success or failure, also carries an x-tokens-request-id header. Keep it; it's what support needs.
One exception to know about if you use Claude Code or the Anthropic SDK. On /v1/messages, errors from the gateway itself (authentication, caps, balance) use the shape above, but errors passed back from an upstream provider use Anthropic's shape: {"type": "error", "error": {"type": "...", "message": "..."}}, with no code field. For those, the status and the x-tokens-request-id header are what you have.
401: the key wasn't accepted#
| Code | Meaning |
|---|---|
missing_api_key | No key in the request. Send Authorization: Bearer <key> or x-api-key: <key>. |
invalid_api_key | A key was sent, but it doesn't match an active key. |
The usual causes are mundane: an environment variable that isn't set in the shell the agent runs in, a key with a trailing newline from copy-paste, or a key that was rotated. Rotation replaces the secret immediately, so an agent still holding the old one gets 401 on its next request. Keys look like tok_live_ followed by 48 hex characters; if yours doesn't, it's been truncated somewhere.
Don't retry a 401. It won't fix itself.
402: the account can't pay for this request#
| Code | Meaning |
|---|---|
insufficient_credits | Your wallet balance doesn't cover the request |
no_funding | No active subscription and no funded wallet |
outstanding_debt | Earlier usage left a negative balance that has to be cleared first |
member_cap_reached | On a team account, your member monthly cap is used up; your admin raises it |
Every 402 message links to billing. Note that a low balance doesn't always mean a 402: when the balance covers part of a request, the gateway may lower max_tokens to what it can pay for, down to 16 tokens. If answers suddenly come back truncated, check the balance.
403: the key isn't allowed to do this#
| Code | Meaning | Fix |
|---|---|---|
model_not_allowed_on_key | The key has an allowed-models list and this model isn't on it. The message lists the allowed ones. | Use an allowed model, or a different key |
monthly_spend_cap_exceeded | The key reached its monthly spend cap | A new key with a higher cap, or wait for next month |
tier_permission_denied | Your plan doesn't include this model and you have no wallet balance to pay for it as you go | Add funds or change plan |
key_inactive | The key was revoked or suspended | Create a new key |
key_expired | The key had an expiry date that has passed | Create a new key |
account_suspended | The account is suspended | Contact support |
Caps and allow-lists can't be edited on an existing key, so the fix for the first two is almost always a new key. Spend caps for coding agents explains how to plan keys so this happens on purpose rather than by accident.
There's one 403 that isn't about you: if the code is upstream_auth_error, the gateway's own credentials for a provider were rejected. The same code can come with a 401 status. Nothing in your setup will fix it; open a ticket with the request ID.
400, 404 and 413: the request itself#
| Status and code | Meaning |
|---|---|
404 unsupported_endpoint | The path isn't one the gateway serves |
404 model_not_found | The model ID doesn't match an active model |
400 model_not_available | The model exists but isn't available to you right now |
400 invalid_request | The upstream provider rejected the request body |
| 413 | The body is over 10 MB |
unsupported_endpoint almost always means a base URL mistake. OpenAI-style clients need https://tokens.bd/v1; Anthropic-style clients, including Claude Code, need https://tokens.bd, because they add /v1/messages themselves. Mix them up and you get /v1/v1/... or a missing /v1. The other cause is calling an endpoint the gateway doesn't offer: images, audio, files, batches, assistants, fine-tuning and moderations all return this code.
For model_not_found, compare your ID against GET /v1/models character by character. IDs are provider/model aliases, like deepseek/deepseek-v4.1-flash. Agents that reference models as <provider-id>/<model-id> (OpenCode's tokens/deepseek/..., for example) should send only the part after their own provider ID; if the error message shows a doubled prefix, the agent config is the place to look.
invalid_request is the vaguest one, because the gateway returns a generic message rather than the provider's original text. Common causes are parameters the model doesn't accept (a temperature outside its range, a tool format it doesn't support) and conversations that exceed the context window. Try the same request with a minimal body; if that works, add fields back until it breaks.
429: too much, too fast#
| Code | Meaning | Retry-After |
|---|---|---|
rate_limited | Over your per-minute request limit (60 by default) | Seconds until you can retry |
concurrency_limit | Too many requests in flight at once for your plan | 2 seconds |
window_exhausted | A plan usage window (5-hour, weekly or monthly) is used up | Seconds until the window resets, which can be hours |
rate_limit_exceeded | An upstream provider rate-limited the request, on every source the gateway tried | Passed through when the provider sends one |
There are no X-RateLimit-* headers to watch; Retry-After on the 429 is the signal. The first two clear in seconds and are safe to retry after the wait. window_exhausted is different in kind: retrying in a loop just wastes requests until the reset time, so surface it to a human. And concurrency_limit from a coding agent often means too many parallel subagents or tool calls; lowering the agent's parallelism fixes it more reliably than retrying.
5xx: something upstream failed#
| Status and code | Meaning |
|---|---|
502 upstream_unreachable | No provider source could be reached |
503 no_upstream_available | No source is currently configured or healthy for this model |
504 upstream_timeout | The provider stopped responding mid-request |
Other 5xx, upstream_error | The provider returned a server error |
Before any of these reach you, the gateway has already tried: on a 429, 502, 503, 504 or connection error from one source, it fails over to the next source for that model automatically. So a 5xx means every option failed for that request. A couple of retries with backoff is reasonable; a tight loop isn't. If it persists, check the status page, which shows the inference API and upstream providers separately.
Long requests by themselves aren't a problem: the gateway waits up to 600 seconds for response headers. If you're seeing timeouts, check your own client's timeout first, since many SDKs default to something shorter.
A retry function that honours Retry-After#
The rules: retry 429 and 5xx and network failures; never retry other 4xx; use Retry-After when present; otherwise back off exponentially with jitter; give up after a few attempts; and treat a very long Retry-After as "stop" rather than sleeping for an hour.
const RETRYABLE = new Set([429, 500, 502, 503, 504]);
export async function callWithRetry(url, body, { attempts = 4, maxWaitS = 60 } = {}) {
for (let i = 0; i < attempts; i++) {
let res;
try {
res = await fetch(url, {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.TOKENS_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify(body),
});
} catch (err) {
if (i === attempts - 1) throw err; // network error: retry, then give up
await sleep(backoff(i));
continue;
}
if (res.ok) return res.json();
const requestId = res.headers.get("x-tokens-request-id");
const err = await res.json().catch(() => ({}));
const code = err.error?.code ?? `http_${res.status}`;
if (!RETRYABLE.has(res.status) || i === attempts - 1) {
throw new Error(`${res.status} ${code} (request ${requestId}): ${err.error?.message ?? ""}`);
}
const retryAfter = Number(res.headers.get("retry-after"));
const waitS = Number.isFinite(retryAfter) && retryAfter > 0 ? retryAfter : backoff(i) / 1000;
if (waitS > maxWaitS) {
throw new Error(
`${res.status} ${code}: retry in ${waitS}s is too long (request ${requestId})`
);
}
await sleep(waitS * 1000);
}
}
const backoff = (i) => Math.min(30_000, 1000 * 2 ** i) * (0.5 + Math.random() / 2);
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));import os, random, time
import requests
RETRYABLE = {429, 500, 502, 503, 504}
def backoff(i: int) -> float:
return min(30.0, 2 ** i) * (0.5 + random.random() / 2)
def call_with_retry(url: str, body: dict, attempts: int = 4, max_wait_s: float = 60):
headers = {"Authorization": f"Bearer {os.environ['TOKENS_API_KEY']}"}
for i in range(attempts):
try:
res = requests.post(url, json=body, headers=headers, timeout=600)
except requests.ConnectionError:
if i == attempts - 1:
raise
time.sleep(backoff(i))
continue
if res.ok:
return res.json()
request_id = res.headers.get("x-tokens-request-id")
try:
err = res.json().get("error", {})
except ValueError:
err = {}
code = err.get("code", f"http_{res.status_code}")
if res.status_code not in RETRYABLE or i == attempts - 1:
raise RuntimeError(f"{res.status_code} {code} (request {request_id}): {err.get('message', '')}")
retry_after = res.headers.get("retry-after")
wait_s = float(retry_after) if retry_after and retry_after.isdigit() else backoff(i)
if wait_s > max_wait_s:
raise RuntimeError(f"{res.status_code} {code}: retry in {wait_s:.0f}s is too long (request {request_id})")
time.sleep(wait_s)The maxWaitS cutoff is what keeps a window_exhausted from turning into an hour-long sleep. Two caveats for streaming. Only retry if the stream failed before any content arrived; once tokens have streamed, they've been generated and metered, so a retry pays for the answer twice. And if you use the official OpenAI or Anthropic SDKs, they have their own retry logic; configure that rather than wrapping it in a second layer of retries.
What to put in a support ticket#
If an error doesn't make sense after the tables above, open a ticket. These details let support find the request in seconds instead of guessing:
- The
x-tokens-request-idfrom the response headers (also in the error body asrequest_id). This is the most useful single item. - The time, with your time zone.
- The endpoint and model ID.
- The HTTP status and
code. - Whether the request was streaming, and whether it failed at the start or partway through.
- The client: which agent or SDK, and its version.
Never paste your API key into a ticket. If you want to tie a request to your own logs, send an x-request-id header with your own ID; the gateway echoes it back on the response alongside its own.
The full list of codes is in the errors docs, and per-minute and concurrency limits are explained in rate limits.