This page lists every API error code the Tokens gateway returns, what each one means, and whether to retry. Branch on the code field in your code; messages are for humans and can change.
Error JSON shape#
Errors raised by the gateway look like OpenAI's:
{
"error": {
"message": "Prepaid wallet balance is insufficient for this request. Top up your wallet in the dashboard: https://tokens.bd/dashboard/billing",
"type": "insufficient_quota",
"code": "insufficient_credits",
"param": null,
"request_id": "8f0c7a4e-2b1d-4c55-9a51-3f7e2d9b6c10"
}
}| Field | Notes |
|---|---|
code | Stable, machine-readable. Use this. |
type | Broad class: authentication_error, permission_denied_error, insufficient_quota, rate_limit_error, invalid_request_error, api_error, server_error or tokens_error. |
message | Human-readable. Billing errors include a link to billing. |
request_id | Same value as the x-tokens-request-id header. Missing on unsupported_endpoint; use the header there. |
On /v1/messages, errors from the upstream provider use Anthropic's shape instead, {"type": "error", "error": {"type": "...", "message": "..."}}, while gateway errors keep the shape above. Upstream error messages are replaced with a generic message so provider internals don't leak; the HTTP status and code still tell you the category.
Error code reference#
| Status | Code | Meaning | What to do |
|---|---|---|---|
| 400 | invalid_request | Missing model (or no Content-Type: application/json), n outside 1 to 4, or the upstream rejected the request body | Fix the request. For upstream rejections, check the model supports the parameters you sent |
| 400 | model_not_available | The model is in the catalog but can't be served right now | Pick another model from GET /v1/models |
| 401 | missing_api_key | No Authorization: Bearer or x-api-key header | Set TOKENS_API_KEY in the environment that runs the tool |
| 401 | invalid_api_key | Key not recognized | Recopy the key or create a new one |
| 401/403 | upstream_auth_error | The upstream provider refused the gateway's credentials. Not your key | Retry later or switch model; report it with the request id |
| 402 | insufficient_credits | Plan credits and wallet can't cover the request | Top up or renew in billing, or lower max_tokens |
| 402 | no_funding | No active plan and no wallet balance | Subscribe or add funds |
| 402 | outstanding_debt | Earlier usage left a negative balance | Top up to clear it |
| 402 | member_cap_reached | Team accounts (if enabled): your member monthly cap is reached | Ask your team admin |
| 403 | key_inactive | Key revoked or rotated | Use the current secret |
| 403 | key_expired | Key past its expiry date | Create a new key |
| 403 | account_suspended | Account suspended | Contact support |
| 403 | model_not_allowed_on_key | Model not in this key's allowed list | Use an allowed model or another key |
| 403 | monthly_spend_cap_exceeded | Key's monthly spend cap reached (the check includes this request's worst-case cost) | Lower max_tokens, wait for the new month, or use another key |
| 403 | tier_permission_denied | Your plan doesn't include this model and you have no wallet balance | Upgrade, or add wallet funds for pay-as-you-go |
| 404 | model_not_found | Model id unknown or inactive (also returned when the upstream doesn't know the model) | Check the exact id with GET /v1/models |
| 404 | unsupported_endpoint | Path or method isn't one of the supported endpoints | See models and usage for the endpoint list |
| 413 | request_entity_too_large | Body over 10 MB | Trim context or attachments |
| 429 | rate_limited | Requests-per-minute limit reached | Wait Retry-After seconds |
| 429 | concurrency_limit | Too many requests in flight on your account | Wait Retry-After (2 s) or reduce parallelism |
| 429 | window_exhausted | A plan usage window (5-hour, weekly or monthly) is used up | Wait for the reset (Retry-After) or upgrade the plan |
| 429 | rate_limit_exceeded | The upstream provider rate limited us after failover was exhausted | Retry with backoff; honor Retry-After if present |
| 500 | lookup_failed, admission_error, catalog_error, usage_unavailable | Internal error on our side | Retry once or twice with backoff, then contact support |
| 502 | upstream_unreachable | Couldn't connect to any upstream for this model | Retry with backoff; check status |
| 5xx | upstream_error | The upstream returned a server error (status passed through) | Retry with backoff |
| 503 | no_upstream_available | No upstream is configured for this model right now | Try another model; check status |
| 503 | model_not_priced | The model has no price configured, so it can't be billed | Try another model and report it |
| 504 | upstream_timeout | Upstream didn't respond within 600 s, or the connection broke after the request was sent | Retry once; consider a smaller request |
The gateway already fails over to another upstream source on 429, 502, 503, 504 and connection errors before returning anything to you. An upstream error you see means every configured source for that model failed or the last one did.
Request ids and support tickets#
Every response, success or error, carries two headers:
| Header | Value |
|---|---|
x-tokens-request-id | The gateway's id for this request. Always generated by us. |
x-request-id | Your own x-request-id if you sent one, otherwise the same as x-tokens-request-id |
Sending your own x-request-id lets you correlate our id with your logs. To read the header with the OpenAI Python SDK:
raw = client.chat.completions.with_raw_response.create(
model="deepseek/deepseek-v4.1-flash",
messages=[{"role": "user", "content": "ping"}],
)
print(raw.headers.get("x-tokens-request-id"))
completion = raw.parse()With curl, add -i to print headers. When you open a ticket in support, include the x-tokens-request-id, the time (with timezone), the endpoint, the model and the status code. Never include the API key. We store usage metadata by request id, not prompt content, so the id is what lets us find your request. More in support.
Retry guidance#
| Retry | Codes |
|---|---|
Yes, after Retry-After | rate_limited, concurrency_limit, rate_limit_exceeded |
| Yes, with exponential backoff | upstream_unreachable, upstream_error, upstream_timeout, all 500s |
| Only after the reset time, not in a loop | window_exhausted (Retry-After can be hours) |
| No, fix something first | All 400, 401, 402, 403, 404 and 413 errors |
Backoff that works in practice: start around 1 second, double each attempt, add random jitter, cap at 30 seconds, and stop after 4 or 5 attempts. If Retry-After is present, wait at least that long. A retry is a new request and is billed if it succeeds, and the OpenAI and Anthropic SDKs already retry some of these errors on their own, so check max_retries before stacking your own loop on top. A full backoff example is in rate limits, and agent-specific fixes are in troubleshooting.