Skip to content
Tutorials10 min read

Reading AI Gateway Errors: 401, 402, 403, 404, 429 and 5xx

Tokens Team
Engineering
3 Oct 2026
On this page

An error from an AI gateway can come from three places: your request, your account, or a model provider somewhere behind the gateway. The HTTP status narrows it down, and the code field pins it. This guide goes through the AI gateway errors you'll see from Tokens, status by status, with what each one means and whether retrying makes any sense.

text
401  your key            fix the key, don't retry
402  your balance        top up, don't retry
403  your key's limits   new key or different model, don't retry
404  your URL or model   fix the request, don't retry
429  slow down           wait for Retry-After, then retry
5xx  upstream trouble    back off and retry, a few times

The shape of an error#

Every error the gateway itself produces has the same JSON body, in OpenAI's style:

json
{
  "error": {
    "message": "Model 'deepseek/deepseek-v4.1-flash' is not permitted on this API key. Permitted models: ...",
    "type": "permission_denied_error",
    "code": "model_not_allowed_on_key",
    "param": null,
    "request_id": "6f1c2a9e-..."
  }
}

Branch on code, not on message. Messages are written for people and can change; codes are the contract. Every response, success or failure, also carries an x-tokens-request-id header. Keep it; it's what support needs.

One exception to know about if you use Claude Code or the Anthropic SDK. On /v1/messages, errors from the gateway itself (authentication, caps, balance) use the shape above, but errors passed back from an upstream provider use Anthropic's shape: {"type": "error", "error": {"type": "...", "message": "..."}}, with no code field. For those, the status and the x-tokens-request-id header are what you have.

401: the key wasn't accepted#

CodeMeaning
missing_api_keyNo key in the request. Send Authorization: Bearer <key> or x-api-key: <key>.
invalid_api_keyA key was sent, but it doesn't match an active key.

The usual causes are mundane: an environment variable that isn't set in the shell the agent runs in, a key with a trailing newline from copy-paste, or a key that was rotated. Rotation replaces the secret immediately, so an agent still holding the old one gets 401 on its next request. Keys look like tok_live_ followed by 48 hex characters; if yours doesn't, it's been truncated somewhere.

Don't retry a 401. It won't fix itself.

402: the account can't pay for this request#

CodeMeaning
insufficient_creditsYour wallet balance doesn't cover the request
no_fundingNo active subscription and no funded wallet
outstanding_debtEarlier usage left a negative balance that has to be cleared first
member_cap_reachedOn a team account, your member monthly cap is used up; your admin raises it

Every 402 message links to billing. Note that a low balance doesn't always mean a 402: when the balance covers part of a request, the gateway may lower max_tokens to what it can pay for, down to 16 tokens. If answers suddenly come back truncated, check the balance.

403: the key isn't allowed to do this#

CodeMeaningFix
model_not_allowed_on_keyThe key has an allowed-models list and this model isn't on it. The message lists the allowed ones.Use an allowed model, or a different key
monthly_spend_cap_exceededThe key reached its monthly spend capA new key with a higher cap, or wait for next month
tier_permission_deniedYour plan doesn't include this model and you have no wallet balance to pay for it as you goAdd funds or change plan
key_inactiveThe key was revoked or suspendedCreate a new key
key_expiredThe key had an expiry date that has passedCreate a new key
account_suspendedThe account is suspendedContact support

Caps and allow-lists can't be edited on an existing key, so the fix for the first two is almost always a new key. Spend caps for coding agents explains how to plan keys so this happens on purpose rather than by accident.

There's one 403 that isn't about you: if the code is upstream_auth_error, the gateway's own credentials for a provider were rejected. The same code can come with a 401 status. Nothing in your setup will fix it; open a ticket with the request ID.

400, 404 and 413: the request itself#

Status and codeMeaning
404 unsupported_endpointThe path isn't one the gateway serves
404 model_not_foundThe model ID doesn't match an active model
400 model_not_availableThe model exists but isn't available to you right now
400 invalid_requestThe upstream provider rejected the request body
413The body is over 10 MB

unsupported_endpoint almost always means a base URL mistake. OpenAI-style clients need https://tokens.bd/v1; Anthropic-style clients, including Claude Code, need https://tokens.bd, because they add /v1/messages themselves. Mix them up and you get /v1/v1/... or a missing /v1. The other cause is calling an endpoint the gateway doesn't offer: images, audio, files, batches, assistants, fine-tuning and moderations all return this code.

For model_not_found, compare your ID against GET /v1/models character by character. IDs are provider/model aliases, like deepseek/deepseek-v4.1-flash. Agents that reference models as <provider-id>/<model-id> (OpenCode's tokens/deepseek/..., for example) should send only the part after their own provider ID; if the error message shows a doubled prefix, the agent config is the place to look.

invalid_request is the vaguest one, because the gateway returns a generic message rather than the provider's original text. Common causes are parameters the model doesn't accept (a temperature outside its range, a tool format it doesn't support) and conversations that exceed the context window. Try the same request with a minimal body; if that works, add fields back until it breaks.

429: too much, too fast#

CodeMeaningRetry-After
rate_limitedOver your per-minute request limit (60 by default)Seconds until you can retry
concurrency_limitToo many requests in flight at once for your plan2 seconds
window_exhaustedA plan usage window (5-hour, weekly or monthly) is used upSeconds until the window resets, which can be hours
rate_limit_exceededAn upstream provider rate-limited the request, on every source the gateway triedPassed through when the provider sends one

There are no X-RateLimit-* headers to watch; Retry-After on the 429 is the signal. The first two clear in seconds and are safe to retry after the wait. window_exhausted is different in kind: retrying in a loop just wastes requests until the reset time, so surface it to a human. And concurrency_limit from a coding agent often means too many parallel subagents or tool calls; lowering the agent's parallelism fixes it more reliably than retrying.

5xx: something upstream failed#

Status and codeMeaning
502 upstream_unreachableNo provider source could be reached
503 no_upstream_availableNo source is currently configured or healthy for this model
504 upstream_timeoutThe provider stopped responding mid-request
Other 5xx, upstream_errorThe provider returned a server error

Before any of these reach you, the gateway has already tried: on a 429, 502, 503, 504 or connection error from one source, it fails over to the next source for that model automatically. So a 5xx means every option failed for that request. A couple of retries with backoff is reasonable; a tight loop isn't. If it persists, check the status page, which shows the inference API and upstream providers separately.

Long requests by themselves aren't a problem: the gateway waits up to 600 seconds for response headers. If you're seeing timeouts, check your own client's timeout first, since many SDKs default to something shorter.

A retry function that honours Retry-After#

The rules: retry 429 and 5xx and network failures; never retry other 4xx; use Retry-After when present; otherwise back off exponentially with jitter; give up after a few attempts; and treat a very long Retry-After as "stop" rather than sleeping for an hour.

const RETRYABLE = new Set([429, 500, 502, 503, 504]);

export async function callWithRetry(url, body, { attempts = 4, maxWaitS = 60 } = {}) {
  for (let i = 0; i < attempts; i++) {
    let res;
    try {
      res = await fetch(url, {
        method: "POST",
        headers: {
          Authorization: `Bearer ${process.env.TOKENS_API_KEY}`,
          "Content-Type": "application/json",
        },
        body: JSON.stringify(body),
      });
    } catch (err) {
      if (i === attempts - 1) throw err; // network error: retry, then give up
      await sleep(backoff(i));
      continue;
    }

    if (res.ok) return res.json();

    const requestId = res.headers.get("x-tokens-request-id");
    const err = await res.json().catch(() => ({}));
    const code = err.error?.code ?? `http_${res.status}`;

    if (!RETRYABLE.has(res.status) || i === attempts - 1) {
      throw new Error(`${res.status} ${code} (request ${requestId}): ${err.error?.message ?? ""}`);
    }

    const retryAfter = Number(res.headers.get("retry-after"));
    const waitS = Number.isFinite(retryAfter) && retryAfter > 0 ? retryAfter : backoff(i) / 1000;
    if (waitS > maxWaitS) {
      throw new Error(
        `${res.status} ${code}: retry in ${waitS}s is too long (request ${requestId})`
      );
    }
    await sleep(waitS * 1000);
  }
}

const backoff = (i) => Math.min(30_000, 1000 * 2 ** i) * (0.5 + Math.random() / 2);
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));

The maxWaitS cutoff is what keeps a window_exhausted from turning into an hour-long sleep. Two caveats for streaming. Only retry if the stream failed before any content arrived; once tokens have streamed, they've been generated and metered, so a retry pays for the answer twice. And if you use the official OpenAI or Anthropic SDKs, they have their own retry logic; configure that rather than wrapping it in a second layer of retries.

What to put in a support ticket#

If an error doesn't make sense after the tables above, open a ticket. These details let support find the request in seconds instead of guessing:

  • The x-tokens-request-id from the response headers (also in the error body as request_id). This is the most useful single item.
  • The time, with your time zone.
  • The endpoint and model ID.
  • The HTTP status and code.
  • Whether the request was streaming, and whether it failed at the start or partway through.
  • The client: which agent or SDK, and its version.

Never paste your API key into a ticket. If you want to tie a request to your own logs, send an x-request-id header with your own ID; the gateway echoes it back on the response alongside its own.

The full list of codes is in the errors docs, and per-minute and concurrency limits are explained in rate limits.

Was this page helpful?

Still stuck? Open a support ticket

Use the coding models you already know, through one API

One key for OpenAI- and Anthropic-compatible tools. Pay in BDT or USD, and keep the coding agent you already use.

Create an account