Four separate limits can stop a request before it reaches a model: requests per minute, concurrent requests, plan usage windows, and a key's monthly spend cap. Upstream providers have their own rate limits on top. This page explains each one, the error it returns, and how a client should react.
Rate limits at a glance#
| Limit | Scope | Default | Error | Retry-After |
|---|---|---|---|---|
| Requests per minute | Account (all keys combined) | 60 RPM, or your plan's value | 429 rate_limited | Seconds until the next minute |
| Concurrent requests | Account | Plan's limit; 10 with a plan, 3 without | 429 concurrency_limit | 2 |
| Usage windows | Subscription | Set by the plan | 429 window_exhausted | Seconds until the window resets |
| Monthly spend cap | One key | None unless set at key creation | 403 monthly_spend_cap_exceeded | Not sent |
| Upstream provider limit | Provider | Set by the provider | 429 rate_limit_exceeded | Passed through if the provider sent one |
Your plan's exact numbers are shown in billing and explained in plans and wallet.
Requests per minute#
The per-minute limit counts requests per account in fixed one-minute buckets that start on the clock minute. Every key on the account shares the same bucket, so creating more keys doesn't raise it. The dashboard playground has its own separate limit of 10 RPM and doesn't use up your API allowance.
When you go over, you get 429 rate_limited with Retry-After set to the seconds remaining in the current minute, so the wait is never more than 60 seconds. The error message mentions "this key", but the limit is per account.
Requests rejected for insufficient funds (402 insufficient_credits or no_funding) or an exhausted usage window still count toward the minute. A client that retries a 402 in a tight loop will also hit the per-minute limit. Don't retry 402s at all.
GET /v1/models and GET /v1/tokens/usage don't count.
Concurrency limit#
Concurrency is the number of requests in flight at once on your account. A streaming request occupies a slot from admission until the stream finishes, so a coding agent that opens several streams in parallel, or a script that fans out with asyncio.gather, can hit this before the per-minute limit.
The 429 concurrency_limit response carries Retry-After: 2. The fix is usually to cap parallelism on your side, for example with a semaphore sized below your plan's limit, rather than retrying harder. Close or abort streams you no longer need; an abandoned stream still holds its slot until it ends.
Usage windows#
Subscription plans can define usage windows, each limited in credits (expressed in USD) or in requests:
| Window | type in the API | Resets |
|---|---|---|
| 5-hour session | session_5h | 5 hours after the request that opened the window |
| Weekly | weekly | Mondays at 00:00 UTC |
| Monthly | monthly | At the end of the subscription period |
When any window is used up, requests return 429 window_exhausted and Retry-After is the number of seconds until that window resets. That can be hours, so don't retry automatically: show the reset time to the user, or stop the job. Check the remaining amounts before starting a long agent run:
curl -s https://tokens.bd/v1/tokens/usage -H "Authorization: Bearer $TOKENS_API_KEY"The response shape is documented in models and usage. Alerts at 50, 75, 90 and 100 percent are covered in usage and alerts.
Monthly spend caps per key#
A key can have a monthly spend cap in USD, set when you create it. Spend is counted per calendar month (UTC). Before each request the gateway checks the key's spend so far plus the worst-case cost of the new request, based on its max_tokens (or 8,192 output tokens if unset). Near the cap, a request with a large max_tokens can be refused while a smaller one would pass.
The error is 403 monthly_spend_cap_exceeded, not a 429, because waiting a few seconds won't help. Caps can't be edited after creation; if you need a higher one, create a new key. If you only change one setting on a key that goes into a shared CI system or an agent you don't watch, make it the spend cap. More in API keys.
Retry-After and rate limit headers#
There are no X-RateLimit-Limit or X-RateLimit-Remaining headers. The only rate-limit header is Retry-After, in seconds, on 429 responses. To see remaining budget ahead of time, poll GET /v1/tokens/usage.
Client-side backoff example#
The OpenAI and Anthropic SDKs retry some 429 and 5xx errors on their own (max_retries, 2 by default). The example below turns that off and handles it explicitly, so window_exhausted isn't retried and Retry-After is honored.
import os
import random
import time
import openai
from openai import OpenAI
client = OpenAI(
base_url="https://tokens.bd/v1",
api_key=os.environ["TOKENS_API_KEY"],
max_retries=0,
)
RETRYABLE = {429, 500, 502, 503, 504}
def create_with_backoff(max_attempts: int = 5, **kwargs):
delay = 1.0
for attempt in range(1, max_attempts + 1):
try:
return client.chat.completions.create(**kwargs)
except openai.APIStatusError as e:
if (
e.status_code not in RETRYABLE
or e.code == "window_exhausted"
or attempt == max_attempts
):
raise
try:
wait = float(e.response.headers.get("retry-after", delay))
except ValueError:
wait = delay
request_id = e.response.headers.get("x-tokens-request-id")
print(f"{e.status_code} {e.code} (request {request_id}), retrying in {wait:.1f}s")
except openai.APIConnectionError:
if attempt == max_attempts:
raise
wait = delay
time.sleep(min(wait, 60) + random.uniform(0, 0.5 * delay))
delay = min(delay * 2, 30)
resp = create_with_backoff(
model="deepseek/deepseek-v4.1-flash",
messages=[{"role": "user", "content": "Say hello."}],
max_tokens=50,
)
print(resp.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://tokens.bd/v1",
apiKey: process.env.TOKENS_API_KEY,
maxRetries: 0,
});
const RETRYABLE = new Set([429, 500, 502, 503, 504]);
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
function retryAfterMs(err: InstanceType<typeof OpenAI.APIError>): number | undefined {
const h: unknown = err.headers; // a Headers object or a plain record, depending on SDK version
const value =
h instanceof Headers
? h.get("retry-after")
: (h as Record<string, string | undefined> | undefined)?.["retry-after"];
return value ? Number(value) * 1000 : undefined;
}
async function createWithBackoff(
body: OpenAI.Chat.ChatCompletionCreateParamsNonStreaming,
maxAttempts = 5
) {
let delay = 1000;
for (let attempt = 1; ; attempt++) {
try {
return await client.chat.completions.create(body);
} catch (err) {
if (!(err instanceof OpenAI.APIError) || attempt >= maxAttempts) throw err;
const connectionError = err instanceof OpenAI.APIConnectionError;
if (
!connectionError &&
(!RETRYABLE.has(err.status ?? 0) || err.code === "window_exhausted")
) {
throw err;
}
const wait = (!connectionError && retryAfterMs(err)) || delay;
await sleep(Math.min(wait, 60_000) + Math.random() * 0.5 * delay);
delay = Math.min(delay * 2, 30_000);
}
}
}
const resp = await createWithBackoff({
model: "deepseek/deepseek-v4.1-flash",
messages: [{ role: "user", content: "Say hello." }],
max_tokens: 50,
});
console.log(resp.choices[0].message.content);Every successful retry is a billed request, so keep the attempt count low. The full list of retryable and non-retryable codes is in errors.