Skip to content

Rate Limits

Requests per minute, concurrency, plan usage windows and per-key monthly spend caps: how each limit works, the errors they return, and how to back off correctly.

On this page

Four separate limits can stop a request before it reaches a model: requests per minute, concurrent requests, plan usage windows, and a key's monthly spend cap. Upstream providers have their own rate limits on top. This page explains each one, the error it returns, and how a client should react.

Rate limits at a glance#

LimitScopeDefaultErrorRetry-After
Requests per minuteAccount (all keys combined)60 RPM, or your plan's value429 rate_limitedSeconds until the next minute
Concurrent requestsAccountPlan's limit; 10 with a plan, 3 without429 concurrency_limit2
Usage windowsSubscriptionSet by the plan429 window_exhaustedSeconds until the window resets
Monthly spend capOne keyNone unless set at key creation403 monthly_spend_cap_exceededNot sent
Upstream provider limitProviderSet by the provider429 rate_limit_exceededPassed through if the provider sent one

Your plan's exact numbers are shown in billing and explained in plans and wallet.

Requests per minute#

The per-minute limit counts requests per account in fixed one-minute buckets that start on the clock minute. Every key on the account shares the same bucket, so creating more keys doesn't raise it. The dashboard playground has its own separate limit of 10 RPM and doesn't use up your API allowance.

When you go over, you get 429 rate_limited with Retry-After set to the seconds remaining in the current minute, so the wait is never more than 60 seconds. The error message mentions "this key", but the limit is per account.

Requests rejected for insufficient funds (402 insufficient_credits or no_funding) or an exhausted usage window still count toward the minute. A client that retries a 402 in a tight loop will also hit the per-minute limit. Don't retry 402s at all.

GET /v1/models and GET /v1/tokens/usage don't count.

Concurrency limit#

Concurrency is the number of requests in flight at once on your account. A streaming request occupies a slot from admission until the stream finishes, so a coding agent that opens several streams in parallel, or a script that fans out with asyncio.gather, can hit this before the per-minute limit.

The 429 concurrency_limit response carries Retry-After: 2. The fix is usually to cap parallelism on your side, for example with a semaphore sized below your plan's limit, rather than retrying harder. Close or abort streams you no longer need; an abandoned stream still holds its slot until it ends.

Usage windows#

Subscription plans can define usage windows, each limited in credits (expressed in USD) or in requests:

Windowtype in the APIResets
5-hour sessionsession_5h5 hours after the request that opened the window
WeeklyweeklyMondays at 00:00 UTC
MonthlymonthlyAt the end of the subscription period

When any window is used up, requests return 429 window_exhausted and Retry-After is the number of seconds until that window resets. That can be hours, so don't retry automatically: show the reset time to the user, or stop the job. Check the remaining amounts before starting a long agent run:

bash
curl -s https://tokens.bd/v1/tokens/usage -H "Authorization: Bearer $TOKENS_API_KEY"

The response shape is documented in models and usage. Alerts at 50, 75, 90 and 100 percent are covered in usage and alerts.

Monthly spend caps per key#

A key can have a monthly spend cap in USD, set when you create it. Spend is counted per calendar month (UTC). Before each request the gateway checks the key's spend so far plus the worst-case cost of the new request, based on its max_tokens (or 8,192 output tokens if unset). Near the cap, a request with a large max_tokens can be refused while a smaller one would pass.

The error is 403 monthly_spend_cap_exceeded, not a 429, because waiting a few seconds won't help. Caps can't be edited after creation; if you need a higher one, create a new key. If you only change one setting on a key that goes into a shared CI system or an agent you don't watch, make it the spend cap. More in API keys.

Retry-After and rate limit headers#

There are no X-RateLimit-Limit or X-RateLimit-Remaining headers. The only rate-limit header is Retry-After, in seconds, on 429 responses. To see remaining budget ahead of time, poll GET /v1/tokens/usage.

Client-side backoff example#

The OpenAI and Anthropic SDKs retry some 429 and 5xx errors on their own (max_retries, 2 by default). The example below turns that off and handles it explicitly, so window_exhausted isn't retried and Retry-After is honored.

import os
import random
import time

import openai
from openai import OpenAI

client = OpenAI(
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    max_retries=0,
)

RETRYABLE = {429, 500, 502, 503, 504}


def create_with_backoff(max_attempts: int = 5, **kwargs):
    delay = 1.0
    for attempt in range(1, max_attempts + 1):
        try:
            return client.chat.completions.create(**kwargs)
        except openai.APIStatusError as e:
            if (
                e.status_code not in RETRYABLE
                or e.code == "window_exhausted"
                or attempt == max_attempts
            ):
                raise
            try:
                wait = float(e.response.headers.get("retry-after", delay))
            except ValueError:
                wait = delay
            request_id = e.response.headers.get("x-tokens-request-id")
            print(f"{e.status_code} {e.code} (request {request_id}), retrying in {wait:.1f}s")
        except openai.APIConnectionError:
            if attempt == max_attempts:
                raise
            wait = delay
        time.sleep(min(wait, 60) + random.uniform(0, 0.5 * delay))
        delay = min(delay * 2, 30)


resp = create_with_backoff(
    model="deepseek/deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Say hello."}],
    max_tokens=50,
)
print(resp.choices[0].message.content)

Every successful retry is a billed request, so keep the attempt count low. The full list of retryable and non-retryable codes is in errors.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.