Skip to content

Usage, limits and alerts

Track spend and remaining allowance on Tokens from the dashboard, a CSV export, the GET /v1/tokens/usage endpoint or the CLI, set up usage alerts, and handle rate limits and Retry-After correctly.

On this page

Coding agents can spend a lot in one long session, so it pays to know where your usage stands before you start, not after. This page covers the four ways to check usage on Tokens, the email alerts, and the rate limits that apply to every account.

Read the usage dashboard#

The Usage page in the dashboard shows:

  • Totals for spend and number of requests, plus your recent burn rate.
  • Usage limits: each usage window on your plan, how much of it is used and when it resets.
  • A chart of daily spend and requests over the last 7, 14 or 30 days, with a breakdown sorted by spend.
  • Recent activity: the latest requests with model, tokens and cost.

The dashboard Overview has a shorter version of the same numbers, and the Wallet page lists every top-up and charge.

Export usage as CSV#

Export CSV on the Usage page downloads your request records for the last 30 days, one row per request:

ColumnMeaning
Record IDInternal id of the usage record
Request IDThe x-tokens-request-id of that request
DateWhen it happened (UTC, ISO 8601)
ModelThe model id you called
Input Tokens, Output Tokens, Cache Read TokensToken counts for the request
Cost (USD)What it cost, to six decimal places
SourceWhere it came from, such as v1 for API calls

The Request ID column is useful when you need to ask support about one specific call.

Check usage from code with GET /v1/tokens/usage#

GET /v1/tokens/usage returns your plan, usage windows, wallet balance and the calling key's limits. It uses the same API key as inference and is not metered, so you can poll it from scripts, status bars or CI.

bash
curl -s https://tokens.bd/v1/tokens/usage \
  -H "Authorization: Bearer $TOKENS_API_KEY"

An example response (values are illustrative):

json
{
  "object": "tokens.usage",
  "plan": { "name": "Example Plan", "tier": "monthly", "periodEnd": "2026-10-31T00:00:00.000Z" },
  "windows": [
    {
      "type": "session_5h",
      "label": "5-Hour Session",
      "unit": "usd",
      "limit": 5,
      "used": 1.85,
      "remaining": 3.15,
      "percentUsed": 37,
      "resetsAt": "2026-10-03T14:20:00.000Z"
    },
    {
      "type": "weekly",
      "label": "Weekly Ceiling",
      "unit": "usd",
      "limit": 25,
      "used": 9.4,
      "remaining": 15.6,
      "percentUsed": 38,
      "resetsAt": "2026-10-05T00:00:00.000Z"
    }
  ],
  "wallet": { "balanceUsd": 12.5 },
  "key": { "monthlySpendCapUsd": 20, "allowedModels": null }
}
FieldMeaning
planYour active plan, or null if you're pay-as-you-go only
windows[].typesession_5h, weekly or monthly
windows[].unitusd for credit-based windows (in dollars), requests for request-count windows
windows[].percentUsedRounded percentage of the window used
windows[].resetsAtWhen the window resets (UTC)
wallet.balanceUsdWallet balance in USD
key.monthlySpendCapUsdThis key's monthly cap, or null if it has none
key.allowedModelsThis key's allowed models, or null if it can use all of yours

To print just the windows with jq:

bash
curl -s https://tokens.bd/v1/tokens/usage -H "Authorization: Bearer $TOKENS_API_KEY" \
  | jq -r '.windows[] | "\(.label): \(.percentUsed)% (resets \(.resetsAt))"'

Check usage with the Tokens CLI#

If you set up your agents with the Tokens CLI, it reads the same endpoint:

bash
node tokens.mjs usage
node tokens.mjs usage --json

The first prints your plan, a progress bar for each window with its reset time, your wallet balance and the key's cap. --json prints the raw response shown above.

Set up usage alerts#

Under Notifications in the dashboard you can choose which emails you get:

  • Usage warnings at 50%, 75% and 90% of a plan limit (each threshold can be switched on or off), and when a limit is reached. Each alert is sent once per threshold, not on every request.
  • Low balance, when your wallet drops below $5. On by default.
  • Renewal reminders before your plan period ends. Plans don't renew automatically, so leave this on.
  • Billing receipts and payment failures.

Tip

Before you leave an agent running unattended, check two things: that usage warnings are on, and that the key it uses has a monthly spend cap. A cap is set when you create a key; see API keys.

Rate limits and concurrency#

Two limits apply to every account, separately from usage windows and credits:

LimitDefaultError
Requests per minute, per user60 RPM (your plan can set a different value); the dashboard playground has its own 10 RPM429 rate_limited
Requests in flight at once, per accountSet by your plan; 3 without a plan429 concurrency_limit

Rate limits apply to your account, not to each key. Creating more keys doesn't raise them.

Tokens doesn't send X-RateLimit-* headers. Use the Retry-After header on 429 responses instead:

CodeWhat Retry-After tells you
rate_limitedSeconds until the next minute starts
concurrency_limit2 seconds
window_exhaustedSeconds until the usage window resets (can be hours)
rate_limit_exceededComes from the upstream provider, after Tokens' automatic failover had no other source left to try. Back off and retry, or switch models

Handle Retry-After in your code#

Coding agents already retry 429s. In your own code, wait for Retry-After on short limits and stop on window_exhausted, because sleeping for hours inside a request loop is rarely what you want:

retry.py
import os
import time
from openai import OpenAI, RateLimitError

client = OpenAI(
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    max_retries=0,  # we handle retries below
)

def ask(messages, attempts=5):
    for attempt in range(attempts):
        try:
            return client.chat.completions.create(
                model="deepseek/deepseek-v4.1-flash",
                messages=messages,
            )
        except RateLimitError as err:
            if err.code == "window_exhausted":
                raise  # resets in hours; report it instead of sleeping
            wait = float(err.response.headers.get("retry-after", 2 ** attempt))
            time.sleep(wait)
    raise RuntimeError("Still rate limited after retries")

print(ask([{"role": "user", "content": "One-line summary of HTTP 429."}]).choices[0].message.content)

To cut concurrency errors, limit how many requests your script sends in parallel to your plan's concurrency limit. More detail is in rate limits and errors.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.