Coding agents can spend a lot in one long session, so it pays to know where your usage stands before you start, not after. This page covers the four ways to check usage on Tokens, the email alerts, and the rate limits that apply to every account.
Read the usage dashboard#
The Usage page in the dashboard shows:
- Totals for spend and number of requests, plus your recent burn rate.
- Usage limits: each usage window on your plan, how much of it is used and when it resets.
- A chart of daily spend and requests over the last 7, 14 or 30 days, with a breakdown sorted by spend.
- Recent activity: the latest requests with model, tokens and cost.
The dashboard Overview has a shorter version of the same numbers, and the Wallet page lists every top-up and charge.
Export usage as CSV#
Export CSV on the Usage page downloads your request records for the last 30 days, one row per request:
| Column | Meaning |
|---|---|
| Record ID | Internal id of the usage record |
| Request ID | The x-tokens-request-id of that request |
| Date | When it happened (UTC, ISO 8601) |
| Model | The model id you called |
| Input Tokens, Output Tokens, Cache Read Tokens | Token counts for the request |
| Cost (USD) | What it cost, to six decimal places |
| Source | Where it came from, such as v1 for API calls |
The Request ID column is useful when you need to ask support about one specific call.
Check usage from code with GET /v1/tokens/usage#
GET /v1/tokens/usage returns your plan, usage windows, wallet balance and the calling key's limits. It uses the same API key as inference and is not metered, so you can poll it from scripts, status bars or CI.
curl -s https://tokens.bd/v1/tokens/usage \
-H "Authorization: Bearer $TOKENS_API_KEY"An example response (values are illustrative):
{
"object": "tokens.usage",
"plan": { "name": "Example Plan", "tier": "monthly", "periodEnd": "2026-10-31T00:00:00.000Z" },
"windows": [
{
"type": "session_5h",
"label": "5-Hour Session",
"unit": "usd",
"limit": 5,
"used": 1.85,
"remaining": 3.15,
"percentUsed": 37,
"resetsAt": "2026-10-03T14:20:00.000Z"
},
{
"type": "weekly",
"label": "Weekly Ceiling",
"unit": "usd",
"limit": 25,
"used": 9.4,
"remaining": 15.6,
"percentUsed": 38,
"resetsAt": "2026-10-05T00:00:00.000Z"
}
],
"wallet": { "balanceUsd": 12.5 },
"key": { "monthlySpendCapUsd": 20, "allowedModels": null }
}| Field | Meaning |
|---|---|
plan | Your active plan, or null if you're pay-as-you-go only |
windows[].type | session_5h, weekly or monthly |
windows[].unit | usd for credit-based windows (in dollars), requests for request-count windows |
windows[].percentUsed | Rounded percentage of the window used |
windows[].resetsAt | When the window resets (UTC) |
wallet.balanceUsd | Wallet balance in USD |
key.monthlySpendCapUsd | This key's monthly cap, or null if it has none |
key.allowedModels | This key's allowed models, or null if it can use all of yours |
To print just the windows with jq:
curl -s https://tokens.bd/v1/tokens/usage -H "Authorization: Bearer $TOKENS_API_KEY" \
| jq -r '.windows[] | "\(.label): \(.percentUsed)% (resets \(.resetsAt))"'Check usage with the Tokens CLI#
If you set up your agents with the Tokens CLI, it reads the same endpoint:
node tokens.mjs usage
node tokens.mjs usage --jsonThe first prints your plan, a progress bar for each window with its reset time, your wallet balance and the key's cap. --json prints the raw response shown above.
Set up usage alerts#
Under Notifications in the dashboard you can choose which emails you get:
- Usage warnings at 50%, 75% and 90% of a plan limit (each threshold can be switched on or off), and when a limit is reached. Each alert is sent once per threshold, not on every request.
- Low balance, when your wallet drops below $5. On by default.
- Renewal reminders before your plan period ends. Plans don't renew automatically, so leave this on.
- Billing receipts and payment failures.
Tip
Before you leave an agent running unattended, check two things: that usage warnings are on, and that the key it uses has a monthly spend cap. A cap is set when you create a key; see API keys.
Rate limits and concurrency#
Two limits apply to every account, separately from usage windows and credits:
| Limit | Default | Error |
|---|---|---|
| Requests per minute, per user | 60 RPM (your plan can set a different value); the dashboard playground has its own 10 RPM | 429 rate_limited |
| Requests in flight at once, per account | Set by your plan; 3 without a plan | 429 concurrency_limit |
Rate limits apply to your account, not to each key. Creating more keys doesn't raise them.
Tokens doesn't send X-RateLimit-* headers. Use the Retry-After header on 429 responses instead:
| Code | What Retry-After tells you |
|---|---|
rate_limited | Seconds until the next minute starts |
concurrency_limit | 2 seconds |
window_exhausted | Seconds until the usage window resets (can be hours) |
rate_limit_exceeded | Comes from the upstream provider, after Tokens' automatic failover had no other source left to try. Back off and retry, or switch models |
Handle Retry-After in your code#
Coding agents already retry 429s. In your own code, wait for Retry-After on short limits and stop on window_exhausted, because sleeping for hours inside a request loop is rarely what you want:
import os
import time
from openai import OpenAI, RateLimitError
client = OpenAI(
base_url="https://tokens.bd/v1",
api_key=os.environ["TOKENS_API_KEY"],
max_retries=0, # we handle retries below
)
def ask(messages, attempts=5):
for attempt in range(attempts):
try:
return client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
messages=messages,
)
except RateLimitError as err:
if err.code == "window_exhausted":
raise # resets in hours; report it instead of sleeping
wait = float(err.response.headers.get("retry-after", 2 ** attempt))
time.sleep(wait)
raise RuntimeError("Still rate limited after retries")
print(ask([{"role": "user", "content": "One-line summary of HTTP 429."}]).choices[0].message.content)To cut concurrency errors, limit how many requests your script sends in parallel to your plan's concurrency limit. More detail is in rate limits and errors.
Related#
- Plans, credits and wallet: what happens when a window or balance runs out
- API keys: monthly spend caps per key
- Models and usage endpoints