A coding agent that gets stuck doesn't stop on its own. It retries the failing test, re-reads the same files, and sends a little more context each time. This tutorial covers spend caps for coding agents on Tokens: the limits you can stack, how to set them up per project, a script that checks your remaining budget, and what each limit looks like when it trips.
key cap $20 / month -> 403 monthly_spend_cap_exceeded
model list 2 models -> 403 model_not_allowed_on_key
plan window 5-hour -> 429 window_exhausted + Retry-After
wallet $0 -> 402 insufficient_creditsWhy agent loops get expensive#
Each turn of an agent loop sends the whole conversation again: system prompt, tool definitions, every file it has read, every command output. The new part of each turn is small; the re-sent part keeps growing.
Some rough arithmetic shows the shape. Say a session starts at 20,000 input tokens and each turn adds 5,000 (a file read, a test log). Turn 50 sends 265,000 tokens. Over all 50 turns, the input adds up to 7,125,000 tokens. At a list price of $0.30 per million input tokens (DeepSeek V4.1 Flash at peak, checked October 2026), that's about $2.14 of input. At $2.00 per million (Claude Sonnet 5.5's list input price), it's about $14.25. The growth is roughly quadratic in the number of turns, so a loop that runs 150 turns instead of 50 costs far more than three times as much.
Prompt caching softens this a lot when the provider supports it, because the repeated prefix is billed at the cache-read rate; prompt caching explained covers when that applies. But caching doesn't help with a loop that shouldn't be running at all. A cap does.
The limits you can stack on Tokens#
| Limit | Where it lives | What happens when it's hit |
|---|---|---|
| Monthly spend cap (USD) | Per API key, set at creation | 403 monthly_spend_cap_exceeded |
| Allowed-models list | Per API key, set at creation | 403 model_not_allowed_on_key |
| Plan usage windows | Your plan: rolling 5-hour session, weekly, monthly | 429 window_exhausted, with Retry-After |
| Wallet balance | Your account | 402 insufficient_credits |
| Requests per minute | Your account (60 by default) | 429 rate_limited |
| Concurrent requests | Your account (plan's limit; 10 with a plan, 3 without, by default) | 429 concurrency_limit |
The key-level limits are the ones you control directly, and they're the ones that contain damage to a single agent or project. The account-level ones apply to everything at once.
Set a monthly spend cap when you create the key#
In API keys, creating a key asks for a name, an optional monthly spend cap in USD, and an optional list of allowed models. You need a verified email and either a plan or a positive wallet balance to create one.
The important constraint: the cap and the allow-list can't be edited after the key is created. If you want a different cap, you create a new key, move the agent to it, and revoke the old one. Once you know that, the sensible pattern is one key per agent or project, created with its limits from the start:
claude-code-laptop: your interactive sessions, a cap you'd be comfortable losing in a bad weekopenclaw-home: an always-on agent, a smaller cap and only cheap modelsci-review-bot: an unattended job, the tightest cap of all
Active keys are limited by your plan (three by default), so decide which agents genuinely need their own. Interactive tools you watch can share one; anything that runs without you should have its own.
Rotation is immediate
Rotating a key issues a new secret and the old one stops working at once. There's no grace period, so update the agent's config in the same sitting.
Restrict models with an allow-list#
An allow-list stops an agent from wandering onto an expensive model: a /model switch in a chat session, a typo in a config file, a teammate's experiment. A request for anything outside the list fails with 403 model_not_allowed_on_key, and the error message names the models the key does allow, which makes the fix obvious.
For unattended agents, I'd pick one cheap default and one fallback and allow only those two. For your own interactive key, an allow-list is optional; you're there to notice.
Usage windows and alerts#
If your plan has usage windows, they cap consumption over time across all your keys: a rolling 5-hour session, a weekly window, a monthly window, measured in credits or requests depending on the plan. When one is used up, requests fail with 429 window_exhausted, and the Retry-After header holds the number of seconds until it resets. That can be hours. A client that retries a few seconds later and gives up is behaving correctly; one that retries forever is wasting your rate limit.
Under Notifications in the dashboard you can get alerts at 50, 75, 90 and 100 percent of a usage window, and a low-balance alert when your wallet drops below a threshold ($5 by default). Turn both on before you leave an agent running overnight. The usage and alerts docs have the details.
Poll GET /v1/tokens/usage from a script#
GET /v1/tokens/usage returns your plan, usage windows, wallet balance and the calling key's limits. It uses the same key as inference, isn't billed, and doesn't count toward your per-minute request limit. It's built for exactly this: a pre-flight check before a long agent run, or a cron job that warns you.
The quick version with curl and jq:
curl -s https://tokens.bd/v1/tokens/usage \
-H "Authorization: Bearer $TOKENS_API_KEY" \
| jq '{wallet: .wallet.balanceUsd, cap: .key.monthlySpendCapUsd, windows: [.windows[] | {type, percentUsed, resetsAt}]}'A guard script that exits non-zero when you're close to a limit, so you can chain it in front of an agent run:
// Usage: node check-budget.mjs && claude
// Node 18+. Reads TOKENS_API_KEY from the environment.
const MAX_WINDOW_PERCENT = 90;
const MIN_WALLET_USD = 2;
const res = await fetch("https://tokens.bd/v1/tokens/usage", {
headers: { Authorization: `Bearer ${process.env.TOKENS_API_KEY}` },
});
if (!res.ok) {
console.error(
`usage check failed: HTTP ${res.status} (request ${res.headers.get("x-tokens-request-id")})`
);
process.exit(2);
}
const usage = await res.json();
const problems = [];
for (const w of usage.windows) {
if (w.percentUsed >= MAX_WINDOW_PERCENT) {
problems.push(`${w.label}: ${w.percentUsed}% used, resets ${w.resetsAt}`);
}
}
// wallet is null when there is no wallet; only check it when there's no plan to fall back on
if (!usage.plan && usage.wallet && usage.wallet.balanceUsd < MIN_WALLET_USD) {
problems.push(`wallet: $${usage.wallet.balanceUsd.toFixed(2)} left`);
}
if (problems.length > 0) {
console.error(`Budget check failed:\n ${problems.join("\n ")}`);
process.exit(1);
}
console.log("Budget OK");Three things about the response shape that the script relies on. windows only contains windows that have usage in their current period, so an empty array on a fresh plan is normal. wallet is null if there's no wallet. And key.monthlySpendCapUsd is the cap itself, not how much of it the key has used so far; the endpoint doesn't report per-key spend to date.
Polling once a minute is plenty. The responses are sent with Cache-Control: no-store, so you always get current numbers. If you'd rather not write code, node tokens.mjs usage from the Tokens CLI prints the same data as bars.
Set max_tokens on every request#
A spend cap limits the month. max_tokens limits a single response, and agents leave it unset more often than you'd think. Reasoning models count their thinking as output, so an unbounded request to a model that thinks hard can produce a surprising bill on its own.
Most agents have a setting for it:
| Agent | Setting |
|---|---|
| Claude Code | CLAUDE_CODE_MAX_OUTPUT_TOKENS |
| OpenClaw | maxTokens per model (if unset, OpenClaw sends no output limit at all) |
| OpenCode | limit.output per model |
| Crush | default_max_tokens per model |
| Your own code | max_tokens / max_completion_tokens (chat), max_output_tokens (Responses), max_tokens (Messages) |
The gateway uses the requested output limit to reserve cost before the request runs. When your balance is low, it may lower max_tokens to what the balance covers, down to a floor of 16 tokens. So if answers start coming back oddly short, check your balance before you blame the model.
What it looks like when a limit trips#
All errors use the same JSON shape, with a machine-readable code. A key that has hit its monthly cap:
{
"error": {
"message": "Monthly spend cap of $20.00 for this API key has been reached. ...",
"type": "permission_denied_error",
"code": "monthly_spend_cap_exceeded",
"param": null,
"request_id": "6f1c2a9e-..."
}
}The fix is a new key with a higher cap, or waiting for the month to roll over. Retrying won't help, and a well-behaved client shouldn't retry a 403.
An exhausted plan window comes back as a 429 with a Retry-After header:
HTTP/1.1 429 Too Many Requests
Retry-After: 5400
x-tokens-request-id: 6f1c2a9e-...
{"error":{"message":"Your plan's ... limit has been reached. Resets in 5400s. ...","type":"rate_limit_error","code":"window_exhausted","param":null,"request_id":"6f1c2a9e-..."}}Inside an agent, these usually surface as a red error line with the message text. Claude Code, OpenCode and the rest will show you the message; the code is what to search for. Reading gateway errors covers every status code, plus a retry function that respects Retry-After and leaves 4xx errors alone.
If you only change one thing after reading this, create a capped key for your next unattended agent at API keys. For turning per-turn numbers into a monthly estimate, see estimate your monthly token bill; the API keys docs cover key creation in reference form.
Third-party list prices checked on 2026-10-03.
Sources: https://api-docs.deepseek.com/quick_start/pricing · https://platform.claude.com/docs/en/about-claude/pricing · https://docs.openclaw.ai/concepts/model-providers/custom-providers