Skip to content

Cutting your spend

Practical ways to spend less on Tokens, from the choices that change the bill most to the ones that prevent surprises: model choice, max_tokens, context size, caching, plan allowances, background models, cancelling streams, and spend caps.

On this page

A request costs the tokens you send plus the tokens the model writes, each at the model's price per million tokens. Everything on this page changes one of three things: which price you pay, how many tokens go in, or how many come out. The sections run roughly from the choices that usually change a bill the most to the ones that mainly prevent surprises. Your own usage will differ, so measure before and after each change instead of trusting any ordering, including this one.

Measure first#

You cannot cut what you cannot see. Three places show what a request cost:

  • Every response carries a usage object with the input and output token counts. Log it.
  • The Usage page in the dashboard lists recent requests with model, tokens and cost, and charts daily spend with a breakdown sorted by spend. Start with whichever model is on top.
  • The CSV export has one row per request with input, output and cache-read tokens and the cost in USD. See usage, limits and alerts.

Look for two patterns: one model that accounts for most of the spend, and requests whose input is far larger than the output. The first points to model choice, the second to context size.

1. Pick the right model for the job#

The biggest difference is usually the price per token of the model itself. Model prices on Tokens span a wide range, and many everyday jobs (commit messages, classification, test scaffolding, small edits) do not need the most expensive model.

  • Use a strong model for hard, long-running work, and a cheap one for bulk or simple work.
  • Run the same real task on two models and compare cost per finished task, not price per token. A pricier model that finishes in fewer turns can cost less.
  • Models that always think before answering add output tokens to every call.

Choosing a model lists representative models with their list prices, and /models has your actual price for each one. Models that your plan includes can also be cheaper for you than pay-as-you-go, see section 5.

2. Set max_tokens#

max_tokens is the most the model may write. Output tokens cost more than input tokens on most models, and a runaway answer is the easiest way to overspend.

  • Set it to what a good answer needs, not to the largest value the model allows. A classification label needs tens of tokens, a code review a few thousand.
  • If you leave it out, Tokens assumes 8,192 output tokens when it checks your balance, key cap and plan limits. Setting a smaller value lets requests pass near a cap that a larger one would not.
  • On the Responses API the field is max_output_tokens; some newer OpenAI models use max_completion_tokens on Chat Completions. The gateway reads all three.
  • When finish_reason is length, the answer was cut off by your limit. If that happens often, the limit is too low for the task; do not just double it everywhere.
  • If your wallet cannot cover the max_tokens you asked for, Tokens may lower it to what the balance covers, never below 16. The answer then comes back shorter. See plans, credits and wallet.
max_tokens.py
import os
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

resp = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Label this commit message as fix, feat or chore: 'handle empty cart'"}],
    max_tokens=10,
)
print(resp.choices[0].message.content, resp.usage.completion_tokens)

3. Trim the context#

Chat APIs are stateless. Each request carries the whole conversation, so a long session pays for the same early messages again on every turn. Coding agents add file contents and tool output on top, which is why input tokens dominate long agent sessions.

What helps:

  • Send less per turn. Drop tool results the model no longer needs, send the relevant file or function instead of the whole tree, and keep the system prompt short.
  • Start new conversations for new tasks instead of continuing an old one.
  • Summarize old turns. Replace the first part of a long conversation with a short summary, and keep the last few messages as they are.
  • Do not paste what a tool can fetch. Agents that search and read selectively usually send less than one that is given everything up front.
  • Remember the price of a large window. A 1M-token context is available on many models, but every token you fill is billed on every request. Some providers charge a higher rate above a size threshold, noted in choosing a model.

To see how large your inputs really are, read usage.prompt_tokens from a response, or use token counting before you send.

trim_history.py
def trim_history(messages, keep_last=6):
    """Keep the system prompt and the most recent messages."""
    system = [m for m in messages if m["role"] == "system"]
    rest = [m for m in messages if m["role"] != "system"]
    return system + rest[-keep_last:]

This is a blunt tool: cutting turns can remove facts the model needs. Summarize instead of dropping when the early turns matter.

4. Use prompt caching#

When the start of your prompt is the same from request to request (the system prompt, tool definitions, a big file), the provider can read it from a cache at a lower price. Cached input is shown in the usage record as cache-read tokens and priced at the model's cache-read rate where one exists. How to structure prompts for it is in prompt caching; the saving depends on the model and your traffic, so check the cache-read tokens in your usage rows.

5. Use your plan's allowances and free models#

How a plan treats each model changes what a request costs you. From plans, credits and wallet:

  • A plan can give a model its own allowance. It is a cap inside your plan's usage, not extra on top: using the model draws from the plan's credits and from the model's allowance at the same time.
  • A free model costs nothing, does not use your credits or any allowance, and keeps working after your plan's credits run out. If a free model is good enough for a job (see the plan's page on pricing), route that job to it.
  • A deal lowers a model's price for plan usage, so the same credits go further.
  • When a model's allowance is used up you get 429 model_limit_reached until the reset. On a plan with pay-as-you-go fallback, requests to that model are instead paid from your wallet at the full catalog price, with no deal applied. That is a quiet way to spend more than you planned, so watch it.

6. Give background calls a cheap model#

Several tools make small calls in the background (titles, summaries, quick edits) in addition to the main conversation. If those use the same expensive model as your main work, they add up. Where a tool lets you set a second model, set a cheap one:

  • Claude Code: ANTHROPIC_DEFAULT_HAIKU_MODEL backs the haiku alias and also runs background tasks. If you leave it unset, background tasks use the main model. CLAUDE_CODE_SUBAGENT_MODEL sets the model for subagents. See Claude Code.
  • OpenCode: small_model takes a cheaper model for lightweight tasks. See OpenCode.
  • Crush: you choose a large and a small model. See Crush.

Check the Usage page after a session. If you see a model you did not choose to use, a background setting is the usual cause.

7. Cancel streams you no longer need#

Closing the connection ends the request on Tokens' side, and you are billed for what had been produced up to that point. This is how Tokens meters a request (checked in the gateway code, October 2026):

  • If you close the connection, the gateway passes the cancellation on to the provider and settles the request with the tokens counted so far.
  • When the provider has already reported usage, that report is used. When it has not yet (many providers send usage only at the end of a stream), Tokens estimates from the length of the text already sent. Such a record is marked as estimated.
  • Input tokens are not refunded. The prompt was already read.
  • An abandoned stream that is never closed keeps one of your concurrency slots until it ends, so close streams you do not read to the end. See rate limits.

In practice: if a user stops a response, abort the HTTP request instead of ignoring the remaining chunks.

cancel_stream.py
import os
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

stream = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Explain how HTTP caching works."}],
    max_tokens=800,
    stream=True,
)
text = ""
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        text += chunk.choices[0].delta.content
    if len(text) > 400:  # enough, stop paying for more
        stream.close()
        break
print(text)

8. Understand windows, plan and wallet#

Plans and the wallet limit spend in different ways, and knowing which one is running tells you what to watch.

  • A plan gives a credit allowance, and some plans spread it over a 5-hour session, a week, or a month. When a window is used up, requests return 429 window_exhausted until it resets, whatever your wallet holds. That is a built-in brake on a heavy day.
  • The wallet is prepaid pay-as-you-go. No usage window limits it, only its balance. A plan with fallback turns to the wallet when credits run out, at full price.
  • A key's monthly spend cap limits one key only, see section 9.

Read where you stand with GET /v1/tokens/usage. It is free, not metered, and does not count towards your per-minute limit.

bash
curl -s https://tokens.bd/v1/tokens/usage -H "Authorization: Bearer $TOKENS_API_KEY" \
  | jq -r '.windows[] | "\(.label): \(.percentUsed)% used, resets \(.resetsAt)"'

A window that has not been used yet in its current period is left out of the response, so an empty list after a reset is normal. Use the response to stop a batch job before it hits a window:

gate.py
import os
import sys
import requests

resp = requests.get(
    "https://tokens.bd/v1/tokens/usage",
    headers={"Authorization": f"Bearer {os.environ['TOKENS_API_KEY']}"},
    timeout=10,
)
resp.raise_for_status()
usage = resp.json()

for window in usage["windows"]:
    if window["percentUsed"] >= 80:
        sys.exit(f"{window['label']} is {window['percentUsed']}% used, resets {window['resetsAt']}")

balance = (usage.get("wallet") or {}).get("balanceUsd")
if balance is not None and balance < 5:
    sys.exit(f"Wallet balance is ${balance:.2f}")
print("OK to start the job")

The 80 and 5 are example thresholds; pick your own.

9. Caps and alerts#

Limits turn a mistake into an error instead of a bill.

  • Monthly spend cap per key. Set it when you create the key; it cannot be edited later. A key used by CI, a script or an unattended agent should always have one. A request that would take the key over the cap is refused with 403 monthly_spend_cap_exceeded. See API keys and, for several keys, one key per customer.
  • Allowed models per key. A key restricted to one cheap model cannot call an expensive one by mistake.
  • Email alerts. Under Notifications you can get warnings at 50%, 75% and 90% of a plan limit, and a low-balance alert when the wallet drops below $5. Check that they are on.

Checklist#

  1. Find the model that takes most of your spend on the Usage page.
  2. Ask whether a cheaper or free model can do that job, and test it on real tasks.
  3. Set max_tokens on every request.
  4. Check input size with usage.prompt_tokens, then trim, summarize or cache what repeats.
  5. Set a cheap model for background and small tasks in your tools.
  6. Cancel streams you no longer need.
  7. Put a monthly cap on every key that runs unattended, and keep the alerts on.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.