# Cutting your spend

> Practical ways to spend less on Tokens, from the choices that change the bill most to the ones that prevent surprises: model choice, max_tokens, context size, caching, plan allowances, background models, cancelling streams, and spend caps.

A request costs the tokens you send plus the tokens the model writes, each at the model's price per million tokens. Everything on this page changes one of three things: which price you pay, how many tokens go in, or how many come out. The sections run roughly from the choices that usually change a bill the most to the ones that mainly prevent surprises. Your own usage will differ, so measure before and after each change instead of trusting any ordering, including this one.

## Measure first

You cannot cut what you cannot see. Three places show what a request cost:

- **Every response** carries a `usage` object with the input and output token counts. Log it.
- **The Usage page** in the dashboard lists recent requests with model, tokens and cost, and charts daily spend with a breakdown sorted by spend. Start with whichever model is on top.
- **The CSV export** has one row per request with input, output and cache-read tokens and the cost in USD. See [usage, limits and alerts](/docs/usage-and-alerts).

Look for two patterns: one model that accounts for most of the spend, and requests whose input is far larger than the output. The first points to model choice, the second to context size.

## 1. Pick the right model for the job

The biggest difference is usually the price per token of the model itself. Model prices on Tokens span a wide range, and many everyday jobs (commit messages, classification, test scaffolding, small edits) do not need the most expensive model.

- Use a strong model for hard, long-running work, and a cheap one for bulk or simple work.
- Run the same real task on two models and compare **cost per finished task**, not price per token. A pricier model that finishes in fewer turns can cost less.
- Models that always think before answering add output tokens to every call.

[Choosing a model](/docs/choosing-a-model) lists representative models with their list prices, and [/models](/models) has your actual price for each one. Models that your plan includes can also be cheaper for you than pay-as-you-go, see section 5.

## 2. Set `max_tokens`

`max_tokens` is the most the model may write. Output tokens cost more than input tokens on most models, and a runaway answer is the easiest way to overspend.

- Set it to what a good answer needs, not to the largest value the model allows. A classification label needs tens of tokens, a code review a few thousand.
- If you leave it out, Tokens assumes 8,192 output tokens when it checks your balance, key cap and plan limits. Setting a smaller value lets requests pass near a cap that a larger one would not.
- On the Responses API the field is `max_output_tokens`; some newer OpenAI models use `max_completion_tokens` on Chat Completions. The gateway reads all three.
- When `finish_reason` is `length`, the answer was cut off by your limit. If that happens often, the limit is too low for the task; do not just double it everywhere.
- If your wallet cannot cover the `max_tokens` you asked for, Tokens may lower it to what the balance covers, never below 16. The answer then comes back shorter. See [plans, credits and wallet](/docs/plans-and-wallet).

```python title="max_tokens.py"
import os
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

resp = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Label this commit message as fix, feat or chore: 'handle empty cart'"}],
    max_tokens=10,
)
print(resp.choices[0].message.content, resp.usage.completion_tokens)
```

## 3. Trim the context

Chat APIs are stateless. Each request carries the whole conversation, so a long session pays for the same early messages again on every turn. Coding agents add file contents and tool output on top, which is why input tokens dominate long agent sessions.

What helps:

- **Send less per turn.** Drop tool results the model no longer needs, send the relevant file or function instead of the whole tree, and keep the system prompt short.
- **Start new conversations** for new tasks instead of continuing an old one.
- **Summarize old turns.** Replace the first part of a long conversation with a short summary, and keep the last few messages as they are.
- **Do not paste what a tool can fetch.** Agents that search and read selectively usually send less than one that is given everything up front.
- **Remember the price of a large window.** A 1M-token context is available on many models, but every token you fill is billed on every request. Some providers charge a higher rate above a size threshold, noted in [choosing a model](/docs/choosing-a-model).

To see how large your inputs really are, read `usage.prompt_tokens` from a response, or use [token counting](/docs/token-counting) before you send.

```python title="trim_history.py"
def trim_history(messages, keep_last=6):
    """Keep the system prompt and the most recent messages."""
    system = [m for m in messages if m["role"] == "system"]
    rest = [m for m in messages if m["role"] != "system"]
    return system + rest[-keep_last:]
```

This is a blunt tool: cutting turns can remove facts the model needs. Summarize instead of dropping when the early turns matter.

## 4. Use prompt caching

When the start of your prompt is the same from request to request (the system prompt, tool definitions, a big file), the provider can read it from a cache at a lower price. Cached input is shown in the usage record as cache-read tokens and priced at the model's cache-read rate where one exists. How to structure prompts for it is in [prompt caching](/docs/prompt-caching); the saving depends on the model and your traffic, so check the cache-read tokens in your usage rows.

## 5. Use your plan's allowances and free models

How a plan treats each model changes what a request costs you. From [plans, credits and wallet](/docs/plans-and-wallet):

- A plan can give a model **its own allowance**. It is a cap inside your plan's usage, not extra on top: using the model draws from the plan's credits and from the model's allowance at the same time.
- A **free model** costs nothing, does not use your credits or any allowance, and keeps working after your plan's credits run out. If a free model is good enough for a job (see the plan's page on [pricing](/pricing)), route that job to it.
- A **deal** lowers a model's price for plan usage, so the same credits go further.
- When a model's allowance is used up you get `429 model_limit_reached` until the reset. On a plan with pay-as-you-go fallback, requests to that model are instead paid from your wallet at the **full catalog price**, with no deal applied. That is a quiet way to spend more than you planned, so watch it.

## 6. Give background calls a cheap model

Several tools make small calls in the background (titles, summaries, quick edits) in addition to the main conversation. If those use the same expensive model as your main work, they add up. Where a tool lets you set a second model, set a cheap one:

- **Claude Code**: `ANTHROPIC_DEFAULT_HAIKU_MODEL` backs the `haiku` alias and also runs background tasks. If you leave it unset, background tasks use the main model. `CLAUDE_CODE_SUBAGENT_MODEL` sets the model for subagents. See [Claude Code](/docs/claude-code).
- **OpenCode**: `small_model` takes a cheaper model for lightweight tasks. See [OpenCode](/docs/opencode).
- **Crush**: you choose a large and a small model. See [Crush](/docs/crush).

Check the Usage page after a session. If you see a model you did not choose to use, a background setting is the usual cause.

## 7. Cancel streams you no longer need

Closing the connection ends the request on Tokens' side, and you are billed for what had been produced up to that point. This is how Tokens meters a request (checked in the gateway code, October 2026):

- If you close the connection, the gateway passes the cancellation on to the provider and settles the request with the tokens counted so far.
- When the provider has already reported usage, that report is used. When it has not yet (many providers send usage only at the end of a stream), Tokens estimates from the length of the text already sent. Such a record is marked as estimated.
- Input tokens are not refunded. The prompt was already read.
- An abandoned stream that is never closed keeps one of your concurrency slots until it ends, so close streams you do not read to the end. See [rate limits](/docs/rate-limits).

In practice: if a user stops a response, abort the HTTP request instead of ignoring the remaining chunks.

```python title="cancel_stream.py"
import os
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

stream = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Explain how HTTP caching works."}],
    max_tokens=800,
    stream=True,
)
text = ""
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        text += chunk.choices[0].delta.content
    if len(text) > 400:  # enough, stop paying for more
        stream.close()
        break
print(text)
```

## 8. Understand windows, plan and wallet

Plans and the wallet limit spend in different ways, and knowing which one is running tells you what to watch.

- **A plan** gives a credit allowance, and some plans spread it over a 5-hour session, a week, or a month. When a window is used up, requests return `429 window_exhausted` until it resets, whatever your wallet holds. That is a built-in brake on a heavy day.
- **The wallet** is prepaid pay-as-you-go. No usage window limits it, only its balance. A plan with fallback turns to the wallet when credits run out, at full price.
- A key's **monthly spend cap** limits one key only, see section 9.

Read where you stand with `GET /v1/tokens/usage`. It is free, not metered, and does not count towards your per-minute limit.

```bash
curl -s https://tokens.bd/v1/tokens/usage -H "Authorization: Bearer $TOKENS_API_KEY" \
  | jq -r '.windows[] | "\(.label): \(.percentUsed)% used, resets \(.resetsAt)"'
```

A window that has not been used yet in its current period is left out of the response, so an empty list after a reset is normal. Use the response to stop a batch job before it hits a window:

```python title="gate.py"
import os
import sys
import requests

resp = requests.get(
    "https://tokens.bd/v1/tokens/usage",
    headers={"Authorization": f"Bearer {os.environ['TOKENS_API_KEY']}"},
    timeout=10,
)
resp.raise_for_status()
usage = resp.json()

for window in usage["windows"]:
    if window["percentUsed"] >= 80:
        sys.exit(f"{window['label']} is {window['percentUsed']}% used, resets {window['resetsAt']}")

balance = (usage.get("wallet") or {}).get("balanceUsd")
if balance is not None and balance < 5:
    sys.exit(f"Wallet balance is ${balance:.2f}")
print("OK to start the job")
```

The 80 and 5 are example thresholds; pick your own.

## 9. Caps and alerts

Limits turn a mistake into an error instead of a bill.

- **Monthly spend cap per key.** Set it when you create the key; it cannot be edited later. A key used by CI, a script or an unattended agent should always have one. A request that would take the key over the cap is refused with `403 monthly_spend_cap_exceeded`. See [API keys](/docs/api-keys) and, for several keys, [one key per customer](/docs/one-key-per-customer).
- **Allowed models per key.** A key restricted to one cheap model cannot call an expensive one by mistake.
- **Email alerts.** Under Notifications you can get warnings at 50%, 75% and 90% of a plan limit, and a low-balance alert when the wallet drops below $5. Check that they are on.

## Checklist

1. Find the model that takes most of your spend on the Usage page.
2. Ask whether a cheaper or free model can do that job, and test it on real tasks.
3. Set `max_tokens` on every request.
4. Check input size with `usage.prompt_tokens`, then trim, summarize or cache what repeats.
5. Set a cheap model for background and small tasks in your tools.
6. Cancel streams you no longer need.
7. Put a monthly cap on every key that runs unattended, and keep the alerts on.

---
Page: https://tokens.bd/docs/cutting-costs
