"How much will this cost per month?" is the first question a team lead asks about coding agents, and per-token prices don't answer it on their own. This post gives you a method to estimate your monthly token bill from your own usage, works through it on two models, and shows where the estimate is most likely to be wrong.
monthly_usd = days * sessions * (
fresh_in * p_in
+ cached_in * p_cache
+ out * p_out
) / 1_000_000Step 1: measure tokens per session, don't guess#
Every other number in the estimate multiplies this one, so measure it. Pick five to ten real tasks of the kind you do every week, such as fixing a failing test, adding an endpoint or reviewing a pull request, and run them with your agent as you normally would. Then read the token counts:
- In the API response.
usageon each response gives prompt and completion tokens, and cached tokens where the provider reports them. On OpenAI-format streams, setstream_options.include_usage: trueor you won't get usage at all. - In the dashboard. Usage analytics covers 7, 14 or 30 days, and you can export CSV to work with in a spreadsheet.
- From the command line.
node tokens.mjs usage --jsonwith the Tokens CLI, orGET /v1/tokens/usagewith your key.
A "session" is one task from start to finish, not one request. A single agent session can make dozens of requests.
Step 2: split input from output#
Coding agents use far more input than output. Every turn resends the system prompt, tool definitions, files already read and the conversation so far, and gets back a short reply or a tool call. Don't assume a ratio. Read it from your measurements, because it varies a lot with the agent, the task and how much the agent reads before it writes.
This split matters because output is typically priced several times higher than input. On Claude Sonnet 5.5 it's five times ($2 in, $10 out); on DeepSeek V4.1 Flash it's four times. A model that looks cheap on output may not be the cheapest for an agent that mostly reads.
Step 3: find your cache share#
Of the input tokens, how many were served from cache? It's cached_tokens (OpenAI format) or cache_read_input_tokens (Anthropic format) divided by total input. Long sessions with a stable prefix can have a high cache share. Short one-off requests have almost none. Prompt caching explained covers what pushes this number up or down.
Cache share is the input that moves the estimate most, so measure it separately for each model. Each provider caches differently.
Step 4: sessions per day, working days per month#
Count sessions per developer per day for a typical week, not your busiest day. Multiply by working days. Use 22 if you don't have a better figure, and adjust for your team's actual schedule.
Worked example: two models, one workload#
These are assumptions to show the arithmetic, not measurements. Replace them with your own figures from steps 1 to 4.
- Per session: 1.5M input tokens, 80% of them cached, and 60K output tokens.
- Volume: 6 sessions a day for 22 working days, so 132 sessions a month for one developer.
List prices checked 3 October 2026 (USD per million tokens):
| Input | Cached input | Output | |
|---|---|---|---|
| DeepSeek V4.1 Flash (peak) | $0.30 | $0.006 | $1.20 |
| Claude Sonnet 5.5 | $2.00 | $0.20 | $10.00 |
DeepSeek V4.1 Flash, per session:
- Fresh input: 0.3M × $0.30 = $0.090
- Cached input: 1.2M × $0.006 = $0.0072
- Output: 0.06M × $1.20 = $0.072
- Total: about $0.169 a session, or about $22.33 a month.
Claude Sonnet 5.5, per session:
- Fresh input: 0.3M × $2.00 = $0.60
- Cached input: 1.2M × $0.20 = $0.24
- Output: 0.06M × $10.00 = $0.60
- Total: $1.44 a session, or $190.08 a month.
For Sonnet, part of the fresh input is written to cache at $2.50 rather than $2.00, which adds up to $0.15 a session at most. That's small enough to leave out of an estimate.
DeepSeek prices by time of day. In Dhaka, its peak hours are 07:00 to 10:00 and 12:00 to 16:00 on weekdays, which is most of a working day, so budget at the peak price and treat off-peak use as a bonus. Whether the off-peak discount reaches you depends on who you buy from.
How sensitive the estimate is to caching#
Now change one assumption: suppose only 50% of input is cached instead of 80%.
| Cache share | DeepSeek V4.1 Flash / month | Claude Sonnet 5.5 / month |
|---|---|---|
| 80% | $22.33 | $190.08 |
| 50% | $39.80 | $297.00 |
| 0% | $68.90 | $475.20 |
On these assumptions, going from 80% to zero caching multiplies the Sonnet bill by 2.5 and the DeepSeek bill by about 3.1. That's why you should measure cache share and not assume it.
Mixing models#
If you send 80% of sessions to the cheap model and escalate 20% to the stronger one, the estimate is a weighted sum: 0.8 × $22.33 + 0.2 × $190.08 ≈ $55.88 per developer per month. Frontier coding models, October 2026 describes this escalation pattern. Escalated sessions usually start with a cold cache, so use a lower cache share for that 20%.
Add a margin, then check it after a month#
Real months include failed attempts, retries, long debugging sessions and the week somebody points the agent at the whole monorepo. We'd add about a quarter on top of the estimate. That's a judgment call, not a statistic. After your first month, compare the estimate with the usage export and adjust the inputs, not just the total.
If you pay in taka, convert at the rate shown at checkout. On Tokens, the wallet is held in USD and the rate is locked when you pay (Paying in BDT).
Plans, usage windows and the shape of your usage#
A monthly total isn't the whole picture if you're on a plan with usage windows. Plans can limit usage over a rolling 5-hour session, a week or a month. When a window is used up, requests return 429 window_exhausted, with Retry-After set to the seconds until the window resets.
A team that does most of its agent work in one afternoon block can hit a 5-hour window while still well under its monthly allowance. Look at your busiest 5-hour stretch as well as the monthly sum. Plans and wallet explains how plan credits and the pay-as-you-go wallet work together.
Turn the estimate into limits#
An estimate is only useful if something stops you going far past it. On Tokens:
- Per-key monthly spend cap. Set it in USD when you create the key. Requests beyond it get
403 monthly_spend_cap_exceeded. Caps can't be edited later; to change one, create a new key with the new cap. See API keys. - Usage alerts at 50%, 75%, 90% and 100%, plus a low-balance alert (default $5). See Usage and alerts.
- Low-balance behaviour. When your balance is low, the gateway may lower
max_tokensto what the balance covers. An agent that suddenly returns cut-off answers may simply be low on funds.
For a quick first number before you've measured anything, the cost calculator on pricing uses the catalog's per-model prices. Then replace its defaults with your own measurements.
List prices checked 3 October 2026. They change; the maker's page is the authority for list prices, and the model catalog for prices on Tokens.
Sources: DeepSeek pricing · Anthropic pricing