Skip to content
Cost & Billing7 min read

How to estimate your monthly token bill for coding agents

Tokens Team
Engineering
3 Oct 2026
On this page

"How much will this cost per month?" is the first question a team lead asks about coding agents, and per-token prices don't answer it on their own. This post gives you a method to estimate your monthly token bill from your own usage, works through it on two models, and shows where the estimate is most likely to be wrong.

python
monthly_usd = days * sessions * (
    fresh_in  * p_in
  + cached_in * p_cache
  + out       * p_out
) / 1_000_000

Step 1: measure tokens per session, don't guess#

Every other number in the estimate multiplies this one, so measure it. Pick five to ten real tasks of the kind you do every week, such as fixing a failing test, adding an endpoint or reviewing a pull request, and run them with your agent as you normally would. Then read the token counts:

  • In the API response. usage on each response gives prompt and completion tokens, and cached tokens where the provider reports them. On OpenAI-format streams, set stream_options.include_usage: true or you won't get usage at all.
  • In the dashboard. Usage analytics covers 7, 14 or 30 days, and you can export CSV to work with in a spreadsheet.
  • From the command line. node tokens.mjs usage --json with the Tokens CLI, or GET /v1/tokens/usage with your key.

A "session" is one task from start to finish, not one request. A single agent session can make dozens of requests.

Step 2: split input from output#

Coding agents use far more input than output. Every turn resends the system prompt, tool definitions, files already read and the conversation so far, and gets back a short reply or a tool call. Don't assume a ratio. Read it from your measurements, because it varies a lot with the agent, the task and how much the agent reads before it writes.

This split matters because output is typically priced several times higher than input. On Claude Sonnet 5.5 it's five times ($2 in, $10 out); on DeepSeek V4.1 Flash it's four times. A model that looks cheap on output may not be the cheapest for an agent that mostly reads.

Step 3: find your cache share#

Of the input tokens, how many were served from cache? It's cached_tokens (OpenAI format) or cache_read_input_tokens (Anthropic format) divided by total input. Long sessions with a stable prefix can have a high cache share. Short one-off requests have almost none. Prompt caching explained covers what pushes this number up or down.

Cache share is the input that moves the estimate most, so measure it separately for each model. Each provider caches differently.

Step 4: sessions per day, working days per month#

Count sessions per developer per day for a typical week, not your busiest day. Multiply by working days. Use 22 if you don't have a better figure, and adjust for your team's actual schedule.

Worked example: two models, one workload#

These are assumptions to show the arithmetic, not measurements. Replace them with your own figures from steps 1 to 4.

  • Per session: 1.5M input tokens, 80% of them cached, and 60K output tokens.
  • Volume: 6 sessions a day for 22 working days, so 132 sessions a month for one developer.

List prices checked 3 October 2026 (USD per million tokens):

InputCached inputOutput
DeepSeek V4.1 Flash (peak)$0.30$0.006$1.20
Claude Sonnet 5.5$2.00$0.20$10.00

DeepSeek V4.1 Flash, per session:

  • Fresh input: 0.3M × $0.30 = $0.090
  • Cached input: 1.2M × $0.006 = $0.0072
  • Output: 0.06M × $1.20 = $0.072
  • Total: about $0.169 a session, or about $22.33 a month.

Claude Sonnet 5.5, per session:

  • Fresh input: 0.3M × $2.00 = $0.60
  • Cached input: 1.2M × $0.20 = $0.24
  • Output: 0.06M × $10.00 = $0.60
  • Total: $1.44 a session, or $190.08 a month.

For Sonnet, part of the fresh input is written to cache at $2.50 rather than $2.00, which adds up to $0.15 a session at most. That's small enough to leave out of an estimate.

DeepSeek prices by time of day. In Dhaka, its peak hours are 07:00 to 10:00 and 12:00 to 16:00 on weekdays, which is most of a working day, so budget at the peak price and treat off-peak use as a bonus. Whether the off-peak discount reaches you depends on who you buy from.

How sensitive the estimate is to caching#

Now change one assumption: suppose only 50% of input is cached instead of 80%.

Cache shareDeepSeek V4.1 Flash / monthClaude Sonnet 5.5 / month
80%$22.33$190.08
50%$39.80$297.00
0%$68.90$475.20

On these assumptions, going from 80% to zero caching multiplies the Sonnet bill by 2.5 and the DeepSeek bill by about 3.1. That's why you should measure cache share and not assume it.

Mixing models#

If you send 80% of sessions to the cheap model and escalate 20% to the stronger one, the estimate is a weighted sum: 0.8 × $22.33 + 0.2 × $190.08 ≈ $55.88 per developer per month. Frontier coding models, October 2026 describes this escalation pattern. Escalated sessions usually start with a cold cache, so use a lower cache share for that 20%.

Add a margin, then check it after a month#

Real months include failed attempts, retries, long debugging sessions and the week somebody points the agent at the whole monorepo. We'd add about a quarter on top of the estimate. That's a judgment call, not a statistic. After your first month, compare the estimate with the usage export and adjust the inputs, not just the total.

If you pay in taka, convert at the rate shown at checkout. On Tokens, the wallet is held in USD and the rate is locked when you pay (Paying in BDT).

Plans, usage windows and the shape of your usage#

A monthly total isn't the whole picture if you're on a plan with usage windows. Plans can limit usage over a rolling 5-hour session, a week or a month. When a window is used up, requests return 429 window_exhausted, with Retry-After set to the seconds until the window resets.

A team that does most of its agent work in one afternoon block can hit a 5-hour window while still well under its monthly allowance. Look at your busiest 5-hour stretch as well as the monthly sum. Plans and wallet explains how plan credits and the pay-as-you-go wallet work together.

Turn the estimate into limits#

An estimate is only useful if something stops you going far past it. On Tokens:

  • Per-key monthly spend cap. Set it in USD when you create the key. Requests beyond it get 403 monthly_spend_cap_exceeded. Caps can't be edited later; to change one, create a new key with the new cap. See API keys.
  • Usage alerts at 50%, 75%, 90% and 100%, plus a low-balance alert (default $5). See Usage and alerts.
  • Low-balance behaviour. When your balance is low, the gateway may lower max_tokens to what the balance covers. An agent that suddenly returns cut-off answers may simply be low on funds.

For a quick first number before you've measured anything, the cost calculator on pricing uses the catalog's per-model prices. Then replace its defaults with your own measurements.

List prices checked 3 October 2026. They change; the maker's page is the authority for list prices, and the model catalog for prices on Tokens.

Sources: DeepSeek pricing · Anthropic pricing

Models mentioned

Was this page helpful?

Still stuck? Open a support ticket

Use the coding models you already know, through one API

One key for OpenAI- and Anthropic-compatible tools. Pay in BDT or USD, and keep the coding agent you already use.

Create an account