Skip to content
Engineering7 min read

What a 1M-token context window changes for coding agents (and what it doesn't)

Tokens Team
Engineering
3 Oct 2026
On this page

Most of the models a coding agent might use in October 2026 accept around a million tokens of context. That changes what an agent can keep in view: whole modules, long logs, a full day's session. It doesn't change the fact that every token you send is billed on every turn. This post is about what 1M-token context changes for coding agents, what it leaves the same, and how to work out what a long-context turn costs.

python
def turn_cost(prompt, cached, out, r):
    # r: per-1M rates for this prompt's tier
    fresh = prompt - cached
    usd = (fresh * r["in"]
           + cached * r["cache"]
           + out * r["out"])
    return usd / 1_000_000

What a million-token context actually changes#

  • Fewer forced summaries. With a 128K or 200K window, a long agent session has to be compacted or restarted, and details get lost in the summary. With 1M, a session can run for much longer before that happens.
  • Whole-module reasoning. A refactor that touches thirty files can have all thirty in context at once, along with their tests, rather than the agent paging them in and out.
  • Raw evidence, not excerpts. You can paste a full CI log, a long stack trace with its surrounding output, or a large API schema without trimming it first.

Those are real improvements. They're also the reason sessions get expensive without anyone noticing.

The context window is not the same on every model#

"1M" is a headline figure. The details differ, and they matter when you plan to use the whole window:

ModelContext (as published)Max outputNote
Claude Sonnet 5.51M128KSame price across the full window
GPT-6 Luna1,050,000128KPrice rises above 272K input
Gemini 3.8 Flash1,048,576 input65,536Smallest output limit here
Kimi K31,048,576131,072 defaultOutput can be raised to the full window
Qwen 3.8 Max1M131,072Input capped at 991,808 (983,616 when thinking)
MiniMax M31M128K recommendedOnly 512K guaranteed; price rises above 512K
DeepSeek V4.1 Flash1M384KPeak and off-peak rates
Grok 4.7500Knot listedPrice rises at 200K

Cost per turn: every token, every time#

Model APIs are stateless. On every turn, the agent sends the system prompt, the tool definitions, every file it has read and the entire conversation so far, and pays for all of it. The formula at the top of this post covers a single turn. To get the cost of a session, add it up over every turn.

Here's a hypothetical session to show the shape. The context starts at 50K tokens and grows by 10K each turn as the agent reads files and gets tool output back. Over 40 turns:

  • The last turn sends 440K tokens.
  • The whole session sends 9.8 million input tokens. Because the context is resent each time, the total grows roughly with the square of the session length.

On Claude Sonnet 5.5 at $2 per million, that's $19.60 of input if nothing is cached.

With caching it's very different. If each turn reads the previous prefix from cache at $0.20 and writes only its new 10K at the 5-minute rate of $2.50, the same session costs about $1.87 in reads plus $1.10 in writes, roughly $2.97. Output tokens aren't included in either figure. Prompt caching explained goes through how this works with each provider.

The number of turns is what drives the cost, more than the size of the window. A 1M window just lets the session run long enough for the number of turns to hurt.

Price tiers above 200K, 272K and 512K tokens#

Several providers charge a higher rate once a request's prompt passes a threshold. Alibaba says explicitly that its tiered Qwen models bill all tokens in the request at the tier of the input size. The examples below assume the same for OpenAI and xAI, which is how their price tables read, but check the exact rule on their pages.

ModelLower tierUpper tier
GPT-6 Lunaup to 272K: $0.10 in / $0.01 cached / $0.50 outabove 272K: $0.20 / $0.02 / $0.75
Grok 4.7under 200K: $2.00 / $6.00200K and over: $4.00 / $12.00
MiniMax M3up to 512K: $0.30 / $0.06 / $1.20above 512K: $0.60 / $0.12 / $2.40
Qwen 3.7 Plusup to 256K: $0.40 / $1.60256K to 1M: $1.20 / $4.80

The jumps are sharp. Using the formula, with 4K output tokens and no caching:

  • Qwen 3.7 Plus. A 250K prompt costs about $0.106. A 260K prompt costs about $0.331. That's 4% more context for 3.1 times the cost.
  • Grok 4.7. A 190K prompt costs about $0.40. A 210K prompt costs about $0.89.
  • GPT-6 Luna. In the 40-turn session above, the first 23 turns stay under 272K and the last 17 go over it. Input comes to about $1.59, against $0.98 if every turn had stayed in the lower tier.

Not every provider does this. Anthropic bills Claude Sonnet 5.5 at the standard rate across the full 1M window. Kimi, Z.ai's GLM models and Qwen 3.8 Max and 3.8 Flash have flat prices. If your sessions regularly run past 200K, that's a good reason to prefer one of them.

Here's one large turn priced on three models: a 600K-token prompt with 550K of it cached, and 4K output tokens.

ModelNo cache550K cached
Claude Sonnet 5.5$1.24$0.25
MiniMax M3 (above 512K tier)$0.37$0.11
GPT-6 Luna (above 272K tier)$0.12$0.02
Grok 4.7over its 500K limit

Latency: long prompts take time to process#

Before a model produces its first output token, it has to process the whole prompt. A 600K-token prompt has more to process than a 60K one, and the wait before the first token reflects that. We haven't published measurements, and we won't guess at numbers for you. Time your own requests at the context sizes you actually use.

Going through Tokens adds a gateway hop on top. The gateway allows up to 600 seconds for response headers, so long requests aren't cut off early. Request bodies are limited to 10 MB (larger ones get a 413). A million tokens of text is several megabytes once it's JSON-encoded, and base64-encoded images add up quickly, so a very large request can hit that limit.

Why retrieval still matters with a 1M window#

A large window doesn't remove the need to choose what goes into it:

  • Cost scales with what you send, per turn, as shown above. Keeping 300K tokens of loosely related files in context costs money on every turn they sit there.
  • Thresholds punish drift. A session that creeps past 200K or 272K tokens can double its rate without anyone deciding it should.
  • More context isn't more attention. A model with the right three files in front of it has less to sort through than one with three hundred. This is an opinion, not a measurement, but it matches how most agent tools are designed: they search first and read second.
  • The window isn't always fully there. MiniMax only guarantees 512K, and Qwen 3.8 Max caps input slightly below 1M.

The practical approach is to keep a stable prefix (system prompt, tool definitions, project conventions) at the start so it caches, retrieve files when they're needed, trim long tool output before it goes back to the model, and restart or compact sessions on purpose instead of waiting until they're forced.

Keep long-context spending visible#

On Tokens, set a monthly spend cap when you create the key (API keys) and turn on usage alerts at 50/75/90/100% (Usage and alerts). If your wallet runs low, the gateway may lower max_tokens to what your balance can cover. In a long agent session, that can show up as a response cut off partway. Check usage in responses, or GET /v1/tokens/usage, to see where the tokens are going.

List prices checked 3 October 2026. Tier thresholds and rates change; the maker's page is the authority.

Sources: Anthropic pricing · Claude Sonnet 5.5 overview · OpenAI pricing · xAI models · Gemini API pricing · Kimi pricing · Alibaba Model Studio pricing · MiniMax pricing · DeepSeek pricing · Z.ai pricing

Was this page helpful?

Still stuck? Open a support ticket

Use the coding models you already know, through one API

One key for OpenAI- and Anthropic-compatible tools. Pay in BDT or USD, and keep the coding agent you already use.

Create an account