Most of the models a coding agent might use in October 2026 accept around a million tokens of context. That changes what an agent can keep in view: whole modules, long logs, a full day's session. It doesn't change the fact that every token you send is billed on every turn. This post is about what 1M-token context changes for coding agents, what it leaves the same, and how to work out what a long-context turn costs.
def turn_cost(prompt, cached, out, r):
# r: per-1M rates for this prompt's tier
fresh = prompt - cached
usd = (fresh * r["in"]
+ cached * r["cache"]
+ out * r["out"])
return usd / 1_000_000What a million-token context actually changes#
- Fewer forced summaries. With a 128K or 200K window, a long agent session has to be compacted or restarted, and details get lost in the summary. With 1M, a session can run for much longer before that happens.
- Whole-module reasoning. A refactor that touches thirty files can have all thirty in context at once, along with their tests, rather than the agent paging them in and out.
- Raw evidence, not excerpts. You can paste a full CI log, a long stack trace with its surrounding output, or a large API schema without trimming it first.
Those are real improvements. They're also the reason sessions get expensive without anyone noticing.
The context window is not the same on every model#
"1M" is a headline figure. The details differ, and they matter when you plan to use the whole window:
| Model | Context (as published) | Max output | Note |
|---|---|---|---|
| Claude Sonnet 5.5 | 1M | 128K | Same price across the full window |
| GPT-6 Luna | 1,050,000 | 128K | Price rises above 272K input |
| Gemini 3.8 Flash | 1,048,576 input | 65,536 | Smallest output limit here |
| Kimi K3 | 1,048,576 | 131,072 default | Output can be raised to the full window |
| Qwen 3.8 Max | 1M | 131,072 | Input capped at 991,808 (983,616 when thinking) |
| MiniMax M3 | 1M | 128K recommended | Only 512K guaranteed; price rises above 512K |
| DeepSeek V4.1 Flash | 1M | 384K | Peak and off-peak rates |
| Grok 4.7 | 500K | not listed | Price rises at 200K |
Cost per turn: every token, every time#
Model APIs are stateless. On every turn, the agent sends the system prompt, the tool definitions, every file it has read and the entire conversation so far, and pays for all of it. The formula at the top of this post covers a single turn. To get the cost of a session, add it up over every turn.
Here's a hypothetical session to show the shape. The context starts at 50K tokens and grows by 10K each turn as the agent reads files and gets tool output back. Over 40 turns:
- The last turn sends 440K tokens.
- The whole session sends 9.8 million input tokens. Because the context is resent each time, the total grows roughly with the square of the session length.
On Claude Sonnet 5.5 at $2 per million, that's $19.60 of input if nothing is cached.
With caching it's very different. If each turn reads the previous prefix from cache at $0.20 and writes only its new 10K at the 5-minute rate of $2.50, the same session costs about $1.87 in reads plus $1.10 in writes, roughly $2.97. Output tokens aren't included in either figure. Prompt caching explained goes through how this works with each provider.
The number of turns is what drives the cost, more than the size of the window. A 1M window just lets the session run long enough for the number of turns to hurt.
Price tiers above 200K, 272K and 512K tokens#
Several providers charge a higher rate once a request's prompt passes a threshold. Alibaba says explicitly that its tiered Qwen models bill all tokens in the request at the tier of the input size. The examples below assume the same for OpenAI and xAI, which is how their price tables read, but check the exact rule on their pages.
| Model | Lower tier | Upper tier |
|---|---|---|
| GPT-6 Luna | up to 272K: $0.10 in / $0.01 cached / $0.50 out | above 272K: $0.20 / $0.02 / $0.75 |
| Grok 4.7 | under 200K: $2.00 / $6.00 | 200K and over: $4.00 / $12.00 |
| MiniMax M3 | up to 512K: $0.30 / $0.06 / $1.20 | above 512K: $0.60 / $0.12 / $2.40 |
| Qwen 3.7 Plus | up to 256K: $0.40 / $1.60 | 256K to 1M: $1.20 / $4.80 |
The jumps are sharp. Using the formula, with 4K output tokens and no caching:
- Qwen 3.7 Plus. A 250K prompt costs about $0.106. A 260K prompt costs about $0.331. That's 4% more context for 3.1 times the cost.
- Grok 4.7. A 190K prompt costs about $0.40. A 210K prompt costs about $0.89.
- GPT-6 Luna. In the 40-turn session above, the first 23 turns stay under 272K and the last 17 go over it. Input comes to about $1.59, against $0.98 if every turn had stayed in the lower tier.
Not every provider does this. Anthropic bills Claude Sonnet 5.5 at the standard rate across the full 1M window. Kimi, Z.ai's GLM models and Qwen 3.8 Max and 3.8 Flash have flat prices. If your sessions regularly run past 200K, that's a good reason to prefer one of them.
Here's one large turn priced on three models: a 600K-token prompt with 550K of it cached, and 4K output tokens.
| Model | No cache | 550K cached |
|---|---|---|
| Claude Sonnet 5.5 | $1.24 | $0.25 |
| MiniMax M3 (above 512K tier) | $0.37 | $0.11 |
| GPT-6 Luna (above 272K tier) | $0.12 | $0.02 |
| Grok 4.7 | over its 500K limit |
Latency: long prompts take time to process#
Before a model produces its first output token, it has to process the whole prompt. A 600K-token prompt has more to process than a 60K one, and the wait before the first token reflects that. We haven't published measurements, and we won't guess at numbers for you. Time your own requests at the context sizes you actually use.
Going through Tokens adds a gateway hop on top. The gateway allows up to 600 seconds for response headers, so long requests aren't cut off early. Request bodies are limited to 10 MB (larger ones get a 413). A million tokens of text is several megabytes once it's JSON-encoded, and base64-encoded images add up quickly, so a very large request can hit that limit.
Why retrieval still matters with a 1M window#
A large window doesn't remove the need to choose what goes into it:
- Cost scales with what you send, per turn, as shown above. Keeping 300K tokens of loosely related files in context costs money on every turn they sit there.
- Thresholds punish drift. A session that creeps past 200K or 272K tokens can double its rate without anyone deciding it should.
- More context isn't more attention. A model with the right three files in front of it has less to sort through than one with three hundred. This is an opinion, not a measurement, but it matches how most agent tools are designed: they search first and read second.
- The window isn't always fully there. MiniMax only guarantees 512K, and Qwen 3.8 Max caps input slightly below 1M.
The practical approach is to keep a stable prefix (system prompt, tool definitions, project conventions) at the start so it caches, retrieve files when they're needed, trim long tool output before it goes back to the model, and restart or compact sessions on purpose instead of waiting until they're forced.
Keep long-context spending visible#
On Tokens, set a monthly spend cap when you create the key (API keys) and turn on usage alerts at 50/75/90/100% (Usage and alerts). If your wallet runs low, the gateway may lower max_tokens to what your balance can cover. In a long agent session, that can show up as a response cut off partway. Check usage in responses, or GET /v1/tokens/usage, to see where the tokens are going.
List prices checked 3 October 2026. Tier thresholds and rates change; the maker's page is the authority.
Sources: Anthropic pricing · Claude Sonnet 5.5 overview · OpenAI pricing · xAI models · Gemini API pricing · Kimi pricing · Alibaba Model Studio pricing · MiniMax pricing · DeepSeek pricing · Z.ai pricing