Skip to content
Engineering9 min read

Prompt caching explained: how cache-hit pricing works across providers

Tokens Team
Engineering
3 Oct 2026
On this page

Prompt caching is the biggest single cost lever for a coding agent, and the one people understand least. This post explains how prompt caching works, how each major provider prices cache hits and cache writes, what quietly stops your requests from hitting the cache, and what a gateway can and can't do about it.

json
"usage": {
  "prompt_tokens": 182400,
  "completion_tokens": 1210,
  "prompt_tokens_details": {
    "cached_tokens": 176128
  }
}

That's the shape of the answer you're after (example numbers): most of a long prompt served from cache, billed at a fraction of the input price.

How prompt caching works#

Providers keep the processed form of recent prompts for a short time. When a new request starts with exactly the same tokens as one they've already processed, they reuse the work for that shared prefix and bill those tokens at the cached-input rate. Everything after the first token that differs is processed and billed normally.

Two consequences follow from that:

  • Only the prefix counts. If the first 100K tokens match and then something changes, you get 100K cached tokens. If the very first line differs, you get none.
  • Order matters more than content. The same text in a different order is a different prefix. For Anthropic's API, the order is tools, then system, then messages. Stable content has to come first.

Cache-hit pricing by provider#

List prices checked 3 October 2026, USD per million tokens, from each maker's own pricing page.

ModelInputCache readCache writeRead as share of input
DeepSeek V4.1 Flash$0.30 peak / $0.15 off-peak$0.006 / $0.003not listed2%
DeepSeek V4 Pro$1.32 / $0.66$0.044 / $0.022not listed3.3%
MiMo V2.6 Pro$0.435$0.0036not listedunder 1%
Claude Sonnet 5.5$2.00$0.20$2.50 (5 min), $4.00 (1 h)10%
GPT-6 Luna (up to 272K)$0.10$0.01not listed10%
GPT-5.6 Sol (promo, up to 272K)$4.00$0.40not listed10%
Gemini 3.8 Flash (to 31 Dec 2026)$0.75$0.075not listed10%
Kimi K3$3.00$0.30$3.00 (5 min), $6.00 (1 h)10%
Qwen 3.8 Flash$0.15$0.016$0.20 (explicit)about 11%
Qwen 3.8 Max$2.00$0.25 implicit / $0.17 explicit$2.50 (explicit)12.5% / 8.5%
GLM-5.3$1.40$0.26not listed (storage "limited-time free")about 19%
MiniMax M3 (up to 512K)$0.30$0.06not listed20%
Grok 4.7 (under 200K)$2.00$0.50not listed25%

The range is wide. A cached token on DeepSeek V4.1 Flash costs 2% of a fresh one. On Grok 4.7 it costs 25%. If your workload is mostly re-reading the same context, the cache-read column tells you more about your bill than the input column.

Claude: explicit breakpoints, a write premium and two TTLs#

On Anthropic's API you mark what to cache with cache_control blocks. There are a few rules:

  • Up to four breakpoints per request. There's also a top-level option that caches the last cacheable block automatically.
  • Two TTLs. The default is 5 minutes, and 1 hour is available.
  • A minimum prefix length. Prefixes shorter than the model's minimum (512 to 4,096 tokens depending on the model) silently don't cache.

The response tells you what happened: cache_creation_input_tokens (written), cache_read_input_tokens (served from cache) and input_tokens (everything else, at full price).

Writes cost more than plain input, so it's worth checking when they pay off. On Sonnet 5.5:

  • 5-minute write: $0.50 per million above normal input. Each later read saves $1.80 per million, so the first reuse pays for it.
  • 1-hour write: $2.00 per million above input. It needs two reads before it pays off.

Use 1 hour when you expect long gaps between turns, such as a developer reading a diff before replying.

Kimi K3: writes are listed separately#

Moonshot lists cache writes as a separate charge: $3.00 per million for a 5-minute TTL and $6.00 for 1 hour. Reads are $0.30. Check Moonshot's docs for exactly which tokens count as written before you model it. The direction is the same as Claude's, though: the longer TTL costs more, and it only pays off if the gaps between turns really need it.

OpenAI GPT: a cached-input price, no write price listed#

OpenAI's pricing table lists a cached-input rate at 10% of input for GPT-6 Luna and GPT-5.6 Sol, and no separate write charge. Above 272K input tokens, the cached rate moves to the higher tier along with everything else ($0.02 for Luna).

Qwen: implicit or explicit caching#

Alibaba has two modes:

  • Implicit caching needs no request changes. Hits cost about 20% of input under Alibaba's general rule, though Qwen 3.8 Max lists $0.25 (12.5%).
  • Explicit caching charges 125% of input to create the cache, and hits cost about 10%.

On Qwen 3.8 Max, explicit creation costs $0.50 per million more than input, and each explicit read saves $0.08 compared with an implicit hit. So explicit caching only pays off against implicit after about seven reads of the same prefix. For a long agent session that's easy to reach. For one-off prompts it isn't.

DeepSeek: tiny cache price, time-of-day rates#

DeepSeek's cache hits cost $0.006 per million at peak and $0.003 off-peak on V4.1 Flash. The half-price off-peak rule applies to cache hits as well. At these rates the cache makes the input side of a long session almost free. Output becomes most of the bill.

Why coding agents get the most out of caching#

A coding agent resends almost the same thing on every turn. The system prompt and tool definitions don't change, and the files it has already read sit in the same place in the conversation. Each turn adds a little at the end. That's the ideal shape for prefix caching.

Here's a hypothetical session. It has a 100K-token stable prefix (system prompt, tools, project files), adds 6K tokens per turn, and runs for 30 turns. That's 5.61 million input tokens in total. Assuming every turn lands within the cache TTL:

ModelNo cachingWith caching
Claude Sonnet 5.5$11.22about $1.75 ($0.685 writes + $1.067 reads)
DeepSeek V4.1 Flash (peak)$1.68about $0.11

Output isn't included. The arithmetic: 274K tokens are new at some point in the session (written once), and the remaining 5.34 million are reads.

To get this, put the stable material first and the volatile material last:

  1. System prompt and tool definitions, unchanged between turns.
  2. Project context (conventions file, key source files), in a fixed order.
  3. The conversation, appended to, never edited.

What breaks cache hits#

These are the usual causes when cached_tokens stays at zero:

  • A timestamp or request ID in the system prompt. It changes the first few hundred tokens on every request, so nothing after it can match.
  • Tool definitions in a different order, or an MCP server that adds or removes tools between turns.
  • Rewriting history. Summarizing or compacting earlier turns gives you a new prefix. Sometimes that's the right trade, but it does cost a cache miss.
  • Switching models. Caches are kept per model. A request escalated from a cheap model to a frontier one starts cold.
  • Idle time longer than the TTL. Five minutes goes quickly when someone is reviewing a diff.

Prompt caching through a gateway: what you control and what you don't#

Tokens forwards your request body to the upstream provider. It doesn't add or remove cache markers, so cache_control blocks in a /v1/messages request go upstream as you wrote them. Whether a request hits the cache is decided by the upstream that serves it.

On billing, the gateway reads cached-token counts from the upstream's usage (prompt_tokens_details.cached_tokens in the OpenAI format, cache_read_input_tokens and cache_creation_input_tokens in the Anthropic format). It bills them at the cache-read and cache-write prices listed for that model in the catalog. If a model has no cache price listed, those tokens are billed at its normal input price.

What you can't control through any gateway:

  • The cache itself. TTLs, eviction and minimum lengths are the provider's.
  • Failover. If a request is retried on a different upstream source after a 429 or 5xx, that source hasn't seen your prefix, so expect a miss on that request.
  • Usage on OpenAI-format streams. It only arrives if you set stream_options.include_usage: true. Without it, you can't see your cache hits.

To check what's happening, look at usage in responses, then compare cached and uncached input over a day in the dashboard's usage analytics. Usage and alerts covers where those numbers are shown. If you're deciding which model to cache against, the cache-read column above is a good place to start, alongside the cheap-models guide. For long sessions, what 1M-token context changes shows why the savings grow with session length.

List prices checked 3 October 2026. Cache rules and prices change; the maker's page is the authority.

Sources: Anthropic pricing · Anthropic prompt caching · DeepSeek pricing · OpenAI pricing · Kimi pricing · Alibaba Model Studio pricing · Z.ai pricing · MiniMax pricing · xAI models · Gemini API pricing · Xiaomi MiMo

Was this page helpful?

Still stuck? Open a support ticket

Use the coding models you already know, through one API

One key for OpenAI- and Anthropic-compatible tools. Pay in BDT or USD, and keep the coding agent you already use.

Create an account