Skip to content

Prompt caching

How prompt caching works through Tokens: what the provider does, what the gateway passes through, how cache reads and writes show up in usage, and how they are priced and billed.

On this page

Prompt caching lets a provider reuse the work it already did on the start of your prompt. A request that repeats a long system prompt, tool list or conversation history is billed at a lower rate for the repeated part. For a coding agent, which resends almost the same context on every turn, it is the largest cost lever there is.

The cache belongs to the upstream provider. Tokens does not run a cache of its own. What Tokens does is forward your request, read the cache counts the provider reports, and bill those tokens at the cache prices set for the model. This page separates the two: what is the provider's rule, and what Tokens does.

Who does what#

QuestionAnswer
Who decides whether a request hits the cache?The provider that serves it. Tokens cannot force a hit.
Who sets cache lifetime and minimum prompt size?The provider, per model.
Does Tokens change cache markers in my request?Not when the request goes to the provider in its own format. When the gateway translates between formats, cache markers are dropped.
Does Tokens change the usage numbers?No. The usage object in the response is the provider's. When the gateway translates formats it maps the cache fields (see below).
Who prices cache reads and writes?Tokens, per model. A model with no cache price set is billed at its normal input price.
Where do I see what I paid?Usage shows cached tokens per request and the cost of the request.

Two kinds of caching#

Automatic caching. You change nothing. The provider detects that the start of your prompt matches an earlier request and reuses it. OpenAI's models work this way, and so does DeepSeek's. The provider's documentation says the cache works on a prefix: a request hits only if its beginning is identical to an earlier one. DeepSeek describes its cache as best effort and does not promise a hit. Prompts shorter than the provider's minimum are not cached; OpenAI lists 1,024 tokens for its newest models.

Explicit caching with cache_control. On Anthropic's Messages API you mark the content you want cached with a cache_control block. Anthropic's rules, checked October 2026 in its prompt caching documentation:

  • A cache breakpoint covers everything before it, in this order: tools, system, then messages.
  • You can set up to four breakpoints per request. A top-level cache_control field places one automatically on the last cacheable block.
  • The default lifetime is 5 minutes, refreshed each time the cached content is used. "ttl": "1h" asks for one hour.
  • There is a minimum prompt length that depends on the model (512 to 4,096 tokens in Anthropic's table). A shorter prompt is processed normally, without an error and without caching.
  • A cache entry becomes available only after the first response begins, so parallel requests sent at the same moment do not share it.
  • Any change at one level, for example editing a tool definition, invalidates that level and everything after it.

These are Anthropic's rules for its own models. Other providers that accept cache_control may differ. Check the provider's documentation for the model you use.

Using cache_control through /v1/messages#

Send cache_control exactly as you would to Anthropic. When the model is served by a provider that speaks the Messages API, the request body goes to it as written.

python
import os
import anthropic

client = anthropic.Anthropic(
    base_url="https://tokens.bd",
    api_key=os.environ["TOKENS_API_KEY"],
)

# Must be longer than the model's minimum cacheable length.
project_notes = open("project-notes.txt", encoding="utf-8").read()


def ask(question: str) -> None:
    message = client.messages.create(
        model="deepseek/deepseek-v4.1-flash",
        max_tokens=300,
        system=[
            {
                "type": "text",
                "text": project_notes,
                "cache_control": {"type": "ephemeral"},
            }
        ],
        messages=[{"role": "user", "content": question}],
    )
    usage = message.usage
    print(
        "input:", usage.input_tokens,
        "| cache write:", usage.cache_creation_input_tokens or 0,
        "| cache read:", usage.cache_read_input_tokens or 0,
        "| output:", usage.output_tokens,
    )


ask("Summarize the notes in one sentence.")  # first call: expect a cache write
ask("List three risks mentioned in the notes.")  # within the TTL: expect a cache read

If the second call shows a cache read, the markers reached a provider that honors them. If both calls show zero for cache write and cache read, one of these is true: the prompt is under the model's minimum, the serving provider does not support explicit caching, or the request was translated (next section). Testing with two identical calls is the only reliable check for a given model.

What is passed through and what is dropped#

Each model is served by one or more upstream providers, and each provider speaks the OpenAI format, the Anthropic format, or both. Tokens sends your request in your own format when the provider supports it, and translates when it does not.

You callProvider speaksWhat happens to caching
/v1/messagesAnthropic formatBody forwarded as is. cache_control reaches the provider. Usage comes back in Anthropic fields.
/v1/messagesOpenAI format onlyTranslated to chat completions. cache_control markers are dropped. Automatic caching at the provider can still apply.
/v1/chat/completionsOpenAI formatBody forwarded as is. Automatic caching applies. Fields such as prompt_cache_key go to the provider if you send them.
/v1/chat/completionsAnthropic format onlyTranslated to Messages. Cache markers are not added or carried. The Anthropic provider's usage is mapped back to OpenAI fields.

You cannot choose which provider serves a request. Providers are tried in priority order, and a request moves to the next one only when the first fails or does not answer in time. The next provider has not seen your prefix, so expect a miss on that request.

Cache fields in the response#

The gateway reads these fields from the provider's usage object. Everything else in the response is passed through.

Format and fieldWhat it meansHow Tokens bills it
Chat completions: prompt_tokensAll input tokens, cached ones includedMinus cached, as input
Chat completions: prompt_tokens_details.cached_tokensInput tokens served from cacheCache read
Chat completions: cache_creation_input_tokens (top level)Present when the gateway translated an Anthropic answer: input tokens written to cacheCache write
Messages: input_tokensInput tokens that were neither read from nor written to cacheInput
Messages: cache_read_input_tokensInput tokens served from cacheCache read
Messages: cache_creation_input_tokensInput tokens written to cacheCache write

Two consequences:

  • The two formats count differently. In chat completions prompt_tokens already contains the cached tokens. In Messages, input_tokens does not. Tokens converts both into the same four buckets before pricing, so you do not need to adjust for the difference.
  • A provider that reports caching only in a field Tokens does not read, for example a custom hit counter, has its cached tokens billed as normal input. If the numbers on the usage page show no cached tokens for a model that you know caches, tell support with a request id.

On OpenAI-format streams the usage chunk only reaches you if you set stream_options.include_usage. Billing does not depend on it: the gateway meters the stream either way. See streaming.

How cache tokens are billed#

Every request is split into four buckets, and each has its own price per million tokens:

text
cost = input        x input price
     + cache read   x cache-read price
     + cache write  x cache-write price
     + output       x output price

The prices come from the model's entry in the Tokens catalog. Where a model has no cache price, that bucket is billed at the model's input price, so caching saves nothing for that model. A cache price of zero means the bucket is free. Plan discounts, if your plan has one, apply to the total.

An example with made-up prices: input $2.00, cache read $0.20, output $8.00 per million tokens. A request has 10,000 input tokens, of which 9,880 come from the cache, and produces 300 output tokens.

BucketTokensPrice per millionCost
Input120$2.00$0.000240
Cache read9,880$0.20$0.001976
Output300$8.00$0.002400
Total$0.004616

The same request with no cache hit costs 10,000 x $2.00 per million plus the same output, $0.0224. Real prices for each model are on the model catalog and pricing, and the usage page shows what each request was actually charged.

Providers that charge extra for writing a cache entry (Anthropic charges 1.25 times the input price for a 5-minute write and 2 times for a 1-hour write, on its own API) pass that on only if the Tokens catalog has a cache-write price for the model. Where it does not, writes are billed as input.

Cached requests are still checked against your balance before they run. The gateway reserves a worst-case amount using your max_tokens and an estimate of the input priced at the normal input rate, because it cannot know in advance that the cache will hit. A request with a very large cached prefix can therefore be refused for a low balance even though its final cost would be small. Settlement then charges the real amount. More in chat completions.

See cache hits on the dashboard#

The activity table on Usage shows tokens per request as input · N cached · output. The cached figure is cache reads and cache writes added together. The row's cost is the full charge for the request. A request that was estimated, because the provider sent no usage, carries an Estimated badge.

To see the split between reads and writes for one call, read the usage object in the response itself and log it.

Get more cache hits#

These follow from the providers' rules above.

  • Put stable content first. System prompt, tool definitions and long reference text go at the start, unchanged between requests. Variable content, such as the new user message, goes last.
  • Keep the prefix byte-identical. A timestamp, request id or random value near the top of the system prompt changes every request and defeats the cache for everything after it. Keep tool definitions in a fixed order.
  • Append to history, do not rewrite it. Editing or compacting earlier turns creates a new prefix.
  • Keep the same model. Caches are per model.
  • Stay inside the lifetime. With a 5-minute lifetime, a pause longer than that means the next request pays for a write again. For slow interactive use on Anthropic models, "ttl": "1h" can cost less overall, because a 1-hour write costs more than a 5-minute one.
  • Send the first request alone. Wait for the first response to begin before firing parallel requests that share the prefix.
  • Make the prompt long enough. Short prompts are below the provider's minimum and never cache.

Troubleshooting#

No cached tokens ever. Run the two-call test above. Check, in order: the prompt length against the model's minimum, whether anything at the top of the prompt changes between calls, whether the calls are more than a few minutes apart, and whether you are using /v1/messages with a model that is only served in OpenAI format (markers dropped). If the test shows zero reads on a model you expect to cache, send the two request ids to support.

Cached tokens appear, but the request cost about the same. The model probably has no cache-read price in the catalog, so cached tokens are billed at the input price. Compare the cost on the usage page against the model's prices.

Cache write tokens higher than expected. Each change near the start of the prompt writes a new entry. Look for content that varies per request, and for tool lists that change between turns.

cache_control returns a 400. The provider rejected the request body. Anthropic returns an error if a request combines automatic caching with four explicit breakpoints. Check the provider's limits, then see errors.

A coding agent costs more than you expected. The agent decides what goes in the prompt and where. When it compacts or rewrites earlier turns, the prefix changes and the next request misses the cache. Check your agent's settings for compaction, and see its page under coding agents.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.