TL;DR: Two levers, one bill. On a 50-turn coding-agent session with a 100K-token stable context, prompt caching on Claude Sonnet 5.5 cuts input cost 93% — from $10.00 to $0.74 — with zero quality loss. Downgrading to Claude Haiku 5.5 costs $0.50 uncached and $0.06 with caching, but you trade frontier-model quality for it. Verdict: if the smaller model's quality is good enough for the task, downgrading wins on raw cost; if you need Sonnet's capability, caching is the best lever you have; combining both is cheapest of all.
Every AI coding agent has the same expensive habit: it re-sends its entire context on every turn — system prompt, tool schemas, conversation history — all billed as fresh input tokens. A long session's repeated input is often 70–90% of the bill. Two levers attack it, and developers argue endlessly about which to pull first.
Lever 1: prompt caching. Keep the big model; make repeated tokens cheap (90–98% off). Lever 2: model downgrade. Switch to a smaller, cheaper model; make every token cheap.
This post runs the same workload through both levers with current list prices, so you can pick by math instead of vibes. If you want the deep mechanics of caching first, the earlier explainer Prompt Caching Explained covers the prefix rule, break-even math, and the five cache killers — this one is the head-to-head.
Side A: prompt caching — what it saves#
Caching works on a strict rule: when a request's prefix is byte-identical to one the provider recently processed, the stored computation is reused and the repeated tokens are discounted 90–98%. The figures below were checked 10 October 2026.

| Provider | Cache read discount | Write premium | Setup |
|---|---|---|---|
| Anthropic (Sonnet 5.5) | $0.10/1M — 95% off (halved from $0.20 on 7 Oct 2026) | 1.25x for 5-min TTL | cache_control breakpoints |
| OpenAI (GPT-5.4/5.5) | 90% off | None | Fully automatic |
| DeepSeek V4.1 Flash | $0.003/1M — 98% off | None | Fully automatic |
| Google Gemini | ~90% off implicit, or explicit cached content | Storage rent for explicit | Automatic or managed |
Key economics: on Anthropic, the write premium means a prefix must be reused at least twice within the TTL — a never-reused prefix costs 25% more than not caching. On OpenAI and DeepSeek there's no write tax, so every hit is pure saving. The headline number for the head-to-head: Sonnet 5.5, 50 turns × 100K stable tokens — $10.00 uncached → $0.74 cached (one 5-min write at $2.50/1M = $0.25, plus 49 reads at $0.10/1M = $0.49). Same model, same quality, 93% cheaper.
Side B: model downgrade — what it saves#
Downgrading attacks the base price instead of discounting the volume. Maker list prices, checked 10 October 2026:
| Model | Input / 1M | Output / 1M | Notes |
|---|---|---|---|
| Claude Sonnet 5.5 | $2.00 | $10.00 | Frontier coding quality |
| Claude Haiku 5.5 | $0.10 | $0.50 | 95% cheaper input; smaller model |
| GPT-6 Luna | $0.10 | $0.50 | Budget OpenAI option |
| DeepSeek V4.1 Flash | $0.15 (off-peak) | $0.60 | Cache hits at $0.003/1M |
The same 50-turn session (5M input tokens) on Haiku 5.5 with no caching: 5M × $0.10/1M = $0.50 — 95% cheaper than uncached Sonnet 5.5, and even cheaper than cached Sonnet 5.5's $0.74.
The catch is quality, not price. A smaller model may need more iterations, more debugging passes, or a fallback escalation on hard tasks — and every extra turn re-bills the whole context. Downgrade wins the spreadsheet; caching wins the risk profile. For a per-model price map across the whole market, see Cheapest Coding LLM APIs 2026.
Head-to-head: one session, four ways#
Same workload: a coding agent re-sending a 100K-token stable context every turn, 50 turns, input cost only, maker list prices checked 10 October 2026.
| Strategy | Math | Input cost | vs. uncached Sonnet |
|---|---|---|---|
| Sonnet 5.5, no caching | 5M × $2.00 | $10.00 | baseline |
| Sonnet 5.5 + prompt caching | $0.25 write + $0.49 reads | $0.74 | −93% |
| Haiku 5.5, no caching | 5M × $0.10 | $0.50 | −95% |
| Haiku 5.5 + prompt caching | $0.0125 write + $0.049 reads | $0.06 | −99.4% |

Two surprises in the table. First, raw downgrade ($0.50) beats caching the big model ($0.74) on input cost alone — the 95% list-price cut outweighs the 93% cache discount. Second, the combination ($0.06) is another order of magnitude down: downgrade multiplies the base rate, caching multiplies the volume discount, and they stack.
When each lever wins#
Choose prompt caching when:
- You need the frontier model's reasoning — complex refactors, unfamiliar codebases, agentic debugging where quality compounds across turns.
- Output quality is the product (customer-facing code, reviewed PRs) and rework is expensive.
- Your context is stable and repetitive: system prompts, tool schemas, few-shot examples, long documents referenced across turns.
Choose model downgrade when:
- The task fits the smaller model's quality bar — boilerplate, tests, summaries, simple CRUD, documentation.
- Output tokens dominate your bill: Haiku-class output is 20x cheaper, and caching discounts input only.
- You can route by difficulty: cheap model for easy turns, escalate only when stuck (though note that switching models resets the cache — the new model pays full input price).
Choose both when: the workload is high-volume and quality-tolerant — batch processing, evals, or background agents. Haiku 5.5 cached at ~$0.06 per 50-turn session is roughly 167x cheaper than the naive setup.
Implementing each lever#
Caching on Anthropic is a few lines — mark the stable blocks:
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=1024,
system=[
{"type": "text", "text": SYSTEM_PROMPT,
"cache_control": {"type": "ephemeral"}}, # 5-min cache
],
tools=[{**tool, "cache_control": {"type": "ephemeral"}} for tool in tools],
messages=messages,
)Downgrading is often just a model-string swap — and through the Tokens gateway, the same code runs against 62 models (catalog checked 10 October 2026: Sonnet 5.5 $2.20/$11.00, DeepSeek V4.1 Flash $0.17/$0.66, GPT-6 Luna $0.11/$0.55) with one base_url change, as shown in Use Any LLM With the OpenAI SDK. That makes A/B-testing the downgrade a config change, not a rewrite — and provider usage fields pass through, so you can verify cache hit rates on the same dashboard where the BDT billing lands. Gateway-vs-direct trade-offs are compared in Tokens vs OpenRouter vs direct APIs.
Whichever lever you pull, give the workload its own API key with a monthly spend cap and usage alerts — a runaway loop should warn you before it bills you.
Prompt Caching vs Downgrade FAQ#
Doesn't caching just add complexity a cheaper model avoids?
On OpenAI, DeepSeek, and Gemini's implicit caching, there is no complexity at all — caching is automatic, and your only job is keeping the prefix stable. On Anthropic it's a few lines of cache_control. The complexity argument mostly doesn't survive contact with the savings: 93% off the same model for near-zero effort.
If Haiku 5.5 is $0.50 uncached vs $0.74 cached Sonnet, why not always downgrade? Because input cost isn't the whole bill. If the smaller model produces a worse fix that needs two more debugging turns, or gets escalated to Sonnet mid-session anyway (resetting the cache), the savings evaporate. Downgrade when you've verified quality on your workload; cache when you haven't.
Can I combine caching and downgrading? Yes — they're independent multipliers. Haiku 5.5 with caching lands around $0.06 for the 50-turn session in the table above, 99.4% off the naive baseline. The one interaction: switching models resets the cache, so pick the model per workload, not per turn.
Does output pricing change the verdict? It strengthens the downgrade case for output-heavy workloads. Sonnet 5.5 output is $10.00/1M vs Haiku 5.5's $0.50/1M — a 20x gap that caching (input-only) can't touch. For long generations, downgrade or a hybrid route usually wins.
List prices checked 10 October 2026 at the makers' list prices and the Tokens model catalog. Provider list prices change; the maker's pricing page is the authority, and the catalog page is the authority for Tokens pricing.
Sources: Anthropic cache-read cut, Unite.AI · Haiku 5.5 pricing, FoneArena · Sonnet 5.5 rate card, RohitAI · OpenAI caching explainer, dev.to · DeepSeek pricing, Cabina.AI · DeepSeek V4.1 Flash pricing, Yotta Labs · Provider caching comparison, Tech Insider · GPT-6 Luna vs DeepSeek vs Haiku pricing, Tech Insider



