TL;DR: The cheapest capable coding model in October 2026 is Qwen 3.7 Flash at $0.03/1M input and $0.14/1M output; GPT-6 Luna at $0.10/$0.50 is the cheapest big-lab model, and MiniMax M3 at $0.30/$1.20 is the cheapest model with a published SWE-bench Verified score above 80% (80.5%). The same heavy agent session (2M input + 150K output tokens) costs $0.08 on the cheapest model and $9.08 on the priciest — a 112x spread. But token price isn't task price: Claude Haiku 5.5 matches Luna's $0.10/$0.50 list price while generating roughly 3x more output tokens per reasoning task. Prices below are from the Tokens catalog and official list prices, checked 9 October 2026; BDT figures at ৳123.3 per USD.
A coding agent's token bill is the easiest infrastructure cost to cut, because most of the work doesn't need a frontier model. The spread between the cheapest and priciest coding-capable models is now wider than at any point this year — and October's price war (Anthropic's Haiku 5.5 launch, OpenAI's Luna, DeepSeek's time-of-day pricing) made the price list more complicated than "pick the lowest number."
This post is the number side of the story: a per-1M-token comparison across the coding models you can actually call, the catches hidden in the price lists, and what a real agent session costs on each. If you want the qualitative side — which cheap models hold up in agentic setups — read our companion guide Cheap Coding Models for AI Agents: October 2026 Prices.
How this comparison is built#
Two sources, both checked 9 October 2026:
- The Tokens model catalog — the prices you actually pay through one API key, one base URL. Gateway prices track the makers' list prices closely.
- Official list prices from each maker's pricing page, via reporting from 7–9 October 2026 (Haiku 5.5 launch coverage, provider docs).
All figures are USD per million tokens: input / output. Cached-input pricing is noted where it changes the math, because agents re-send the same context on every turn — see prompt caching explained. Output tokens cost 3–10x more than input at every provider, but agents are input-heavy (they read far more than they write), so the input price usually dominates the bill.
Prices change constantly — OpenAI moved three times in two months this summer. Re-verify before you budget.
The full table: coding-capable models, cheapest first#
Prices from the Tokens catalog, checked 9 October 2026. Every model here is on the catalog with function/tool calling support.
| Model | Input / 1M | Output / 1M | Context | Notes |
|---|---|---|---|---|
| Qwen 3.7 Flash | $0.03 | $0.14 | 1M | Cheapest on the catalog; tiered pricing by context |
| Step 3.5 Flash | $0.10 | $0.33 | 256K | Text-only; small window, good for subagent jobs |
| GPT-6 Luna | $0.11 | $0.55 | 1.1M | Cheapest big-lab model; tier jumps above 272K input |
| Muse Spark 1.3 Contributor | $0.11 | $0.22 | 1M | Lowest output price near this input tier |
| MiMo V2.5 | $0.15 | $0.31 | 1M | Open weights (MIT) |
| MiMo V2.6 Flash | $0.15 | $0.31 | 1M | Same price as V2.5 |
| DeepSeek V4.1 Flash | $0.17 | $0.66 | 1M | Maker list: $0.30/$1.20 peak, half off-peak |
| GLM-5.3 Flash | $0.17 | $0.55 | 1M | Maker list $0.15/$0.50; reasoning always on |
| Qwen 3.8 Flash | $0.18 | $0.52 | 1M | Maker list $0.15/$0.47; flat price across 1M |
| MiniMax M3 | $0.33 | $1.32 | 1M | 80.5% SWE-bench Verified; 2x above 512K input |
| DeepSeek V4 Pro | $0.73 | $2.18 | 1M | Maker peak pricing by time of day |
| GLM-5.3 | $1.54 | $4.84 | 1M | Maker list $1.40/$4.40; reasoning always on |
| Gemini 3.8 Flash | $1.65 | $8.25 | 1M | Maker intro price $0.75/$3.75 through 31 Dec 2026 |
| Claude Sonnet 5.5 | $2.20 | $11.00 | 1M | Maker list $2.00/$10.00; cache reads halved to $0.10 |
| Qwen 3.7 Max | $2.75 | $8.25 | 1M | Heavier Qwen tier |
| Kimi K3 | $3.30 | $16.50 | 1M | Maker list $3.00/$15.00; thinking defaults to max |
| GPT-5.6 Sol | $5.50 | $33.00 | 1.1M | Maker list $5.00/$30.00 |
Also on the catalog at $0.00 / $0.00: Ling 3.0 Flash Sante, Ling 3.1 Flash, and Laguna S 2.1. Free is free, but we haven't verified their coding quality — try them on low-risk tasks before trusting them with real work.

Cost per task, not cost per token#
The table above answers "what does a token cost." Your bill answers "what does a finished task cost," and those diverge in two ways:
1. Token efficiency. Claude Haiku 5.5 (launched 7 October 2026) matches GPT-6 Luna's list price exactly — $0.10/$0.50 for prompts up to 100K tokens. But Artificial Analysis testing found Haiku 5.5 used roughly 162,000 output tokens per task at maximum effort versus about 54,000 for Luna — a 3x efficiency gap. Haiku's new tokenizer also generates roughly 30% more tokens per character than its predecessor's. Anthropic's own "about 75% cheaper than Haiku 4.5 on average" claim (not the 90% headline) already bakes in the tokenizer effect. Two models at the same list price can bill very differently for the same job.
2. Capability per dollar. The cheapest token is expensive if the task fails and retries. The models with published SWE-bench Verified scores above 80% tell the value story:
| Model | SWE-bench Verified | Input / 1M (list) | Output / 1M (list) |
|---|---|---|---|
| Claude Fable 5 | 95.0% | $10.00 | $50.00 |
| Claude Sonnet 5 | 85.2% | $2.00 | $10.00 |
| DeepSeek V4 Pro Max | 80.6% | open weights | — |
| MiniMax M3 | 80.5% | $0.30 | $1.20 |
MiniMax M3 is the cheapest model with a published score above 80% — roughly 15x cheaper per input token than Sonnet 5 at ~5 points of benchmark. That doesn't mean it solves your tasks 15x cheaper (benchmarks aren't your repo), but it's where cost-per-task measurement should start. On Terminal-Bench 4.0: Claude Opus 5.5 scores 66.4%, Haiku 5.5 scores 39.2%.
The honest rule: measure tokens per completed task on your own workload, then multiply by the table. Everything else is a proxy.
The fine print: five pricing catches#
The cheapest headline price is rarely the price you pay. October's price lists hide five mechanisms:
1. The Haiku 5.5 hundred-thousand cliff. Above 100K prompt tokens, Haiku 5.5 jumps 5x: $0.50 input / $2.50 output. Between 100K and 272K prompt tokens, GPT-6 Luna still charges $0.10 input — five times cheaper on input at list price. If your P90 request crosses 100K, budget at the upper tier. (Chunking work and caching the big context are the fixes — see the Haiku 5.5 pricing breakdown.)
2. Luna's 272K step. Above 272K input tokens, GPT-6 Luna costs 2x input / 1.5x output. Still cheap, but the step exists, and agent contexts grow.
3. DeepSeek's clock. DeepSeek prices by time of day: peak is 01:00–04:00 and 06:00–10:00 UTC on weekdays — 07:00–10:00 and 12:00–16:00 in Dhaka (UTC+6) — with off-peak at half price. Schedule heavy jobs for evenings and weekends and the same model costs half. Check the catalog entry, because not every reseller passes the off-peak discount through.
4. Introductory prices with expiry dates. Gemini 3.8 Flash's $0.75/$3.75 runs only through 31 December 2026; from 1 January 2027 it becomes $1.50/$7.50. Budget at the 2027 price if you're building a durable workflow.
5. Always-on reasoning and effort defaults. GLM-5.3 and GLM-5.3 Flash can't turn thinking off — planning tokens bill as output on every request, including trivial ones. Kimi K3 defaults to max thinking effort, so routine turns on a $15/1M-output model run at maximum reasoning unless you lower it. See what reasoning effort levels actually cost for the arithmetic.
One piece of good news: Sonnet 5.5 cache reads were halved ($0.20 → $0.10/1M) alongside the Haiku launch — Anthropic says that makes Sonnet 5.5 about 20% cheaper on typical agentic work, where the same context is re-read every turn.
Worked example: one heavy agent session, six models#
Same session on every model — 2M input tokens, 150K output tokens, nothing cached — using Tokens catalog prices. This is an illustration of the arithmetic, not a measurement; how to estimate your monthly token bill shows how to measure your own.
| Model | Input: 2M | Output: 0.15M | Session total | In BDT |
|---|---|---|---|---|
| Qwen 3.7 Flash | $0.06 | $0.02 | $0.08 | ৳10 |
| Step 3.5 Flash | $0.20 | $0.05 | $0.25 | ৳31 |
| GPT-6 Luna | $0.22 | $0.08 | $0.30 | ৳37 |
| MiniMax M3 | $0.66 | $0.20 | $0.86 | ৳106 |
| Claude Sonnet 5.5 | $4.40 | $1.65 | $6.05 | ৳746 |
| Kimi K3 | $6.60 | $2.48 | $9.08 | ৳1,119 |
$0.08 to $9.08 for the same token volume — a 112x spread. And because Tokens bills in BDT with local payment options (no international card needed — see paying for AI APIs from Bangladesh), the cheapest session above costs less than a cup of cha.
Caching compresses the spread further: cached input on DeepSeek V4.1 Flash is $0.006/1M at list peak (2% of fresh input), and Haiku 5.5 cache reads are a penny per million. The models with the biggest input/output asymmetry reward caching the most.
How to actually spend less#
The table gets you the cheapest token. These get you the cheapest bill:
- Route by task, not by habit. Default to a cheap model (Luna, Qwen 3.8 Flash, MiniMax M3) and escalate to a frontier model only when an objective check fails. The frontier models guide walks through the pattern.
- Cache aggressively. System prompts, tool definitions, and file contents repeat every turn — caching cuts 50–90% off repeated-context costs. That's usually the single biggest lever, bigger than model choice.
- Cap the downside. One buggy loop on Kimi K3 at max effort can burn a budget in an hour. Put thinking-heavy workloads behind per-key spend caps and usage alerts.
- Schedule around the clock. If you're on DeepSeek pricing, run heavy jobs off-peak — evenings and weekends in Dhaka.
- One key, every model. The whole point of the catalog is that switching models is a one-word change, not a new contract. Try the cheap end this week; the experiment costs dollars, not days.
Cheapest coding LLM APIs FAQ#
What is the cheapest LLM API for coding in October 2026? By raw token price, Qwen 3.7 Flash at $0.03/1M input and $0.14/1M output on the Tokens catalog. Among big-lab models, GPT-6 Luna at $0.10/$0.50. Among models with a published SWE-bench Verified score above 80%, MiniMax M3 at $0.30/$1.20 (80.5%). The catalog also lists free tiers (Ling 3.x Flash, Laguna S 2.1) whose coding quality we haven't verified.
Is Claude Haiku 5.5 really 90% cheaper than Haiku 4.5? The 90% figure is the list-price cut for prompts up to 100K tokens ($1.00→$0.10 input, $5.00→$0.50 output). Anthropic's own "about 75% cheaper on average" is closer to real bills — it accounts for workload mix and a new tokenizer that uses roughly 30% more tokens per character. Above 100K prompt tokens, rates jump 5x.
Why do GPT-6 Luna and Claude Haiku 5.5 cost different amounts if both are $0.10/$0.50? Because list price isn't task price. Haiku 5.5 generates roughly 3x more output tokens per reasoning task than Luna (per Artificial Analysis), its >100K prompt tier is 5x the cheap tier, and Luna's tool-calling rules differ (full tool support needs the Responses API). Measure tokens per completed task on your workload.
How do I pay for these APIs in BDT? Tokens bills in Bangladeshi Taka with local payment methods — no international card required. The catalog prices above are in USD; at ৳123.3 per USD (9 October 2026), the cheapest heavy agent session in the worked example costs about ৳10. See paying for AI APIs from Bangladesh.
Prices checked 9 October 2026: Tokens catalog at tokens.bd/models plus official list prices from Anthropic, OpenAI, Google, DeepSeek, Z.ai, Alibaba, MiniMax, StepFun, Xiaomi and Moonshot via 7–9 October 2026 reporting. List prices change frequently — the maker's page is the authority.
Sources: Tokens model catalog · Anthropic Haiku 5.5 pricing via Unite.AI · Haiku 5.5 price analysis via AIWeekly · Haiku 5.5 review via AgentBreaking · LLM API price comparison via Morph · USD/BDT rate via Finnhub



