Skip to content
Model Guides8 min read

Cheap coding models that hold up for agent work (October 2026)

Tokens Team
Engineering
3 Oct 2026
On this page

Most of what a coding agent does is routine work: reading files, running tests, making small edits, writing commit messages. A cheap coding model can handle that work if it calls tools reliably and copes with long context. This guide covers eight low-cost models worth considering in October 2026, with their official list prices and the catches in how each one is priced.

There are no benchmark scores here. We haven't run a controlled comparison, and the makers' own numbers don't tell you how a model behaves in your repository. What follows is the information you can check (price, limits, tool and vision support) plus our reasoning about which jobs each model suits.

python
# 2M tokens in + 150K out, USD, no cache
ds_flash = 2 * 0.30 + 0.15 * 1.20   # 0.78
glm_flash = 2 * 0.15 + 0.15 * 0.50  # 0.375
luna = 2 * 0.10 + 0.15 * 0.50       # 0.275
mimo = 2 * 0.14 + 0.15 * 0.28       # 0.322

What a cheap model needs for agentic coding#

Price per token is only part of it. A model is cheap for agent work if it finishes the task without you paying for it twice. Check these:

  • Tool calling. Every model below supports function calling. One of them has a restriction (GPT-6 Luna, covered below).
  • Context window. Seven of the eight accept around a million tokens. Step 3.5 Flash stops at 256K, which matters if your agent loads large parts of a repository.
  • Cached-input price. Agents send the same system prompt, tool definitions and file contents on every turn. A low cache-read price often matters more than the headline input price. See prompt caching explained.
  • Output price, and reasoning that's always on. Agents read far more than they write, so output is usually the smaller share of the bill. On models that always think (GLM-5.3 Flash, for example), reasoning is generated text and counts towards it.
  • Pricing quirks. Time-of-day rates, tiers by context length and introductory prices all change what you actually pay.

Price comparison (list prices checked 3 October 2026)#

All prices are in USD per million tokens, from each maker's own pricing page.

ModelMakerContextMax outputInputCached inputOutputVision
DeepSeek V4.1 FlashDeepSeek1M384K$0.30 peak / $0.15 off-peak$0.006 / $0.003$1.20 / $0.60Yes
GLM-5.3 FlashZ.ai1M128K$0.15$0.03$0.50Yes
Qwen 3.8 FlashAlibaba1M131,072$0.15$0.016$0.47Yes
MiMo V2.6 FlashXiaomi1M128K$0.14not listed$0.28Yes
MiniMax M3MiniMax1M128K recommended$0.30 (up to 512K input)$0.06$1.20Yes
Step 3.5 FlashStepFun256Knot listed$0.10not listed$0.30No
GPT-6 LunaOpenAI1,050,000128K$0.10 (up to 272K input)$0.01$0.50Yes
Gemini 3.8 FlashGoogle1,048,57665,536$0.75 until 31 Dec 2026$0.075$3.75Yes

The code block at the top shows what one fairly heavy agent session costs on four of these models if nothing is cached: 2 million input tokens and 150,000 output tokens. The session size is an illustration, not a measurement. Measure your own sessions before you trust any total; estimate your monthly token bill shows how.

When each cheap coding model fits#

DeepSeek V4.1 Flash: a good default, if you watch the clock#

DeepSeek prices by time of day. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, excluding Chinese public holidays. Off-peak rates are half the peak rates. In Dhaka (UTC+6), peak is 07:00 to 10:00 and 12:00 to 16:00 on weekdays, which covers much of a normal working day. Evenings and weekends are off-peak.

The cache-hit price is what makes it good for agents. At $0.006 per million tokens at peak, a cached token costs 2% of a fresh one, so re-reading the same context on every turn adds almost nothing. It has 1M context, a 384K output limit, thinking and non-thinking modes, and native image input. DeepSeek says it beats its own V4 Pro on "performance, cost, speed & total runtime". The weights are open.

The older names deepseek-v4-flash and deepseek-v4-flash-vision-exp are now routed to V4.1 Flash and billed at V4.1 Flash prices, so there's no reason to keep using them.

Whether a reseller passes on the off-peak discount is up to them. Many charge one flat rate. Check the model's entry in the catalog.

GLM-5.3 Flash: cheap, multimodal, always reasoning#

Z.ai calls it "stronger intelligence than GLM-5.2 at an exceptionally low cost". It takes text, images, video and files as input, and thinking can't be turned off. That's useful for planning steps, but it adds output tokens to trivial requests. Cached input is $0.03. There's also a faster tier, GLM-5.3 FlashX, at $0.37 / $1.25, which Z.ai describes as running at about 200 tokens per second.

Qwen 3.8 Flash: flat price across the whole window#

Alibaba positions it as a cheap, fast model for coding and agents. It is a mixture-of-experts model with about 6B active parameters and open weights. Unlike Qwen 3.7 Flash, which is priced by context tier, 3.8 Flash charges the same $0.15 / $0.47 at any prompt size up to 1M, so a very long prompt doesn't jump to a higher tier. Cache hits are $0.016 whether they come from implicit or explicit caching.

MiMo V2.6 Flash: lowest output price on the list#

At $0.28 per million output tokens, it's the cheapest model here for output-heavy work such as generating tests or writing docs. Xiaomi released it on 22 September 2026 with tools, vision, 1M context and MIT-licensed open weights. It's also the newest model on this list, so it has the least history in real-world agent setups. That's a reason to try it on low-risk tasks first, not a reason to avoid it.

MiniMax M3: the expensive end of cheap#

MiniMax describes M3 as a "frontier multimodal coding model with 1M context window", running at around 100+ tokens per second, with tool use interleaved with thinking. Two pricing details matter:

  • Requests with more than 512K input tokens cost twice as much: $0.60 / $0.12 / $2.40.
  • MiniMax only guarantees 512K of context. The 1M window isn't promised on every request.

If your sessions stay under 512K, it's priced the same as DeepSeek V4.1 Flash at peak.

Step 3.5 Flash: lowest input price, smallest window#

At $0.10 in and $0.30 out, it's the cheapest model here on input. StepFun quotes speeds of 100 to 350 tokens per second. It's text-only, has a 256K context window and dates from February 2026. It fits subagent jobs that don't need much context: searching code, summarizing a diff, drafting a commit message.

GPT-6 Luna: cheap, with a tool-calling rule#

OpenAI calls it "our most efficient model for focused, high-volume tasks". Up to 272K input tokens it costs $0.10 / $0.01 cached / $0.50. Above that, the request costs $0.20 / $0.02 / $0.75.

The catch for agents is tool calling. On Chat Completions, tools only work with reasoning_effort set to none. Full tool support is on the Responses API. Tokens serves POST /v1/responses, and the Codex CLI setup uses wire_api = "responses" for this reason. If your agent only speaks Chat Completions, test tool calls before relying on Luna.

Gemini 3.8 Flash: priced for now, budget for January#

Google calls it its "most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents". The current $0.75 / $3.75 is an introductory price that runs through 31 December 2026. From 1 January 2027 it becomes $1.50 / $7.50. Its output limit of 65,536 tokens is the lowest of the eight. If you build a workflow around it, budget at the 2027 price.

Set a cheap default model in your agent#

The quickest way to switch a coding agent's default to a cheap model through Tokens is the CLI. It writes the base URL, key and model into the configs of the agents it finds (OpenCode, Claude Code, Codex CLI, Crush) and backs up each file first:

bash
node tokens.mjs setup --base-url https://tokens.bd --model deepseek/deepseek-v4.1-flash

To use a different model from this list, check its exact ID in the model catalog or with GET /v1/models, because aliases are in provider/model form and the list your key can call depends on your plan. The full flag reference is in Tokens CLI, and Choosing a model covers how to decide.

A cheap default model works best when you also have a stronger model for the turns where it gets stuck. Frontier coding models, October 2026 covers the stronger end and a simple "start cheap, then escalate" pattern.

List prices checked 3 October 2026. They change, sometimes weekly; the maker's page is the authority.

Sources: DeepSeek pricing · DeepSeek V4.1 news · Z.ai pricing · GLM-5.3 Flash · Alibaba Model Studio pricing · Qwen 3.8 Flash · Xiaomi MiMo · MiniMax pricing · StepFun pricing · OpenAI pricing · Gemini API pricing

Models mentioned

Was this page helpful?

Still stuck? Open a support ticket

Use the coding models you already know, through one API

One key for OpenAI- and Anthropic-compatible tools. Pay in BDT or USD, and keep the coding agent you already use.

Create an account