Choosing a model is mostly about the job you give it, not a leaderboard. This page groups models by the work they suit, compares a representative set, and explains how model ids work on Tokens. Everything you can buy is in the catalog at /models; the exact list for one key comes from GET /v1/models.
About the prices on this page
Prices below are the providers' own list prices per 1M tokens in USD, checked October 2026. They show how models compare to each other. Your price on Tokens is on /models, and it can differ from these.
Model ids on Tokens#
Every model has an id in provider/model form, for example deepseek/deepseek-v4.1-flash. That id goes in the model field of every request and in your agent's config.
- The id on Tokens is an alias and doesn't always match the provider's own id. DeepSeek's API, for example, calls its current Flash model
deepseek-flash. Always copy ids from /models or from the API. - The catalog shows everything Tokens offers. Your key may see fewer models: plans include different models, and a key can have an allowed-models list.
curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY"Each data[].id in the response is a model this key can call right now.
Compare representative models#
| Model (maker) | Context | Known for | List price in / out |
|---|---|---|---|
| Claude Sonnet 5.5 (Anthropic) | 1M, 128K out | Agentic coding, tool use, image input | $2 / $10 |
| GPT-5.6 Sol (OpenAI) | 1.05M, 128K out | Complex, long-horizon agentic work | $4 / $20 (promotional until at least 2026-11-21; launch price $5 / $30) |
| GPT-6 Luna (OpenAI) | 1.05M, 128K out | Focused, high-volume tasks | $0.10 / $0.50 |
| Gemini 3.8 Flash (Google) | 1M in, 65K out | Long-horizon software engineering; text, image, video, audio and PDF input | $0.75 / $3.75 until 2026-12-31, then $1.50 / $7.50 |
| Grok 4.7 (xAI) | 500K | Coding and knowledge work | $2 / $6 (prompts under 200K) |
| Kimi K3 (Moonshot AI) | 1M | Long-horizon coding and deep reasoning; thinking always on | $3 / $15 |
| GLM-5.3 (Z.ai) | 1M, 128K out | Coding and agents; text only; reasoning always on | $1.40 / $4.40 |
| Qwen 3.8 Max (Alibaba) | 1M, 131K out | Coding, multi-tool orchestration, image and video input | $2 / $6 |
| DeepSeek V4.1 Flash (DeepSeek) | 1M, 384K out | Fast and cheap, image input | $0.30 / $1.20 at peak hours, half off-peak |
| MiniMax M3 (MiniMax) | 1M | Multimodal coding | $0.30 / $1.20 (input up to 512K) |
| Qwen 3.8 Flash (Alibaba) | 1M, 131K out | Cheap, fast coding and agents | $0.15 / $0.47 |
| MiMo V2.6 Flash (Xiaomi) | 1M, 128K out | Low-cost reasoning, vision input | $0.14 / $0.28 |
The "known for" column paraphrases each maker's own positioning. It isn't a benchmark result.
Best model for agentic coding with long sessions#
Agents like Claude Code, Codex CLI and OpenCode call the model dozens of times per task, resending the conversation and file contents each turn. Three things matter more than raw intelligence scores:
- Reliable tool calling. A model that fumbles tool arguments wastes turns.
- Input price, and cached-input price. Long sessions are dominated by input tokens. Claude Sonnet 5.5 lists cached input at $0.20 against $2 uncached; GPT-5.6 Sol lists $0.40 against $4. Caching is done by the upstream provider; where a model has a cache-read price, the catalog shows it.
- Context that holds up. Most current flagships offer around 1M tokens.
Good starting points: Claude Sonnet 5.5, GPT-5.6 Sol, Kimi K3, GLM-5.3, Gemini 3.8 Flash and Qwen 3.8 Max. Kimi K3 and GLM-5.3 always think before answering, which helps on hard problems but adds output tokens to every call.
Tip
Use a strong model for the main agent and a cheap one for side work. Several agents let you set a separate small or fast model for titles, summaries and quick edits. See the guide for your agent, such as Claude Code or OpenCode.
Cheap models for high-volume edits#
For bulk refactors, commit messages, classification, test scaffolding or anything you run thousands of times, price per token decides. From the table: GPT-6 Luna ($0.10 / $0.50), MiMo V2.6 Flash ($0.14 / $0.28), Qwen 3.8 Flash ($0.15 / $0.47) and DeepSeek V4.1 Flash. Z.ai's GLM-5.3 Flash ($0.15 / $0.50) is in the same range and accepts image, video and file input.
One detail on GPT-6 Luna: OpenAI supports its tool calling on Chat Completions only with reasoning_effort set to none; full tool support is on the Responses API. Tokens supports both POST /v1/chat/completions and POST /v1/responses.
Big-context models for repo-wide work#
A 1M-token window is common now, but filling it is not free. At Claude Sonnet 5.5's list price, an 800K-token prompt costs $1.60 in input alone, every time you send it. Some providers also charge more past a threshold:
- GPT-6 Luna and GPT-5.6 models: higher rates above 272K input tokens.
- Grok 4.5 to 4.7: $4 / $12 at 200K prompt tokens and above.
- MiniMax M3: double rates above 512K input.
For most repo questions, a good agent that searches and reads selectively beats pasting the whole tree. When you do need the full window, check the price on /models first. Also note that a request body to Tokens can be at most 10 MB.
Vision and multimodal input#
If your prompts include screenshots, diagrams, PDFs or recordings, check input types before anything else:
- Broadest input: Gemini 3.8 Flash (text, image, video, audio, PDF), Qwen 3.8 Omni Flash (text, image, audio, video), MiMo V2.6 Pro (text, image, video, audio).
- Image input with strong coding: Claude Sonnet 5.5, Qwen 3.8 Max, Kimi K3, MiniMax M3, DeepSeek V4.1 Flash.
- Text only: GLM-5.3, DeepSeek V4 Pro, LongCat 2.0 and Tencent Hy4 Preview. Images sent to these will fail or be ignored.
Images go inside the chat message, in the format your SDK uses. Tokens doesn't offer image generation, audio or file-upload endpoints.
Fast models for interactive use#
Several makers sell a faster-serving version of a model at a higher price. Their speed figures, as published by the makers:
- Kimi K2.7 Code HighSpeed: about 180 tokens/s, at twice the price of K2.7 Code ($1.90 / $8).
- GLM-5.3 FlashX: about 200 tokens/s, $0.37 / $1.25.
- MiMo V2.6 Pro UltraSpeed: "up to 20x faster output" than MiMo V2.6 Pro, $4.35 / $8.70.
- Step 3.5 Flash: 100 to 350 tokens/s, $0.10 / $0.30.
Tokens adds a network hop, so time to first token through Tokens won't beat calling the provider directly. Turn on streaming ("stream": true) so text appears as it is generated; see streaming.
Try two models before you commit#
Pick two candidates and give them the same real task from your codebase. The dashboard Playground is the quickest way to compare single answers. For agent work, point your agent at each model for one session and compare cost per finished task in usage, not price per token. A model with a higher token price that finishes in fewer turns can be the cheaper one.
When you've chosen, set it as the default in your agent: the Tokens CLI asks for a default model during setup, or use the per-agent guides.
Sources (list prices and specs, checked October 2026): Anthropic pricing, OpenAI pricing, Google Gemini pricing, xAI models, DeepSeek pricing, Kimi pricing, Z.ai pricing, Alibaba Model Studio pricing, MiniMax pricing, Xiaomi MiMo, StepFun pricing.