Skip to content
Engineering13 min read

Coding Model Benchmarks: SWE-bench, Terminal-Bench

Md Badsha
Platform Engineer
11 Oct 2026
On this page
Coding Model Benchmarks: SWE-bench, Terminal-Bench

TL;DR: On SWE-bench Verified, the top coding models are statistically tied — independent testing (vals.ai) puts Claude Opus 5 at 97.0%, DeepSeek V4 Pro at 96.4%, GPT-5.6 Sol at 96.2% and GLM-5.3 at 95.4%, all within each other's error bars. The ranking has moved to the harder SWE-bench Pro, where Scale's standardized harness leads at 61.5% (Muse Spark 1.1) — far below vendor-claimed 80%+ runs on custom scaffolds. On the Artificial Analysis Intelligence Index v4.3.2, Claude Opus 5.5 (57.6) leads Claude Sonnet 5.5 (56.0). The value story: DeepSeek V4 Pro delivers ~94.5 Verified points per dollar — about 11x the score-per-dollar of GPT-5.6 Sol — on the 60+ model tokens.bd catalog (prices checked 11 October 2026).

Every coding-model launch post opens with a benchmark table where the new model wins. That is the whole point of the table. This post is the antidote: the same four benchmark suites — SWE-bench Verified, SWE-bench Pro, Terminal-Bench 4.0, and the Artificial Analysis Intelligence Index — with independent scores next to vendor scores, and per-token prices next to all of them. The picture that comes out is not "one model to rule them all." It is cheaper and more interesting than that.

The leaderboard that stopped mattering: SWE-bench Verified#

SWE-bench Verified was the industry's favorite coding benchmark: 500 real GitHub issues, an agent writes a patch, the test suite decides. It worked so well that everyone optimized for it — which is how it died. An analysis of the public results showed the top 8 submissions are a single statistical tier: a paired McNemar test on the 500 shared instances cannot separate #1 from #8 (the first significant separation appears at rank 9). Scores also clustered near the ceiling, and contamination concerns grew — vals.ai archived its independent board on September 5, 2026.

So treat Verified scores as a screening test, not a ranking: anything above ~93% is in the frontier tier, and the decimals between 95.4% and 97.0% are noise, not signal. Here is the independent vals.ai board (mini-swe-agent bash-only harness), with the prices those models carry on the 60+ model tokens.bd catalog where available:

ModelSWE-bench Verified (independent)Input / 1MOutput / 1MOn tokens.bd
Claude Opus 597.0% ±0.76——No
DeepSeek V4 Pro96.4%৳90.75 / $0.73৳272.25 / $2.18Yes
GPT-5.6 Sol96.2% ±0.86৳687.50 / $5.50৳4,125.00 / $33.00Yes
Grok 4.695.6% ±0.92৳275.00 / $2.20৳825.00 / $6.60Yes
GPT-5.6 Terra95.4% ±0.94——No
GLM-5.395.4% ±0.94৳192.50 / $1.54৳605.00 / $4.84Yes
Kimi K393.4%৳412.50 / $3.30৳2,062.50 / $16.50Yes
Qwen3.8-27B (open weights)86.0%৳55.00 / $0.44৳412.50 / $3.30Yes

Catalog prices checked 11 October 2026. Kimi K3's and Qwen3.8-27B's BDT figures are the site's own per-million prices, set separately in taka — exactly what you pay.

Two things to notice. First, the gap between the #1 and #5 score is 1.6 points — smaller than the error bars. Second, the price spread for that same 1.6 points runs from $1.02 to $11.00 per million tokens at an 80/20 input-output mix. Benchmarks have converged; prices have not.

Why the same model scores differently everywhere: the harness effect#

Before the rest of the tables, the most important number in this post: DeepSeek V4.1 Flash scores 90.6% on Terminal-Bench 2.1 in DeepSeek's own "DeepSeek Harness Minimal" scaffold — and 74.5% on the same benchmark in vals.ai's neutral harness. DeepSeek published both numbers, along with a table showing a wide spread across harnesses and benchmarks (Terminal-Bench 2.1: 84.1%–90.6% across six scaffolds; DeepSWE v1.1: 65.5%–74.2%). The scaffold is part of the score.

This is why SWE-bench Pro now exists in two versions that look like different benchmarks:

  • Vendor runs (provider's own scaffold): Claude Fable 5 at 80.0%, Claude Opus 5 at 79.2%, Qwen3.8 Max at 67.7%, GPT-5.6 Sol at 64.6%, GLM-5.2 at 62.1%.
  • Scale's SEAL standardized runs (same scaffold for every model): Muse Spark 1.1 at 61.5%, GPT-5.4 (xHigh) at 59.1%, Claude Opus 4.6 (thinking) at 51.9%.

The 80-vs-61 gap between those two boards is the harness effect in one line. When a vendor quotes you a number, the number is a property of a model–harness pair, not of the model. Independent, same-harness comparisons are the only ones you can sort.

Terminal-Bench 4.0: the agentic test#

Terminal-Bench runs models in a real terminal — file edits, shell commands, package installs, the works. The newest 4.0 suite is where Anthropic's September releases separated from the field, and it is also where vendor and independent numbers diverge most:

ModelTerminal-Bench 4.0 (vendor)Terminal-Bench 4.0 (Artificial Analysis, independent)
Claude Sonnet 5.570.6% (Anthropic)63.6%
Claude Opus 5.566.4% (Anthropic)59.6%
GPT-6 Astra57.9% (per Anthropic's comparison)—
DeepSeek V4.1 Flash31.2% (DeepSeek)—

Vendor figures checked against the Anthropic system-card numbers (9 October 2026) and DeepSeek's model card. The full Sonnet-vs-Opus comparison, with list prices and a worked session bill, is in our Claude Sonnet 5.5 vs Opus 5.5 deep dive from this morning — the short version: Sonnet leads on terminal tasks at half the per-token price.

The honest read: Artificial Analysis's independent runs knock roughly 7 points off both Anthropic vendor claims, and the ordering survives the knockdown. When the ordering survives the harness change, the signal is real.

SWE-bench Pro: the harder ranking that replaced Verified#

Scale AI's SWE-bench Pro draws on ~1,865 tasks from 41 repositories across Python, Go, JavaScript and TypeScript — longer, multi-file engineering work on code the model likely has not seen. It also has the strongest contamination control available: a 276-task commercial set from 18 proprietary startup codebases that are not on the public internet. And the commercial set reshuffles the ranking:

ModelPro public set (Scale SEAL)Pro commercial set
Muse Spark 1.1 (Meta)61.5% ±3.151.5% ±5.5
GPT-5.4 (xHigh)59.1% ±3.643.4% ±6.0
Claude Opus 4.6 (thinking)51.9% ±3.647.1% ±6.1
Gemini 3.1 Pro (thinking)46.1% ±3.632.2% ±5.7
Claude Opus 4.545.9% ±3.623.4% ±5.1

Scale AI standardized leaderboard, read September 14, 2026. The takeaway column is the commercial one: Opus 4.5 loses 22.5 points moving from public repos to private codebases, while Muse Spark 1.1 loses only 10 and leads both boards. If your work happens in a private codebase — and whose doesn't — the commercial column is the one that predicts your experience. Public-set scores on code every model may have trained on are the least informative numbers in AI right now.

General intelligence: the Artificial Analysis Intelligence Index#

For the non-coding picture, Artificial Analysis's Intelligence Index v4.3.2 combines ten graded evaluations (45% of them private, never-published test sets) into one 0–100 score:

ModelAA Intelligence Index v4.3.2Context
Claude Opus 5.557.62Leads, at max effort
Claude Sonnet 5.556.001.6 points behind Opus at half the price
Claude Fable 5.153.4Earlier Mythos-class tier
GPT-6 Astra52.7OpenAI's agent-built flagship
Gemini 4 Argon52.6Google, Sep 30 2026 launch
GPT-6.1 Sol (max)51.8OpenAI's price-performance pick
MiMo V2.6 Pro46Top open-weight model, tied with Grok 4.7

Index snapshot late September / 1 October 2026. The coding story and the general-intelligence story agree on the ordering: Anthropic at the top, a tight cluster of flagships in the low 50s, then open weights. And notice what changed at the bottom of that table — MiMo V2.6 Pro, a 1.02-trillion-parameter open-weight model released September 21, 2026, sits at 46: the highest open-weight score on the board, above several closed models' tiers.

The value table: score per dollar#

Benchmarks without prices are trivia. Here is the independent SWE-bench Verified score divided by what 1M tokens at a typical 80% input / 20% output agent mix costs on the 60+ model tokens.bd catalog — all five from the site's own USD prices, checked 11 October 2026:

ModelVerified scoreCost per 1M (80/20, USD)Cost per 1M (80/20, BDT)Points per dollar
DeepSeek V4 Pro96.4%$1.02৳127.0594.5
GLM-5.395.4%$2.20৳275.0043.4
Grok 4.695.6%$3.08৳385.0031.0
Kimi K393.4%$5.94৳742.5015.7
GPT-5.6 Sol96.2%$11.00৳1,375.008.7

Worked: DeepSeek V4 Pro at 0.8 × $0.73 + 0.2 × $2.18 = $1.02 per million; 96.4 ÷ 1.02 = 94.5 points per dollar. GPT-5.6 Sol at 0.8 × $5.50 + 0.2 × $33.00 = $11.00; 96.2 ÷ 11.00 = 8.7. Ratio: 94.5 ÷ 8.7 ≈ 10.9 — about eleven times the score per dollar, for a model 0.2 points behind on the benchmark. The BDT column uses each model page's own taka price — no exchange-rate math, no card markup.

Horizontal bar chart titled 'Coding value for money: SWE-bench Verified score per dollar': DeepSeek V4 Pro leads at 94.5 points per dollar, followed by GLM-5.3 at 43.4, Grok 4.6 at 31.0, Kimi K3 at 15.7, and GPT-5.6 Sol at 8.7 — computed from independent vals.ai benchmark scores and tokens.bd list prices for a 1M-token 80/20 input-output mix, checked 11 October 2026

The chart makes the frontier's dirty secret visual: the models are within 3 points of each other on the benchmark and an order of magnitude apart on price. That is not a market inefficiency that lasts — it is why cheaper models keep winning share.

Open weights have nearly caught up#

Three results that would have been unthinkable eighteen months ago:

  • GLM-5.3 scores 95.4% on SWE-bench Verified — inside the error bars of the top closed model — and Z.ai has promised open weights.
  • MiMo V2.6 Pro holds the top open-weight spot on the AA Intelligence Index (46), with weights and 7,000+ RL environments released under MIT.
  • DeepSeek V4.1 Flash leads the Vals Index among open-weight models at 57.86% — just ahead of Kimi K3 at 57.81% — while its test runs cost $0.30 each against Kimi K3's $6.47, a ~22x gap ($6.47 ÷ $0.30 ≈ 21.6).

For developers in Bangladesh, that last point matters twice: the open-weight tier is where per-token prices collapse, and it is also where you can self-host instead of paying per token at all. Qwen3.8-27B at 86.0% Verified runs on a 24GB Mac. The frontier premium is now paying for the last few points, not for the capability itself.

How to read a benchmark like an engineer#

Four rules that survive every launch cycle:

  1. Independent beats vendor. Same-harness, third-party runs (vals.ai, Artificial Analysis, Scale SEAL) are comparable; vendor launch tables are model–harness pairs. When a vendor publishes its scaffold's effect size — like DeepSeek's 65.5%–90.6% harness table — believe the range, not the headline.
  2. Match the benchmark to your work. Terminal-Bench if your agent lives in a shell, SWE-bench Pro's commercial set if your code is private, multilingual suites if your stack is not Python-only.
  3. Watch the error bars. On saturated benchmarks, a 1–2 point "lead" is a coin flip. Ask for the confidence interval; if there isn't one, treat the ranking as decorative.
  4. Divide by the price. A model 0.2 points ahead that costs 11x more per token is not ahead. Recompute the value table for your own token mix with the method in how to estimate your monthly token bill.

And the meta-rule: benchmarks measure the past. Run a small eval on your own tasks before you move a workload — every lab's numbers say their model wins, and the harness table is the only one telling the truth.

Coding benchmark FAQ#

Which benchmark should I trust for choosing a coding model? No single one. Use SWE-bench Pro's Scale-standardized board for hard engineering work, Terminal-Bench for shell-heavy agent tasks, and the Artificial Analysis Intelligence Index for general reasoning. Prefer independent, same-harness comparisons over vendor launch tables, and check the confidence intervals — on saturated benchmarks the top models are statistically tied.

Why does the same model score so differently on different leaderboards? Three reasons: the agent harness (DeepSeek V4.1 Flash scores 90.6% on Terminal-Bench 2.1 in its own scaffold vs 74.5% in vals.ai's — a 16-point harness effect), the reasoning-effort setting, and contamination (public code the model may have trained on). Always compare scores from the same harness before drawing conclusions.

Is SWE-bench Verified still useful? As a screen, yes; as a ranking, no. The top 8 submissions are statistically indistinguishable, vals.ai archived its independent board in September 2026, and contamination concerns are documented. The differentiator has moved to SWE-bench Pro, where standardized scores sit 10–30 points below vendor-scaffold claims.

What is the best-value coding model on tokens.bd right now? On score-per-dollar against independent SWE-bench Verified results: DeepSeek V4 Pro at ~94.5 points per dollar (80/20 mix, ৳127.05 per million tokens) — about 11x the value of GPT-5.6 Sol. For a generalist at the lowest absolute price, MiMo V2.6 Pro runs ~৳72 per million tokens at the top open-weight AA Index score. For the premium closed tier, Claude Sonnet 5.5 leads Terminal-Bench at half of Opus 5.5's per-token price. Prices from the model catalog, checked 11 October 2026.

Do these benchmarks measure what my coding agent actually does? Approximately — and the approximation is the harness. Every score is a property of a model–scaffold pair, not the model alone. The benchmarks are the best public proxy available, but a small eval on your own repository is worth more than any leaderboard row. See the docs for wiring any of these models into your agent with one base_url change.

Benchmark figures from independent boards (vals.ai, Artificial Analysis, Scale AI) and vendor launch materials as noted, read 9–11 October 2026. List prices from the tokens.bd catalog, checked 11 October 2026. Benchmarks and prices change — the vendor's page is the authority for scores, and the model catalog for prices on Tokens. Last updated 11 October 2026.

Sources: AI Coding Leaderboard, Oct 2026 (vals.ai scores) · SWE-bench statistical-tie analysis · State of AI, Fall 2026 (AA Index, Terminal-Bench) · SWE-bench Pro private-set leaderboard · SWE-bench Pro leaderboard, Oct 9 2026 · DeepSeek V4.1 Flash benchmarks · SWE-bench leaderboard with prices · MiMo V2.6 Pro AA Index · OpenAI model lineup, Oct 2026 · LLM leaderboard methodology (AA vs LMArena)

Was this page helpful?

Still stuck? Open a support ticket

Use the coding models you already know, through one API

One key for OpenAI- and Anthropic-compatible tools. Pay in BDT or USD, and keep the coding agent you already use.

Create an account

Follow us for more articles