Cheap tokens are not the story: GLM-5.3-Flash pairs an independent intelligence score of 57 with vendor-reported Terminal-Bench 2.1 of 84.3 at $0.0013 per estimated request

Asked:“i want you to aggregate more data about benchmarks and capabilities for each model. Cost is not everything. Account for Intelligence, cost per task etc..”

Evidence matrix for all 25 models on the $10/month OpenCode Go plan: 309 collected benchmark and capability records from vendor cards, leaderboards and reviews (research date 2026-09-03), joined with OpenCode's official prices and estimated monthly request volumes. Scores are shown with their exact benchmark label and a confidence badge — unlike benchmarks are never averaged or compared, and a blank cell means no reliable evidence, not a low score.

The benchmark-evidence matrix

Independent leaderboard / benchmark Official first-party Secondary citation Thin / unlabeled source Blank = no reliable evidence. Hover a cell for the caveat and source.

Terminal-Bench 2.1 cohort: score vs plan volume

Ten models report a Terminal-Bench 2.1 score — the one benchmark version enough models share. Bubble size is the "benchmark-success proxy": ($10 ÷ estimated monthly requests) ÷ score. Directional economics only: publishers and harnesses differ even where the version label matches, and this is not production cost per solved issue.

Three cost concepts, kept separate

A · Externally measured cost per task

Only where a source states it directly. GLM-5.3-Flash: Intelligence Index v4.1.1 score 57 at a reported $0.045/task (discounted) on the Artificial Analysis workload — the strongest independent intelligence-per-cost anchor in the whole roster (docs.z.ai).

B · Subscription cost per estimated request

$10 ÷ OpenCode's official estimated monthly requests, valid only at full utilization of that one model. Shared rolling 5-hour/weekly/monthly caps and the per-model $15/$30/$60 included usage still apply. Range: $0.000044 (Muse Contributor) to $0.0204 (Kimi K3) (opencode.ai).

C · Benchmark-success proxy

(B) ÷ score fraction, computed only inside the exact Terminal-Bench 2.1 cohort: GLM-5.3-Flash ≈ $0.00150, DeepSeek V4 Flash Vision Exp ≈ $0.00063, Qwen3.8 Max ≈ $0.0143, Kimi K3 ≈ $0.0241. Never merged across benchmark versions or harnesses, and never a claim of real cost per solved task.

Routing playbook

Default

GLM-5.3-Flash — best combination of independent intelligence-per-cost evidence, coding/terminal scores (vendor-reported, harness-sensitive) and usable volume (7,900 req/mo at $0.15/$0.50).

Cheap parallel drafts & tests

MiMo-V2.5 (150,400 req/mo, $0.14/$0.28; Xiaomi reports SWE-Bench Pro 71.8 — needs an independent repo test). Optionally Qwen3.8 Flash after verifying the endpoint is the model behind the "Flash-Next" scores.

Long-horizon code escalation

MiMo-V2.5-Pro (SWE-bench Verified 78.9 reported, 16,300 req/mo) or Kimi K2.7 Code (explicit agentic-coding positioning, broad coverage). Compare both against GLM-5.3-Flash on your own repo.

Visual coding / UI / diagrams

DeepSeek V4 Flash Vision Exp — Terminal-Bench 2.1 83.9, DeepSWE 59.3, NL2Repo 57.7, plus multimodal-agent scores. Experimental build; and OpenCode's page dates DeepSeek's ZDR agreement through 2026-08-31, already past — verify before confidential use.

Hard terminal / agentic escalation

Qwen3.8 Max (Terminal-Bench 2.1 86.6, only 810 req/mo, ≈$0.0123/request) or Kimi K3 (Intelligence Index 57, provider-reported SWE-bench Verified 76.8, 490 req/mo, ≈$0.0204/request) — second-attempt models, not daily drivers.

Confidential work

Exclude Muse Spark 1.2 Contributor regardless of score — it permits training on prompts/completions, is not ZDR and is region-limited. Verify DeepSeek ZDR renewal; GPT 5.6 Luna and Grok 4.6 carry 30-day retention in OpenCode's table.

Why one leaderboard number is not intelligence

Benchmark contamination lets a model memorize test items released before its training cutoff. First-party scores skew upward: vendors pick the harness, prompt scaffold and token budget that flatter them — GLM used mini-swe-agent on DeepSWE while Qwen reported the better of two harnesses, so even identical benchmark names hide different tests. pass@k inflates over pass@1; Terminal-Bench 2.0, 2.1 and 3.0 are different exams (Grok 4.6 scores 26 on TB 3.0 while mid-tier models score 80+ on TB 2.1); version aliases like "Qwen3.8 Flash" vs "Flash-Next" may not be the same endpoint; and fresh releases simply lack independent replication — evidence immature, not weak. Real coding quality also includes correctness under review, edit discipline, tool-call reliability, latency, long-context handling and retry rate — none of which a single leaderboard number captures.

Run your own bake-off: pick 20–50 tasks from your repository, stratified across bug fixes, refactors, test writing, frontend work and tool-driven tasks. Run each candidate model at least twice per task. Record: success without human intervention, human review minutes, total tokens, wall-clock time and retries. Then compute production cost per accepted task — that number, not any leaderboard, decides your default.

Findings

GLM-5.3-Flash is the intelligence-per-cost standout: Intelligence Index v4.1.1 of 57 at a reported $0.045/task, with Z.ai-reported DeepSWE v1.1 63.4 and Terminal-Bench 2.1 84.3 — the coding numbers are first-party and harness-sensitive. docs.z.ai
MiMo-V2.5 is a high-volume intelligence bet, not just the cheapest usable model: Xiaomi reports SWE-Bench Pro 71.8, τ3-bench 62.3 and GDPVal-AA 69.5 at 150,400 estimated requests/month ≈ $0.000066 each — but Vals.ai rows pairing "0.0%" with ranks are parsing artifacts and were discarded. mimo.xiaomi.com
Qwen3.8 Flash's SWE-bench Pro 62.5, CoWorkBench 73.9 and 262K→1M context come from social/secondary posts that may describe Qwen3.8-Flash-Next — a strong value challenger at 27,000 req/mo, but confidence stays below GLM-5.3-Flash until endpoint equivalence is verified. x.com
DeepSeek V4 Flash Vision Exp is more than a cheap image add-on — its published text-agent scores beat the text build (Terminal-Bench 2.1 83.9 vs 82.7, DeepSWE 59.3 vs 54.4) with 1M context — but it is experimental and its ZDR agreement needs re-verification. eesel.ai
Evidence depth varies sharply: DeepSeek V4 Flash has 12 records across 10 sources while GLM-5.3-Flash and Hy4 preview arrived with almost none in the base crawl — the gap-fill set closed part of that, and blank matrix cells mark what remains unknown. opencode.ai

All 25 models

Method: 309 benchmark/capability records for 23 + 6 gap-filled OpenCode Go models, collected up to 2026-09-03 from vendor model cards, leaderboards, reviews and release posts; joined with OpenCode's official pricing, included-usage and estimated-monthly-request tables ($10/month plan). Cost per estimated request = $10 ÷ official estimated monthly requests, an utilization-dependent estimate under shared 5-hour/weekly/monthly caps, not a guarantee. Benchmark scores are quoted with their exact label and never averaged or compared across versions/harnesses; obvious extraction artifacts (Vals "0.0%"+rank rows, delta-only figures) were discarded; blank cells mean no reliable evidence. Duplicate corroborating records were cut for space.

This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

Keenable SELECTAsk your own question
Made with Keenable SELECT