Evidence matrix for all 25 models on the $10/month OpenCode Go plan: 309 collected benchmark and capability records from vendor cards, leaderboards and reviews (research date 2026-09-03), joined with OpenCode's official prices and estimated monthly request volumes. Scores are shown with their exact benchmark label and a confidence badge — unlike benchmarks are never averaged or compared, and a blank cell means no reliable evidence, not a low score.
Ten models report a Terminal-Bench 2.1 score — the one benchmark version enough models share. Bubble size is the "benchmark-success proxy": ($10 ÷ estimated monthly requests) ÷ score. Directional economics only: publishers and harnesses differ even where the version label matches, and this is not production cost per solved issue.
Only where a source states it directly. GLM-5.3-Flash: Intelligence Index v4.1.1 score 57 at a reported $0.045/task (discounted) on the Artificial Analysis workload — the strongest independent intelligence-per-cost anchor in the whole roster (docs.z.ai).
$10 ÷ OpenCode's official estimated monthly requests, valid only at full utilization of that one model. Shared rolling 5-hour/weekly/monthly caps and the per-model $15/$30/$60 included usage still apply. Range: $0.000044 (Muse Contributor) to $0.0204 (Kimi K3) (opencode.ai).
(B) ÷ score fraction, computed only inside the exact Terminal-Bench 2.1 cohort: GLM-5.3-Flash ≈ $0.00150, DeepSeek V4 Flash Vision Exp ≈ $0.00063, Qwen3.8 Max ≈ $0.0143, Kimi K3 ≈ $0.0241. Never merged across benchmark versions or harnesses, and never a claim of real cost per solved task.
GLM-5.3-Flash — best combination of independent intelligence-per-cost evidence, coding/terminal scores (vendor-reported, harness-sensitive) and usable volume (7,900 req/mo at $0.15/$0.50).
MiMo-V2.5 (150,400 req/mo, $0.14/$0.28; Xiaomi reports SWE-Bench Pro 71.8 — needs an independent repo test). Optionally Qwen3.8 Flash after verifying the endpoint is the model behind the "Flash-Next" scores.
MiMo-V2.5-Pro (SWE-bench Verified 78.9 reported, 16,300 req/mo) or Kimi K2.7 Code (explicit agentic-coding positioning, broad coverage). Compare both against GLM-5.3-Flash on your own repo.
DeepSeek V4 Flash Vision Exp — Terminal-Bench 2.1 83.9, DeepSWE 59.3, NL2Repo 57.7, plus multimodal-agent scores. Experimental build; and OpenCode's page dates DeepSeek's ZDR agreement through 2026-08-31, already past — verify before confidential use.
Qwen3.8 Max (Terminal-Bench 2.1 86.6, only 810 req/mo, ≈$0.0123/request) or Kimi K3 (Intelligence Index 57, provider-reported SWE-bench Verified 76.8, 490 req/mo, ≈$0.0204/request) — second-attempt models, not daily drivers.
Exclude Muse Spark 1.2 Contributor regardless of score — it permits training on prompts/completions, is not ZDR and is region-limited. Verify DeepSeek ZDR renewal; GPT 5.6 Luna and Grok 4.6 carry 30-day retention in OpenCode's table.
Benchmark contamination lets a model memorize test items released before its training cutoff. First-party scores skew upward: vendors pick the harness, prompt scaffold and token budget that flatter them — GLM used mini-swe-agent on DeepSWE while Qwen reported the better of two harnesses, so even identical benchmark names hide different tests. pass@k inflates over pass@1; Terminal-Bench 2.0, 2.1 and 3.0 are different exams (Grok 4.6 scores 26 on TB 3.0 while mid-tier models score 80+ on TB 2.1); version aliases like "Qwen3.8 Flash" vs "Flash-Next" may not be the same endpoint; and fresh releases simply lack independent replication — evidence immature, not weak. Real coding quality also includes correctness under review, edit discipline, tool-call reliability, latency, long-context handling and retry rate — none of which a single leaderboard number captures.
Run your own bake-off: pick 20–50 tasks from your repository, stratified across bug fixes, refactors, test writing, frontend work and tool-driven tasks. Run each candidate model at least twice per task. Record: success without human intervention, human review minutes, total tokens, wall-clock time and retries. Then compute production cost per accepted task — that number, not any leaderboard, decides your default.
Method: 309 benchmark/capability records for 23 + 6 gap-filled OpenCode Go models, collected up to 2026-09-03 from vendor model cards, leaderboards, reviews and release posts; joined with OpenCode's official pricing, included-usage and estimated-monthly-request tables ($10/month plan). Cost per estimated request = $10 ÷ official estimated monthly requests, an utilization-dependent estimate under shared 5-hour/weekly/monthly caps, not a guarantee. Benchmark scores are quoted with their exact label and never averaged or compared across versions/harnesses; obvious extraction artifacts (Vals "0.0%"+rank rows, delta-only figures) were discarded; blank cells mean no reliable evidence. Duplicate corroborating records were cut for space.