How good is grok?
Evidence snapshot as of 1 September 2026: five authoritative pages from xAI, Vals AI, and Artificial Analysis, expanded into 125 extracted benchmark results for Grok 4.5 and 4.6. Independent testers place Grok 4.6 in line with the top frontier models on intelligence and near the very top on agentic work, at a markedly lower price — while flagging that benchmark strength does not guarantee factual reliability.
Qualitative synthesis of the evidence on a 0–10 scale — an editorial reading, not a normalized scientific score. Hover a row for the benchmarks behind it.
Score in % · xAI's own eval table, 12 Aug 2026 · higher is better · suites are not interchangeable · hover a dot for the exact value
| Model | Benchmark | Result | Tester | Kind | Source |
|---|---|---|---|---|---|
| Grok 4.6 | AA Intelligence Index | 61 | Artificial Analysis | independent | artificialanalysis.ai |
| Grok 4.6 | GDPval-AA v2 Elo | 1753 — behind only Claude Opus 5 | Artificial Analysis | independent | artificialanalysis.ai |
| Grok 4.6 | Terminal-Bench v2.1 | 88.4% | Artificial Analysis | independent | artificialanalysis.ai |
| Grok 4.6 | 𝜏³-Banking | 50.7% — top two with Qwen3.8 Max | Artificial Analysis | independent | artificialanalysis.ai |
| Grok 4.6 | AA-Briefcase Elo | 1577 (debut) | Artificial Analysis | independent | artificialanalysis.ai |
| Grok 4.5 | Vals Index | 65.30% — #8 overall | Vals AI | independent | vals.ai |
| Grok 4.5 | SWE-bench Verified | 86.60% — #4 | Vals AI | independent | vals.ai |
| Grok 4.5 | GPQA Diamond | 92.93% — #5 | Vals AI | independent | vals.ai |
| Grok 4.5 | Vibe Code Bench | 69.00% — #10 (up from #47) | Vals AI | independent | vals.ai |
| Grok 4.5 | Harvey Legal Agent Bench | 12.92% — #2 | Vals AI | independent | vals.ai |
| Grok 4.5 | Cost / task (Vals agent test) | $0.17 — cheapest, 55 tasks resolved | Vals AI | independent | vals.ai |
| Grok 4.6 | CursorBench v3.2 | 69.9% | xAI | vendor | x.ai |
| Grok 4.6 | DeepSWE v1.1 | 65.9% | xAI | vendor | x.ai |
| Grok 4.6 | APEX-Agents | 57.5% | xAI | vendor | x.ai |
| Grok 4.5 | API price & speed | $2 in / $6 out per 1M tokens · 80 TPS | xAI | vendor | x.ai |
Method: synthesis of 5 authoritative source pages (xAI, Vals AI, Artificial Analysis) expanded into 125 extracted model/benchmark rows, snapshot as of 2026-09-01. Independent rows (Vals AI, Artificial Analysis) are prioritized; vendor rows are xAI's own eval table. Benchmark suites use different tasks and scales and are not directly comparable; Elo and index scores are unitless. Duplicate and weakly labeled extractions were cut for space. The 0–10 scorecard is a qualitative editorial synthesis, not a measured quantity. Grok 4.5 pricing does not necessarily apply to Grok 4.6.