Grok is frontier-grade at coding, agents, and price — about 8/10 overall, with one caveat: verify its facts

Asked:How good is grok?

Evidence snapshot as of 1 September 2026: five authoritative pages from xAI, Vals AI, and Artificial Analysis, expanded into 125 extracted benchmark results for Grok 4.5 and 4.6. Independent testers place Grok 4.6 in line with the top frontier models on intelligence and near the very top on agentic work, at a markedly lower price — while flagging that benchmark strength does not guarantee factual reliability.

The scorecard, by use case

Qualitative synthesis of the evidence on a 0–10 scale — an editorial reading, not a normalized scientific score. Hover a row for the benchmarks behind it.

Grok 4.6 against its closest rivals

Score in % · xAI's own eval table, 12 Aug 2026 · higher is better · suites are not interchangeable · hover a dot for the exact value

Independent: Artificial Analysis scores Grok 4.6 at 61 on its Intelligence Index — in line with GPT-5.6 Sol — with a GDPval-AA v2 Elo of 1753, behind only Claude Opus 5, and 88.4% on Terminal-Bench v2.1. artificialanalysis.ai
Independent: Vals AI ranks Grok 4.5 #8 overall on the Vals Index (65.30%), but #4 on SWE-bench Verified (86.60%) and #5 on GPQA Diamond (92.93%) — stronger in code and science than in the average. vals.ai
Value is a genuine edge: on a Vals agent test Grok 4.5 was the cheapest model at $0.17 per test with 55 tasks resolved, and xAI prices it at $2 in / $6 out per million tokens at 80 tokens/sec — 4.6 pricing may differ. vals.ai · x.ai
The caveat: alongside Grok 4.5's accuracy gains, Artificial Analysis measured its hallucination rate rising from 25% to 54% — so consequential factual claims still need source checking. artificialanalysis.ai

Should you use it?

Yes, confidently

Coding and software agents (top-5 independent rankings), long-running agentic workflows, research that leans on live X data, and any workload where cost per task matters — Grok sits on the cost-performance frontier.

Yes, with checks

General reasoning and knowledge work: competitive rather than dominant. For unsupervised, high-stakes factual output, treat Grok as a strong drafter whose claims you verify — measured hallucination rates argue for a human in the loop.

The evidence rows

ModelBenchmarkResultTesterKindSource
Grok 4.6AA Intelligence Index61Artificial Analysisindependentartificialanalysis.ai
Grok 4.6GDPval-AA v2 Elo1753 — behind only Claude Opus 5Artificial Analysisindependentartificialanalysis.ai
Grok 4.6Terminal-Bench v2.188.4%Artificial Analysisindependentartificialanalysis.ai
Grok 4.6𝜏³-Banking50.7% — top two with Qwen3.8 MaxArtificial Analysisindependentartificialanalysis.ai
Grok 4.6AA-Briefcase Elo1577 (debut)Artificial Analysisindependentartificialanalysis.ai
Grok 4.5Vals Index65.30% — #8 overallVals AIindependentvals.ai
Grok 4.5SWE-bench Verified86.60% — #4Vals AIindependentvals.ai
Grok 4.5GPQA Diamond92.93% — #5Vals AIindependentvals.ai
Grok 4.5Vibe Code Bench69.00% — #10 (up from #47)Vals AIindependentvals.ai
Grok 4.5Harvey Legal Agent Bench12.92% — #2Vals AIindependentvals.ai
Grok 4.5Cost / task (Vals agent test)$0.17 — cheapest, 55 tasks resolvedVals AIindependentvals.ai
Grok 4.6CursorBench v3.269.9%xAIvendorx.ai
Grok 4.6DeepSWE v1.165.9%xAIvendorx.ai
Grok 4.6APEX-Agents57.5%xAIvendorx.ai
Grok 4.5API price & speed$2 in / $6 out per 1M tokens · 80 TPSxAIvendorx.ai

Method: synthesis of 5 authoritative source pages (xAI, Vals AI, Artificial Analysis) expanded into 125 extracted model/benchmark rows, snapshot as of 2026-09-01. Independent rows (Vals AI, Artificial Analysis) are prioritized; vendor rows are xAI's own eval table. Benchmark suites use different tasks and scales and are not directly comparable; Elo and index scores are unitless. Duplicate and weakly labeled extractions were cut for space. The 0–10 scorecard is a qualitative editorial synthesis, not a measured quantity. Grok 4.5 pricing does not necessarily apply to Grok 4.6.

This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

Made withKeenable SELECT · 6,308 pages in 6m 34s · Ask your own questionShare:XLinkedInReddit
Made with Keenable SELECT