Grok 4.6 Arrives: Back at the Frontier with 61 Points, and a $0.84-Per-Task Bill
On August 12, xAI released Grok 4.6. The official positioning is clear: building on Grok 4.5, with a particular focus on long-running agents and more ambitious interactive and visual work — the model needs to stay with complex tasks across many steps, whether researching a topic, working across a codebase, or turning an idea into a polished application.
Benchmarks: Matching GPT-5.6 Sol, Behind Claude's Twin Flagships
On the Artificial Analysis Intelligence Index (a composite of nine benchmarks), Grok 4.6 scores 61, matching GPT-5.6 Sol (max), behind Claude Fable 5 (62) and Claude Opus 5 (63). Independent evaluator Artificial Analysis put it this way: this brings xAI "back to the intelligence frontier alongside OpenAI, behind only Anthropic." For reference, Grok 4.5 one month earlier scored 56 — a 5-point gain in one month, and +23 cumulative over Grok 4.3.
Where it pulls ahead is the agentic dimension:
- GDPval-AA v2 (real-world agentic knowledge work): Elo of 1753, behind only Claude Opus 5, with confidence intervals overlapping Claude Fable 5 and Qwen3.8 Max
- τ³-Banking (multi-turn customer service with tool use): 50.7%, among the top two
- Terminal-Bench v2.1: 88.4%, in line with the leading models
- In xAI's self-reported table: CursorBench v3.2 at 69.9%, DeepSWE v1.1 at 65.9%, and FrontierCode v1.1 at 61.3% — all above Grok 4.5's 66.7% / 54% / 56.6%
The other side of the official table deserves attention too: on the newer, harder Terminal-Bench v3.0, Grok 4.6 scores just 26% versus GPT-5.6 Sol's 34.6% — the gap on new-generation benchmarks is still real, and agentic strength does not mean across-the-board leadership.
Training: Putting Grok 4.5 to Work for Grok 4.6
The official training description is worth a close read. Grok 4.6 underwent a longer supplemental training run than Grok 4.5, using curated model-generated data and high-quality engineering data, with an improved optimizer and training recipe. xAI then used Grok 4.5 to regenerate the SFT trajectories across reasoning efforts, agent harnesses, and domains such as STEM, software engineering, and knowledge work, filtering out problematic traces with model-based checks. This was followed by agentic RL across knowledge work, general coding, and domain-specific environments for kernel optimization, web development, and computer-aided design. Older model as data factory, newer model as student — this self-iteration pipeline is becoming standard equipment for frontier labs.
Pricing: A Rare Generation That Gains Intelligence Without Raising Prices
Artificial Analysis specifically noted: at the frontier, intelligence gains usually come with price increases, but Grok 4.6 holds pricing at $2/$6 per million input/output tokens — more than 60% below Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). Measured cost per task is $0.84, the same as Kimi K3 with slightly higher intelligence, placing it on the Intelligence vs. Cost per Task Pareto frontier.
The turn efficiency is even more striking: on AA-Briefcase, the long-horizon knowledge work benchmark, Grok 4.6 completes tasks in roughly 53 turns and 0.5B input tokens on average, versus roughly 103 turns and 2.0B input tokens for Claude Opus 5 (max). Same-tier answers at half the turns and a quarter of the input — and since long-horizon agents accumulate context rapidly, this token-efficiency edge translates into cost savings well beyond the per-token price gap. The only price increase is on cache hits, up from $0.3 to $0.5 per million tokens; the context window stays unchanged at 500k.
So What
Grok 4.6 does not take a single outright first place on any benchmark, but it lands on the cost-performance Pareto frontier for every agentic evaluation in the Intelligence Index. When intelligence scores are packed into a narrow 61–63 band, what decides procurement may no longer be the leaderboard but the token bill at the end of the month — and that is probably the signal from this release truly worth watching.
Sources: xAI official release, Artificial Analysis independent review