Grok 4.6 matches GPT-5.6 Sol's intelligence at 60% lower cost and half the turns
On August 12, 2026, SpaceXAI shipped Grok 4.6, which scores 61 on the Artificial Analysis Intelligence Index — level with GPT-5.6 Sol — at $2/$6 per million tokens, and finishes long-horizon agentic tasks in half the turns of Claude Opus 5. For anyone building agents, the deciding variable is no longer the benchmark, it is the cost and token count burned per task.
August 12, 2026. SpaceXAI — formerly xAI — ships Grok 4.6, a frontier model aimed at long-running agents and interactive and visual work. The number that matters: it scores 61 on the Artificial Analysis Intelligence Index, level with GPT-5.6 Sol, behind Claude Opus 5 (63) and Claude Fable 5 (62). Except the price is elsewhere — $2 per million input tokens and $6 per million output, which is 60% below Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). The frontier battle just moved to new ground: cost per task.
The detail most people miss on first pass: Grok 4.6 does not win on raw score. It wins on economics. At near-equal intelligence, it costs a fraction of the price and finishes long-horizon agentic tasks in half the turns. For a team industrializing agents, that is exactly the metric that decides the end-of-month invoice.
A return to the frontier
The jump is real. Grok 4.6 gains 5 points over Grok 4.5 on the Intelligence Index, a month after that release, and +23 points versus Grok 4.3. Artificial Analysis sums up the trajectory: SpaceXAI is back at the frontier, alongside OpenAI, behind only Anthropic.
The raw numbers, on agentic and coding benchmarks:
- AA Intelligence Index: 61, level with GPT-5.6 Sol Max (61), behind Claude Opus 5 (63) and Claude Fable 5 (62);
- GDPval-AA v2: 1,753 Elo, behind only Claude Opus 5;
- CursorBench v3.2: 69.9%, ahead of GPT-5.6 Sol (67.2%);
- DeepSWE v1.1: 65.9%; FrontierCode v1.1: 61.3%; APEX-Agents: 57.5%.
The nuance worth keeping: Grok 4.6’s best scores land on agentic work, not static reasoning. This is a model built to sustain a chain of tasks, not to shine on a single question.
The release also lands in a crowded field. DeepSeek and Qwen keep shipping open-weight models at aggressive prices, while Google and Anthropic hold the top of the static-reasoning leaderboards. Grok 4.6’s bet is specific: win the agentic economy — long tasks, real code, repeated runs — rather than the one-shot reasoning crown. That is a narrower fight, but a far more commercially consequential one.
The price is the message
The decisive point is not technical, it is commercial. Grok 4.6 holds Grok 4.5’s pricing — $2/$6 per million tokens — at a moment when the previous frontier generation raised prices with every intelligence gain. The comparison that matters for a buyer: Claude Opus 5 at $5/$25 and GPT-5.6 Sol at $5/$30. On the output token, which dominates the bill for reasoning-heavy workloads, Grok 4.6 is four to five times cheaper.
Artificial Analysis puts the measured cost per task at $0.84 — the same as Kimi K3, but for higher intelligence. The result: the model sits on the intelligence-versus-cost Pareto frontier for nearly every agentic evaluation in the index. To put it plainly: nobody sells the token cheaper at this level of intelligence.
Two caveats, to be fair. The cache climbs from $0.3 to $0.5 per million tokens — an increase, but still far below the competition. And the fast variant is billed at twice the standard model’s price — a latency-based pricing mechanism that is becoming the norm at the frontier.
The underlying economics are worth spelling out. An agent working a complex task generates tens of thousands of output tokens per iteration, and re-reads its own context every turn. The bill for an agentic deployment is therefore dominated by two items: the output token (where Grok 4.6 is four to five times cheaper than its equivalents) and the volume of re-read input tokens (where its turn efficiency compounds). That is why $0.84 per task is a more telling number than any benchmark score: it folds price, turn efficiency and answer quality into a single metric.
Efficiency in turns, not tokens
The most instructive number is not the per-token price, it is turn efficiency. On AA-Briefcase, Artificial Analysis’s private benchmark of long-horizon agentic knowledge work, Grok 4.6 reaches an Elo of 1,577 — Fable 5-tier — but resolves tasks in ~53 turns and ~0.5 billion input tokens on average, against ~103 turns and ~2.0 billion tokens for Claude Opus 5 (max).
The consequence is direct. Long-horizon agentic work accumulates context at speed: a model that reaches a comparable answer in half the turns and a quarter of the input tokens has a cost advantage far beyond what its per-token price suggests. That is the real story of this release — and it is exactly what single-score benchmarks hide.
SpaceXAI also flags a behavior change: on long trajectories, the model checks its own work before moving on — self-testing and validation of what it has produced before advancing to the next step. That is the kind of behavior that separates an agent that silently drifts from one that finishes.
What changed under the hood
The recipe follows the current playbook, with an extra step. Grok 4.6 went through a longer supplemental training run than Grok 4.5, using model-generated data curated for reasoning and advanced technical concepts, an improved optimizer and a better training recipe. SpaceXAI then used Grok 4.5 to regenerate the SFT trajectories across reasoning levels, agent harnesses and domains (STEM, software engineering, knowledge work), filtering out problematic traces with model-based checks.
The agentic RL training spans a wide range: knowledge work, general coding, and specialized environments — kernel optimization, web development, computer-aided design. The safety positioning is explicit: Grok 4.6 is calibrated to be useful in vulnerability patching, accelerating the engineering design cycle, and augmenting AI research, with the widest pre-deployment testing suite the company has ever run.
The model is available today in Cursor, Grok Build, the API, and through partners (OpenRouter, Vercel, Cloudflare), with 2× included usage for the first week in Grok Build and Cursor. The context window stays at 500,000 tokens.
Verdict
The Grok 4.6 release confirms a shift: at the frontier, the score benchmark has become a ticket to enter, not a selling point. What separates models now is the economics of the agent — how many turns, how many tokens, how many dollars to finish a real task.
The recommendation is conditional and quantified: if you build long-running agents or code, and your bill is dominated by output tokens, Grok 4.6 is the best intelligence-per-dollar on the market right now — ahead of GPT-5.6 Sol and far ahead of Claude Opus 5 on cost per task. If you want the maximum raw score on static reasoning, Claude Opus 5 (63) keeps the edge, and Claude Fable 5 stays ahead on pure agentic work. But for the majority of industrialized use cases — where agents are relaunched hundreds of times a day — the variable that decides profitability is no longer peak intelligence: it is the number of turns and tokens the agent burns before returning a correct result.
References
- SpaceXAI — Introducing Grok 4.6, August 12, 2026
- Artificial Analysis — Grok 4.6 returns SpaceXAI to the intelligence frontier and leads on cost efficiency, August 12, 2026
- VentureBeat — SpaceXAI debuts Grok 4.6, overtaking Kimi K3 and matching GPT-5.6 Sol, August 12, 2026
- Unite.AI — SpaceXAI Launches Grok 4.6 for Long-Running Agents, August 12, 2026