Grok 4.6 finishes complex tasks in about 53 steps. Claude Opus 5 needs roughly 103. That gap says more about what SpaceXAI built than the headline benchmark score does.
Level with OpenAI, just under Anthropic
On the Artificial Analysis Intelligence Index, Grok 4.6 scores 61 points, tying OpenAI’s GPT-5.6 Sol. Only Anthropic‘s Claude Opus 5 at 63 and Claude Fable 5 at 62 score higher.
That’s a five-point jump over Grok 4.5. It’s also a two-point deficit to the top of the index, which is thin enough that you shouldn’t pick a model on the index alone.
Agentic tasks are where the step count matters
Grok 4.6 does its best work on agentic tasks, the ones where a model runs a multi-step workflow on its own without a human nudging it back on track.
On GDPval-AA v2, a benchmark built to measure real knowledge work done on a computer, it comes in second with an Elo score of 1,753, behind Claude Opus 5.
The step count is the part worth watching. Roughly half the steps for a task means roughly half the tokens burned getting there, and anyone who has watched an agent loop rack up a bill knows that’s not a rounding error.
The price is the real argument
Pricing stays at $2/$6 per million tokens. Claude Opus 5 runs $5/$25 and GPT-5.6 Sol runs $5/$30, which puts Grok 4.6 more than 60 percent cheaper than both.
Output tokens are where that spread bites hardest, and output tokens are exactly what agentic workloads generate in volume. A model that’s a couple of points behind on a benchmark index but a quarter of the output cost is an easy call for most production work.
Where you can run it today
Grok 4.6 is available now through the API, Cursor, Grok Build and partners including OpenRouter, Vercel and Cloudflare.
For the first week, x.ai is doubling the usage quota in Grok Build and Cursor. If you’re running agents that bill by the step, that’s the window to check whether 53 steps holds up on your own workload rather than on a benchmark harness.