Same score, one-fifth the price — that's the whole story

Here's the deal: xAI released Grok 4.6 on August 12. The announcement itself is restrained — one line about building for long-running agents and ambitious interactive and visual work, then a few benchmark tables.

But the numbers in those tables reset the conversation this week. On the Artificial Analysis Intelligence Index, Grok 4.6 scored 61. That is exactly level with OpenAI's GPT-5.6 Sol at max reasoning effort. The only models above it are Anthropic's Claude Opus 5 at 63 and Fable 5 at 62. So this model is sitting in third place in the world right now.

That alone would be a "another frontier model shipped" story. The real story is in the next column over. Grok 4.6 costs $2 per million input tokens and $6 per million output tokens. GPT-5.6 Sol, with the identical index score, costs $5 in and $30 out. On output that's one-fifth. Against Claude Opus 5 at $5/$25, it's about a quarter.

When performance is level and price is a fifth, that's not a performance story — it's a pricing story. And pricing stories stick around a lot longer than benchmark stories. That's why Artificial Analysis headlined its own writeup "Grok 4.6 returns SpaceXAI to the intelligence frontier and leads on cost efficiency." The second half of that sentence is doing more work than the first.

What xAI has actually become

xAI was founded by Elon Musk in 2023. For most of its early life, the narrative was "the chatbot bolted onto X." Grok was known for real-time access to the X timeline and for a less sanded-down voice than its rivals. It was not known for topping leaderboards.

The corporate shape changed along the way too. On x.ai today the company signs its announcements as SpaceXAI, the name that followed the tie-up with SpaceX. That's more than branding. The binding constraint on frontier training is power and capital, and this is a company that now has a parent capable of supplying both. Colossus 2, the gigawatt-class training cluster, is what that combination produced.

The model lineage has compressed dramatically in the last year. Across Grok 4.3 to 4.5 to 4.6, the Intelligence Index score climbed from the high 30s to 61. Artificial Analysis puts it precisely: Grok 4.6 gains 5 points over Grok 4.5 and 23 points over Grok 4.3. And Grok 4.5 shipped on July 16. So that 5-point jump took a little over a month.

There's an important technical detail buried here. Grok 4.6 is not a new foundation model. It reuses the same 1.5-trillion-parameter V9 base as Grok 4.5, with what xAI calls a "supplemental training run" layered on top, followed by fresh supervised fine-tuning and reinforcement learning stages. The company says it used curated data plus an improved optimizer and training recipe, and leaves it there.

That matters because it means five points came out of a model of the same size. Closing a frontier gap through post-training alone, without a new pretraining run, tells you where the remaining headroom in this industry currently sits. It also tells you something about cadence: post-training is far cheaper and far faster than pretraining. As long as this approach keeps yielding points, xAI's release intervals can stay short.

Line the numbers up and you can see where the gap opens

Start with the frontier models side by side.

Model AA Intelligence Index Input $/1M Output $/1M Context
Claude Opus 5 (max) 63 $5 $25
Claude Fable 5 (max) 62
Grok 4.6 (high) 61 $2 $6 500k tokens
GPT-5.6 Sol (max) 61 $5 $30
Kimi K3 (Moonshot AI) below 61
Grok 4.5 56 $2 $6

Read it vertically and the scores are within two points. Read it horizontally and the prices differ by four to five times. That asymmetry is essentially the entire announcement.

The detailed benchmarks point the same direction. Per xAI's own numbers, Grok 4.6 posts 1,753 Elo on GDPval-AA v2, 69.9% on CursorBench v3.2, 65.9% on DeepSWE v1.1, 61.3% on FrontierCode v1.1, and 57.5% on APEX-Agents. Artificial Analysis adds 50.7% on τ³-Banking, 88.4% on Terminal-Bench v2.1, and 1,577 Elo on AA-Briefcase. The GDPval figure sits second only to Claude Opus 5, with confidence intervals that overlap Fable 5 and Qwen3.8 Max — meaning at that benchmark the models are statistically indistinguishable.

But the most interesting number isn't a benchmark score at all. It's cost per task. Artificial Analysis measures Grok 4.6 at $0.84 per task, level with Moonshot AI's open-weight Kimi K3. That's close to the first time a commercial frontier model has matched an open-weight model on effective cost rather than just on sticker price.

And that isn't only about per-token rates. On long-horizon work, Grok 4.6 finishes tasks in roughly 53 turns and ~0.5B input tokens, where Claude Opus 5 takes roughly 103 turns and ~2.0B input tokens. Half the turns, a quarter of the input tokens — then multiply by the per-token gap and you get the final invoice. When xAI writes that the model shows "improved self-testing and verification," this is where that claim shows up numerically. A model that flails less burns fewer tokens.

The price sheet has fine print, though. The developer docs show that Grok 4.6 changes rate above a 200k-token prompt. Under 200k it's $2 input / $0.50 cached input / $6 output. Above 200k it becomes $4 / $1 / $12 — exactly double. So the 500k context window is open, but using its full length is not free. Cache pricing moved too: it was $0.30 per million on Grok 4.5 and is $0.50 now. The headline numbers stayed identical to 4.5 while the cache rate quietly went up.

Speed is the weak spot. Artificial Analysis measures Grok 4.6 (high) at 65.5 output tokens per second with 39.57 seconds to first token. Grok 4.5 ran at 80 tokens per second, so this is a step backward. Out of 188 models tracked, it ranks 6th on intelligence and 80th on speed. This is not a model designed to answer fast — it's designed to think long and finish cheap. Reasoning effort is exposed in four steps (low, medium, high as default, and xhigh), so the intended usage is to dial down when you're in a hurry and up when the task is hard.

Who actually gains from this

xAI gains credibility. The company's last twelve months were remembered more for controversies than for leaderboard positions — political bias fights, deepfake litigation, a sharp drop in app downloads. Against that backdrop, a third-party measurement reading "third-most-intelligent, lowest cost" pulls the conversation back to the product. In enterprise sales, that's an asset with real value.

Developers and startups are the most direct beneficiaries. For anyone building agent products, model API spend is cost of goods sold. Buying comparable quality at a fifth of the price rewrites gross margin outright. In workflows that fire thousands of agent runs a day, the difference between $0.84 and $2–3 per task can decide whether the business closes at all.

Cursor shows up immediately as a distribution partner. Grok 4.6 was available in Cursor and in Grok Build on day one, with a first-week promotion doubling included usage in both. OpenRouter, Vercel, and Cloudflare carry it as well. Given how much xAI has struggled with its own consumer app distribution, riding developer-tool channels is the sensible play.

OpenAI and Anthropic are in an awkward spot. Both price on the premise that the smartest model commands a premium, and the justification for that premium just narrowed to a two-point gap. Benchmark scores are not the whole of real-world quality, of course, and both companies compete on safety posture, enterprise support, and ecosystem depth. Still, when a procurement lead walks in holding this table, there's one more thing to explain.

The open-weight camp is under pressure. A large part of the case for models like Kimi K3 was cost. When a commercial frontier model ties them on cost per task, that argument weakens. Open weights remain the only option for organizations that cannot let data leave the building — but "open weights because it's cheaper" no longer holds on its own.

Infrastructure vendors are the quiet winners. If post-training alone can move performance this much, release cadence shortens, and shorter cadence means more training and inference cycles. And historically, when token prices fall, total token consumption rises rather than falls.

We've seen price shocks before — the results split

Undercutting the frontier on price is not new. Past attempts landed in two different places.

The success case people keep citing is DeepSeek in early 2025. It delivered comparable reasoning at a fraction of the price and reset the industry's price expectations wholesale. Within months, major providers cut rates or introduced new low-cost tiers. The lesson is clear: a price shock moves the market's price curve before it moves any single company's market share. Grok 4.6 could follow the same path.

The failure pattern is just as well documented. Performance parity plus a price advantage is often not enough to dislodge an entrenched workflow. Plenty of cheaper models have climbed the leaderboards while incumbent vendors held their enterprise seats. The reason is inertia, not technology. Prompts are tuned to a specific model, eval pipelines are built around it, legal review is already done. Swapping models costs vastly more than changing one line of API config.

xAI's own history is worth reading here. Grok 4 Fast, released in June 2026, cut costs dramatically and got attention on launch — and in the months that followed, Grok app downloads fell sharply anyway. The company's problem was never price competitiveness; it was trust and distribution. Leading this launch with Cursor, OpenRouter, Vercel, and Cloudflare distribution looks like the lesson being applied.

One more caution. Price cuts cut both ways on brand. Once selling cheap becomes the habit, "the affordable alternative" hardens into a position, and climbing back to premium gets hard. That's a large part of why Anthropic rarely discounts the Opus line. While Grok 4.6 holds performance parity, the problem stays invisible. If the gap reopens in the next generation, price is all that's left.

How the rivals counter

OpenAI's easiest move is tier engineering. Leave the GPT-5.6 Sol headline rate alone and lower the effective price through batch processing, caching, and lower-reasoning tiers. Conceding at the negotiating table without touching the public price sheet is standard grammar in this industry. In parallel, expect OpenAI to push the argument toward Codex and its agent products — selling a finished workflow rather than a per-token rate.

Anthropic holds a different card. Opus 5 is still number one on the index, and Anthropic's edge has always been enterprise trust more than raw benchmarks. Safety evaluations, audit readiness, and regulatory posture are the axes it will keep using to justify a premium. What could sting is the Artificial Analysis efficiency measurement — twice the turns and four times the input tokens on long-horizon tasks. That's an efficiency problem, not a rate problem, and you can't paper over it with a discount.

Google answers with distribution. Gemini is already embedded in Cloud, Workspace, and Android, so it gets consumed inside contracts that were signed for other reasons. On that path, per-token comparisons rarely even happen. It's the most effective way to sidestep a price fight entirely.

Moonshot AI and the open-weight field will go further down-market or get more specialized. Once cost per task is level, competing on price alone in the general-purpose lane is difficult. What remains is on-premise deployment, fine-tuning freedom, and data sovereignty — axes commercial APIs struggle to offer.

The application layer, including Cursor, may be the biggest winner of all. If you keep the architecture model-swappable, supplier price competition lands directly in your margin. Grok 4.6 appearing in Cursor on launch day is that dynamic in action: application-layer companies have nothing but reasons to welcome one more cheap, capable model.

xAI's own next move is already telegraphed. In late July, Musk said a Grok 4.7 at 2.1 trillion parameters would follow within weeks of 4.6. The larger Grok 5 is reportedly training on Colossus 2, but its schedule has slipped repeatedly and there is still no official date. That part rests on Musk's statements and the reporting around them rather than a company announcement, so treat it as a pointer rather than a plan.

So what actually changes

For developers, now is the moment to re-run your cost math. For multi-turn agent workloads especially, look at total cost per task rather than per-token rates — Grok 4.6's advantage lives where the rate and the turn efficiency multiply together. But a near-40-second time to first token is felt directly in conversational UI, so if your users are sitting there watching a cursor blink, be careful about making this the default model.

For enterprise buyers, you just gained leverage. If you have a renewal coming up, the mere existence of a same-score, one-fifth-price alternative is a bargaining chip, even if you never intend to switch. Just make sure the contract accounts for the rate doubling above 200k prompt tokens. In workflows that stuff entire documents into context, that clause moves the real bill considerably.

For investors, margin compression is the thing to watch. As frontier performance commoditizes, pricing power drains out of the model layer and accrues to the application layer selling products on top. Two indicators matter: whether OpenAI and Anthropic actually adjust pricing, and whether large customers actually shift their model mix. Both are far more informative than another change in leaderboard order.

For everyday users, almost nothing changes today. Grok 4.6 landed through the API, Cursor, and Grok Build; the consumer app experience is unchanged. But when model costs fall, the pattern has been that consumer plans and free-tier limits eventually widen. If several chatbots quietly raise their free allowances a few months from now, this kind of price competition is the reason.

For the industry as a whole, the real message isn't the ranking — it's how fast capability is commoditizing. Closing five points in a month, without a new pretraining run, means the shelf life of a frontier lead has gotten short. The cost of holding the lead keeps rising while the premium that lead commands keeps shrinking. That squeeze is the force that defines the next several quarters.

🥄 Three Things You're Probably Wondering

— Should I switch models right now? Depends on the workload. For batch jobs and background agents where nobody is waiting on a screen, the savings are large enough to be worth a trial. For real-time conversational products, that 40-second time to first token is a genuine obstacle. Run an A/B on your own workload — it'll tell you more than any leaderboard.

— Is 61 versus 63 something I'd actually feel? Honestly, too early to say flatly. The Intelligence Index is a composite of several benchmarks, so how a two-point gap manifests varies by task. On GDPval the confidence intervals overlap, but on specific coding and agentic sub-benchmarks the ordering still separates. The only reliable move is to pick the benchmark that resembles the work you're actually assigning.

— Will this pricing hold? No way to know. The common read across the industry is that frontier pricing currently sits below cost in places, funded by the fight for share. Note that even this launch kept the headline rate flat while raising cache hits from $0.30 to $0.50 per million. Watch your actual invoices rather than the advertised numbers.

References

Numbers are as of announcement and may change.