xAI released Grok 4.6 on August 12, 2026. Within a week it was available in GitHub Copilot and on Amazon Bedrock, and the company pushed its Build Mode agent surface to general availability. The model itself is an incremental step over Grok 4.5, which shipped four weeks earlier.
The release is worth attention, but not primarily for the model. The price and the distribution are the story, and the benchmark table xAI published is more mixed than the announcement’s framing suggests.
What xAI shipped on August 12
The announcement describes Grok 4.6 as building on Grok 4.5 "with a particular focus on long-running agents and more ambitious interactive and visual work." At launch it was available in Cursor, in Grok Build, through the xAI API console, and via OpenRouter, Vercel, and Cloudflare, with double the included usage in Grok Build and Cursor for the first week.
The published scores, as xAI reported them:
| Benchmark | Grok 4.6 |
|---|---|
| AA Intelligence Index | 61 |
| GDPVal-AA v2 | 1753 |
| CursorBench v3.2 | 69.9% |
| DeepSWE v1.1 | 65.9% |
| FrontierCode v1.1 (Extended) | 61.3% |
| APEX-Agents | 57.5% |
| APEX-SWE | 56.4% |
| Terminal-Bench v3.0 | 26% |
| AA-Briefcase | 1577 |
| Harvey LAB (Vals) | 15.8% |
The framing deserves unpacking. "Long-running agents" describes a workload where the model plans, calls tools, reads results, and keeps going for minutes or hours without a human in the loop. That is a harder problem than answering well once, because errors compound: a wrong turn on step four is still wrong on step forty, and the model has to notice and recover. It is also the hardest capability to verify from a benchmark table, because the failure modes that matter show up over long horizons rather than in single-shot scoring. If the category is new to you, our explainer on what AI agents are covers the workload.
The price is the strategy
Grok 4.6 costs $2 per million input tokens and $6 per million output tokens, with a fast variant at double that. Those are the same rates xAI set for Grok 4.5 in July, so a month of capability improvement arrived at no additional cost per token.
The model documentation lists the API identifier as grok-4.6 and the context window at 500,000 tokens, which the launch announcement did not mention. That is a large window for the price, and it is the specification most relevant to the agentic workloads the model is aimed at, since a long-running agent accumulates context with every tool call it makes.
That figure is the most consequential number in the release. It places Grok 4.6 well below the per-token rates that flagship models from the larger US labs have typically carried, and the gap is widest on output, which is where agentic workloads spend most of their budget. If you are unsure why output pricing dominates the bill for agent work, our explainer on input versus output tokens covers the arithmetic.
We are deliberately not printing a competitor price table here. Frontier pricing has moved repeatedly this year and any comparison we set in type today would be wrong within weeks. Check current rates on each vendor’s own page before making a procurement decision on this basis.
Where the benchmarks actually improved
One clean comparison exists between Grok 4.5 and Grok 4.6, and it is worth isolating because it is the only apples-to-apples row across the two announcements.
On DeepSWE v1.1, Grok 4.5 scored 53%. Grok 4.6 scores 65.9%. Same benchmark, same version, a gain of nearly thirteen points in four weeks. That is a substantial improvement on a software engineering task and the strongest evidence in the release that this is more than a version bump.
Everything else in the two tables uses different benchmarks or different versions of the same benchmark, so no other direct comparison is available. That is not unusual, and it is not necessarily evasive, but it does mean the "smarter than 4.5" claim rests on one verifiable row plus xAI’s own characterization.
The Terminal-Bench number, and why you cannot compare it
Grok 4.6 posts 26% on Terminal-Bench v3.0. Grok 4.5 posted 83.3% on Terminal Bench 2.1. Those two numbers get placed side by side in coverage of this release and read as a collapse. They should not be.
What can be said honestly is narrower and still useful. 26% on the current generation of terminal evaluation is a low absolute number for a model marketed on long-running agents, and xAI published it anyway rather than omitting the row. Whether that reflects a genuine weakness or an unusually punishing new benchmark cannot be settled until independent results on v3.0 exist for other models. Treat it as an open question, and test terminal workloads yourself before committing.
The Harvey LAB legal reasoning score of 15.8% is similarly low in absolute terms, and similarly lacks a published peer set to judge it against.
Distribution is the real move
The week after launch is more telling than launch day. Per the xAI news index, Grok 4.6 reached GitHub Copilot on August 14 and Amazon Bedrock on August 19, and Build Mode went to general availability the same day.
That matters because it changes who can buy it. A model on Bedrock is purchasable through an existing AWS agreement, billed on an existing invoice, inside an account that has already cleared security review. A model in Copilot reaches developers whose employer has standardized on Microsoft and who will never open an xAI console. Neither route requires a company to form a commercial relationship with xAI at all.
Build Mode going generally available on August 19, the same day as the Bedrock listing, fits the same pattern. It is xAI’s own agent surface, the place a developer would use Grok 4.6 for exactly the long-running work the model is pitched at, and it went from limited access to open access at the moment the model landed in two other companies’ distribution channels.
Set that beside the price and the strategy is legible. xAI is not trying to win the top of the leaderboard. It is trying to be the cheap, adequate, already-approved option inside the procurement paths enterprises already use. Given the company’s ownership structure and the controversies attached to its consumer product, removing the need to sign anything with xAI directly is a rational play.
Choosing whether Grok 4.6 belongs in your stack
The honest recommendation has three parts.
Run your own evaluation on your own workload. The published table has one comparable row against the previous version and no independent replication at all. That is not enough to choose a model on, and it is especially not enough when the marketing claim is about long-running agents and the terminal number is 26%.
Buy it through Bedrock or Copilot if you are going to use it. The per-token rate is the same, procurement is simpler, and you inherit your cloud provider’s data handling terms rather than negotiating new ones.
Size the decision on output tokens. At $6 per million output, an agentic workload that would be uneconomic on a flagship model may be viable here, and that, rather than any benchmark, is the case for Grok 4.6. If your workload is small or latency-sensitive rather than volume-driven, the saving will not be what decides it.
What our earlier coverage does not yet reflect
One disclosure. Our pillar on Grok and the xAI model family was published on July 14, two days before Grok 4.5 shipped, and it describes Grok 4.3 as the current flagship with 4.5 in private beta. That was accurate on its publish date and is not accurate now. The lineage, the compute story, and the company background in that piece still stand. The current-flagship framing does not, and we are updating it separately.
Since that piece went out, xAI has also shipped Grok Bot, Imagine Image 2.0, Grok Voice Think Fast 2.0, Build Mode, and integrations across Google Workspace, Excel, and Outlook. The release cadence has been faster than our coverage, which is worth saying plainly rather than quietly correcting.
Frequently Asked Questions
What is Grok 4.6?
Grok 4.6 is xAI’s language model released on August 12, 2026, positioned as an incremental improvement over Grok 4.5 with a stated focus on long-running agents and interactive or visual work.
How much does Grok 4.6 cost?
$2 per million input tokens and $6 per million output tokens. A faster variant costs double. These are the same rates xAI charged for Grok 4.5, so the capability improvement came at no additional per-token cost.
Is Grok 4.6 better than Grok 4.5?
On the one benchmark reported at the same version across both releases, yes: DeepSWE v1.1 rose from 53% to 65.9%. Every other row in the two announcements uses a different benchmark or a different version, so no further direct comparison is possible from published data.
Why is the Terminal-Bench score only 26%?
Because it is measured on Terminal-Bench v3.0, a newer and harder evaluation than the 2.1 version Grok 4.5 was scored against. The two numbers are not comparable. Whether 26% is weak in absolute terms cannot be judged until other models publish v3.0 results.
Where can I use Grok 4.6?
Cursor, Grok Build, the xAI API console, OpenRouter, Vercel, and Cloudflare at launch. GitHub Copilot added it on August 14 and Amazon Bedrock on August 19.
Can I buy Grok without a direct xAI contract?
Yes. Amazon Bedrock and GitHub Copilot both provide access through an existing agreement with those vendors, which means billing, security review, and data handling terms run through your cloud or developer platform provider rather than through xAI.
Has anyone independently verified these benchmark scores?
Not at the time of writing. Every figure in the launch table is self-reported by xAI. Independent leaderboard results for Grok 4.6 were not available when this piece went out, which is the normal state of affairs in the first weeks after a model launch.
Does Grok 4.6 change what the Grok pillar says?
Yes. Our Grok pillar was published on July 14, before Grok 4.5, and names Grok 4.3 as the flagship. That framing is out of date. The model lineage and company background in that piece remain accurate, and the current-state sections are being updated separately.