xAI released Grok 4.7 on September 21, 2026, at exactly the per-token price Grok 4.6 carried: $2 per million input tokens and $6 per million output. The context window is unchanged at 500,000 tokens. What changed is how much the model writes. In independent testing, Grok 4.7 produced about twice as many output tokens as Grok 4.6 on the same work, so each task costs more even though the rate card did not move.
This piece covers why the price table hides the cost change, how reasoning effort drives the bill, what developers must handle on the Responses API, and where teams can use the model. It also lays out where xAI’s own benchmark figures disagree.
The short version: budget by task, not by token. Artificial Analysis measured about $3.74 per task for Grok 4.7 at its highest effort setting, against $1.86 for Grok 4.6 at the default setting. At the same default setting, Grok 4.7 still costs about 47% more per task. In GitHub Copilot the model is on by default for most organizations. In Microsoft 365 Copilot, Grok stays off until an admin turns it on.
What xAI shipped on September 21
The xAI release notes list the model as grok-4.7 on the xAI API. It accepts text and image input and returns text only. It supports four reasoning effort levels: low, medium, high (the default) and xhigh.
The launch announcement names the main changes. xAI describes a new, larger base model than Grok 4.6 and a longer reinforcement learning run on a harder mix of tasks. It also cites better self-verification, better context management and a new safeguard stack. These are xAI’s own descriptions, not independently tested claims.
Our Grok 4.6 coverage from August covered the price story. Grok 4.7 keeps that price, so the useful questions now are what a task costs and who can switch the model on. For the wider model family, see our Grok AI explainer.
Grok 4.7 pricing and context match Grok 4.6
The release notes carry the same pricing line for both models. Below 200,000 prompt tokens, input costs $2 per million, cached input $0.50 and output $6. Above 200,000 prompt tokens, all three rates double.
| Item | Grok 4.6 (Aug 12) | Grok 4.7 (Sept 21) |
|---|---|---|
| Input, under 200k prompt tokens | $2.00 per 1M | $2.00 per 1M |
| Cached input, under 200k | $0.50 per 1M | $0.50 per 1M |
| Output, under 200k | $6.00 per 1M | $6.00 per 1M |
| All rates, over 200k prompt tokens | $4 / $1 / $12 | $4 / $1 / $12 |
| Context window | 500k tokens | 500k tokens |
There is no context upgrade. Grok 4.6 already had a 500,000 token window, as our August post reported.
Two smaller price points are new. The Grok 4.7 model page says the model is also served at a US regional endpoint that "keeps inference in the United States, with token usage priced at a 10% premium." And xAI’s fast variant, Grok 4.7 Fast, costs twice the standard token rates. The release notes say it is available only through Cursor and Grok Build, not on the public xAI API.
The 200k line is where the bill doubles
The 200,000 token threshold applies to the prompt, and crossing it doubles every rate on the request, including output. If tokens as a billing unit are new to you, our explainer on what tokens are covers how text turns into billable units.
Most chat-style work stays well under that line. Loading a large codebase for a refactor can cross it, and so can a long set of board minutes or grant documents in one request. A long agent session that keeps its full tool history also grows past it.
The model page mentions context compaction for long agent loops. Compaction and splitting work into smaller requests are the practical ways to stay under the line. A 500,000 token window is useful, but the second half of it costs twice as much.
Where the extra cost comes from: output tokens
Artificial Analysis runs every model through the same ten evaluations in its Intelligence Index (currently v4.3.2) and publishes the tokens and money each run takes. That gives a like-for-like view of Grok 4.7 and Grok 4.6 at both of the top effort settings.
| Model and effort | Index score | Output tokens on the index | Cost per index task |
|---|---|---|---|
| Grok 4.6 (high) | 44 | 94M | $1.86 |
| Grok 4.6 (xhigh) | 44 | 97M | $2.32 |
| Grok 4.7 (high) | 46 | 200M | $2.73 |
| Grok 4.7 (xhigh) | 46 | 240M | $3.74 |
Figures are from the Artificial Analysis model pages as read on September 27, 2026. They get re-run, so check them again before quoting them in a budget.
At the default high setting, Grok 4.7 wrote a little over twice the output tokens Grok 4.6 did. Its cost per task rose about 47%, from $1.86 to $2.73. Artificial Analysis describes the xhigh run as "very verbose" compared with a median of 88 million output tokens across models.
Output is the expensive side of the meter at three times the input rate. That is why a model that talks more costs more at an identical price. Our piece on input versus output tokens walks through the arithmetic. The same pattern showed up in Gemini this month: our look at Gemini 3.8 Flash cost per task found a 57% higher run cost at an unchanged rate card.
Reasoning effort is the cost setting to manage
The table has a second finding. On Artificial Analysis’s index, Grok 4.7 scores 46 at both high and xhigh. The xhigh run cost about 37% more per task for no change in the headline score. Grok 4.6 shows the same shape: 44 at both settings, with xhigh about 25% more expensive.
An index score is an average across ten tests, so xhigh may still help on the hardest tasks. But the default is not the cheapest setting, and the top setting is not free.
A practical approach for a small team:
- Start at high, not xhigh: it is the default, and the independent index shows no average gain from going higher.
- Test low and medium on routine work: summaries, content edits and small code fixes may not need more. We found no independent cost data for these two settings, so measure your own.
- Reserve xhigh for tasks where a wrong answer is expensive: a production migration, a tricky bug, a legal or policy summary that someone will act on.
- Track cost per completed task: include the time someone spends fixing the output, not only the API bill.
What the Responses API changes for developers
Two details on the Grok 4.7 model page affect code, not only budgets.
First, encrypted reasoning. On the Responses API, the release notes say Grok 4.7 "always returns reasoning.encrypted_content, even when include does not list it." The model page tells developers to "pass the reasoning items back unchanged in the next request’s input." That keeps the model’s reasoning chain intact across turns in a multi-step conversation. Integrations built for Grok 4.6 should be checked for this. The same page says "Chat Completions is unchanged."
Second, caching. xAI "highly recommend[s]" setting a prompt_cache_key, which routes related requests to the same server. The model page warns that "without it you often pay full input price on a cache-cold server." Cached input bills at $0.50 per million instead of $2, so this setting matters for any workload that repeats a long system prompt or document set.
Where teams can use Grok 4.7 today
How Grok 4.7 gets switched on differs a lot by channel.
| Channel | Grok 4.7 status | Who turns it on |
|---|---|---|
| xAI API | Available as grok-4.7 | Your developers |
| GitHub Copilot | Available, gradual rollout | On by default unless an admin disables it |
| Microsoft 365 Copilot | Grok offered; version not stated | Off until an admin enables it |
| Cursor, Grok Build | Available, including Fast | Each user |
| Amazon Bedrock | Not listed as of Sept 27 | Not applicable |
The Grok 4.7 model card also lists Office add-ins for Word, PowerPoint and Excel, plus gateways including OpenRouter, Vercel, Cloudflare, Snowflake and Databricks Mosaic. xAI’s Word add-in listing in Microsoft’s marketplace says it is available to SuperGrok, Heavy, Business and Enterprise plans. That is xAI’s own add-in, separate from Grok inside Microsoft 365 Copilot.
Amazon Bedrock has offered Grok 4.6 since August 19. As of September 27 we found no Bedrock model card or announcement for Grok 4.7. Teams that chose Bedrock for its procurement path, as our August post suggested, are still on 4.6 there.
GitHub Copilot: on unless an admin turns it off
The GitHub changelog says "Grok 4.7 will be available to Copilot Pro, Pro+, Max, Business, and Enterprise SKUs." It lists Visual Studio Code, Visual Studio, Copilot CLI, Copilot cloud agent, the Copilot app, JetBrains, Xcode and Eclipse. GitHub’s supported models page lists it as generally available.
The line that matters for admins is this one: "Under default model enablement, new models are automatically enabled unless an administrator has turned off the global default or explicitly disables this model."
For a Business or Enterprise organization, that means developers may already see Grok 4.7 in the model picker. If your organization has rules about which AI vendors can see client code, check the Copilot policy settings now. Decide whether Grok 4.7 stays on, and whether the global default for new models should stay on at all.
Microsoft 365 Copilot: admin opt-in with its own data terms
Microsoft’s Learn page on AI providers (updated September 11) describes the setup. In the Microsoft 365 admin center, an admin goes to Copilot, then Settings, then View all, then AI providers for other large language models. From there the admin selects SpaceXAI, accepts the legal terms and chooses which users or groups get access. Users need a Microsoft Copilot license.
The data terms are the part to read slowly. Microsoft says this data "is processed outside all Microsoft managed environments and audit controls." Microsoft’s own customer agreements, including its Data Processing Addendum, do not apply. xAI’s Enterprise Terms and xAI’s Data Processing Addendum govern it instead.
For an association or nonprofit that relies on Microsoft’s terms for member or donor data, turning this on is a data-processing decision, not a feature toggle.
Several details come from secondary sources. Office Watch reports that the option is off by default, covers Word, Excel and PowerPoint, and requires the Frontier early access program. It also reports that tenants in the EU, EFTA and UK are excluded during the preview. Third-party copies of message center notice MC1474124 add government clouds (GCC, GCC High and DoD) to the exclusions. We could not read Microsoft’s own announcement text, so treat these as reported, not confirmed.
xAI’s Grok 4.7 benchmark figures do not agree with each other
xAI published benchmark results in two places on launch day, the announcement and the model card. Several figures differ between them.
| Figure | Announcement | Model card |
|---|---|---|
| EEBench, Grok 4.7 (xhigh) | 64.0% | 66.0% |
| EEBench, Grok 4.6 | 53.0% (high) | 60.0% (xhigh) |
| Terminal-Bench 4.0, Grok 4.7 (xhigh) | 37.6% | 38.0% |
| Knowledge cutoff | May 2026 (docs page) | June 2026 pretraining, data to August |
The EEBench gap for Grok 4.6 may come from the different effort settings. The Grok 4.7 gap cannot, since both documents label it xhigh. Grok 4.6’s DeepSWE score also moved: 65.9% in xAI’s August announcement, 65.2% in the Grok 4.7 table.
The announcement table has two more quirks. Grok 4.7’s DeepSWE result of 71.0% is marked as a high effort run, though the column is headed xHigh. And Grok 4.6 is shown at high while competitors are shown at max.
None of these gaps changes the broad picture. They are a reason to read the table as marketing, as we argued in our piece on vendor benchmarks and independent leaderboards.
Reading xAI’s table against the independent index
In xAI’s own table, Grok 4.7 at xhigh beats Grok 4.6 at high on every row. It leads both Claude Fable 5.1 and GPT-5.6 Sol on EEBench and the Harvey Legal Agent benchmark. It trails Fable 5.1 on CursorBench 4.0 (46.3% against 51.8%) and Terminal-Bench 4.0 (37.6% against 57.9%).
Benchmark versions also changed since August. Our Grok 4.6 post printed CursorBench v3.2 at 69.9% and Terminal-Bench v3.0 at 26%. The new table uses CursorBench 4.0 and Terminal-Bench 4.0, where Grok 4.6 scores 40.4% and 20.3%. Those are different tests, not a decline.
Independent results now exist for both models. In August we noted that none were available for Grok 4.6. Artificial Analysis now scores Grok 4.6 at 44 and Grok 4.7 at 46, ranking Grok 4.7 21st of 211 models listed. On the same leaderboard, Claude Fable 5.1 and GPT-6 Astra score 53 at their max settings, and the top model scores 58. xAI’s August table listed Grok 4.6 at 61 on the same index. We could not confirm which index version that figure used, so do not compare it with 44.
The model card also says Grok 4.7 reaches results "with fewer steps and fewer output tokens than other frontier models." Artificial Analysis’s token counts point the other way.
How to test Grok 4.7 before switching
On the API, switching is a model ID change; the cost is what needs testing. Pick three tasks your team runs every week, such as a code fix, a content rewrite and a document summary. Run each on your current model and on Grok 4.7 at high, then try the most important one at medium and xhigh. Record tokens, cost and the minutes spent correcting the result.
Keep prompts under 200,000 tokens where you can, set prompt_cache_key, and pass the encrypted reasoning back on the Responses API. If the result is better enough to justify roughly 1.5 times the cost per task at the default setting, switch. If not, stay on Grok 4.6. As of September 27, we found no retirement date for it in xAI’s docs.
Frequently Asked Questions
How much does Grok 4.7 cost per million tokens?
$2 for input, $0.50 for cached input and $6 for output, on prompts under 200,000 tokens. Above that, the rates are $4, $1 and $12. These are the same rates as Grok 4.6, per xAI’s release notes.
Is Grok 4.7 more expensive than Grok 4.6?
Not per token. Per task, yes. Artificial Analysis measured $2.73 per index task for Grok 4.7 at high effort against $1.86 for Grok 4.6 at high, because Grok 4.7 wrote about twice as many output tokens.
How large is the Grok 4.7 context window?
500,000 tokens, the same as Grok 4.6. It accepts text and image input and returns text only. Prompts over 200,000 tokens are billed at double the standard rates.
Which reasoning effort levels does Grok 4.7 support?
Low, medium, high and xhigh. High is the default. On Artificial Analysis’s index, high and xhigh both score 46, and xhigh costs about 37% more per task.
What is reasoning.encrypted_content, and do I need to handle it?
It is the model’s reasoning, returned in encrypted form on the Responses API. Grok 4.7 always returns it. xAI says to pass the reasoning items back unchanged in the next request’s input so multi-turn work keeps its reasoning chain. Chat Completions is unchanged.
Can I call Grok 4.7 Fast from the API?
No. Per xAI’s release notes, Grok 4.7 Fast costs twice the standard token rates and is available only through Cursor and Grok Build, not the public xAI API.
Which GitHub Copilot plans include Grok 4.7?
Copilot Pro, Pro+, Max, Business and Enterprise. New models are enabled by default unless an administrator has turned off the global default or disabled the model.
Is Grok 4.7 in Microsoft 365 Copilot?
Microsoft offers Grok models from SpaceXAI in Microsoft 365 Copilot, but it has not said which version. An admin must enable SpaceXAI and accept xAI’s terms, and the data is processed outside Microsoft’s environment and audit controls.
Is Grok 4.7 available on Amazon Bedrock?
As of September 27, 2026, we found no Bedrock listing or announcement for Grok 4.7. Grok 4.6 has been on Bedrock since August 19.
How does Grok 4.7 compare with Claude Fable 5.1?
On Artificial Analysis’s Intelligence Index, Grok 4.7 scores 46 and Claude Fable 5.1 scores 53 at its max setting. In xAI’s self-reported table, Fable 5.1 leads on CursorBench 4.0 and Terminal-Bench 4.0, while Grok 4.7 leads on EEBench and the Harvey Legal Agent benchmark.