Artificial Intelligence (AI)

Gemini 3.7 Flash: The Headline Price Expires on December 31

Gemini 3.7 Flash is priced at seventy five cents per million input tokens only through the end of 2026, after which Google's published rate doubles, and the same promotional pricing and expiry applies to Gemini 3.6 Flash.

Gemini 3.7 Flash shipped on August 13, 2026, three weeks after 3.6 Flash. The coverage has focused on the benchmark jump, which is real. It has almost entirely missed the pricing footnote, which is the part that will actually appear on somebody’s invoice.

The number everyone is quoting is a promotion with an expiry date, and the expiry is four and a half months away.

What shipped, and what Google claims

Google’s model documentation lists gemini-3.7-flash as stable and describes it as "Our latest and most capable Flash model, built for complex coding, agentic workflows, and reliable multi-step execution." Its predecessor, gemini-3.6-flash, is now "Our previous-generation Flash model."

On the benchmark, DeepMind’s model page publishes a FrontierCode 1.1 score of 43.6% for Gemini 3.7 Flash against 34.4% for 3.6 Flash, described as measuring production code quality. That is a 9.2 point absolute gain, or roughly 27 percent relative, in three weeks.

Take that at face value with the usual caution. It is a vendor-run benchmark reported by the vendor, which is not the same as an independent result, and we have written at length about why vendor numbers need their evaluation conditions before they mean anything. The score is a claim, clearly published and clearly attributable, which is better than most.

The price is a promotion, and almost nobody is saying so

Here is the thing that is not a claim, because Google states it plainly in its own price list.

The Gemini API pricing page gives the input rate for Gemini 3.7 Flash as "$0.75 through December 31, 2026. $1.50 starting January 1, 2027." Output is "$3.75 through December 31, 2026. $7.50 starting January 1, 2027."

That is not a rumour, a leak or an analyst estimate. It is the published rate card, and it says the price doubles on a specific date.

Now look at how the model is being discussed. It is "$0.75 per million input tokens." That is how it appears in comparison tables, in cost calculators, in build guides and in the spreadsheets people are using this week to decide which model to standardise on. The expiry is a footnote when it should be the headline, because a 100 percent increase is not a rounding error in a spend model.

Four and a half months. If you are scoping a project that ships in Q4 and runs into next year, the price you are modelling is not the price you will pay for most of the project’s life. Model both numbers, or model the January rate and treat the discount as upside.

Gemini 3.7 Flash and 3.6 Flash cost exactly the same

This is the detail that reframes the release, and it took reading the pricing page rather than the announcement to find it.

Gemini 3.6 Flash carries identical pricing: $0.75 input and $3.75 output through December 31, 2026, then $1.50 and $7.50 from January 1. The promotion is not a launch discount attached to the new model. It covers the Flash tier, and both versions come off it on the same day.

Two consequences, and they point in opposite directions.

In the short term, this is about as clean an upgrade case as this market produces. Same price, same tier, a materially better published coding score. There is no cost argument for staying on 3.6 Flash, which is unusual, because normally the newer model is the more expensive one and you are trading budget for capability.

In the medium term, it means the exposure is bigger than a single model. Anyone on the Gemini Flash tier at all, on either version, is on promotional pricing that ends simultaneously. If you migrated to Flash this year on cost grounds, that decision needs revisiting before the new year rather than after it.

What doubling actually does to a build

Worth making concrete, because "the price doubles" lands differently when it is attached to a workload.

Take an agentic pipeline doing 40 million input tokens and 8 million output tokens a month, which is a modest production workload rather than a heavy one. Today that is $30 of input and $30 of output, so $60 a month. From January 1 the identical workload is $60 and $60, so $120.

The absolute numbers are small, which is exactly the trap. The multiplier does not care about scale. At 400 million and 80 million the same arithmetic runs from $600 to $1,200, and at ten times that from $6,000 to $12,000. Whatever your current Gemini Flash line item is, write the doubled figure next to it and decide now whether that changes any decision, because the alternative is discovering it in a January invoice. The mechanics of how those units are counted are covered in our explainers on tokens and on why input and output are priced differently.

The benchmark, and what it does not tell you

FrontierCode 1.1 measures production code quality, and Gemini 3.7 Flash scoring 43.6% on it is genuinely better than 34.4%.

What that does not establish is whether it is better at your work. A 43.6% score means the model failed the majority of the benchmark’s tasks. That is normal for a hard coding evaluation and it is worth saying out loud, because a 27 percent relative improvement in a headline reads like a solved problem and a sub-50 percent absolute score reads like what it is: meaningful progress on something still difficult.

The other thing it does not tell you is consistency. Google’s own positioning for the model emphasises "reliable multi-step execution," which is a property that shows up in agentic workloads and not in a single-shot code benchmark. If reliability across a long chain is what you need, the FrontierCode delta is suggestive rather than dispositive, and a week of your own evaluation traffic will tell you more than any published score.

The wider pattern: this tier moves fast and it retires fast

Three weeks between Flash releases is a genuinely short cycle, and it comes with a second cost that is easy to miss.

Google’s deprecation schedule retires Gemini 2.5 Pro, Flash and Flash-Lite in October 2026, and three imagen-4.0-* model IDs came off on August 17. gemini-2.5-flash-image shuts down on October 2. The Flash tier is being iterated and pruned on a cadence that makes any specific model ID a temporary dependency.

So the practical planning horizon for a named Gemini Flash model is shorter than a typical project. That is not a criticism of Google, whose deprecation documentation is clearer than most, but it is a design constraint. It is also part of why an abstraction layer between your application and the model has become standard advice, and why the ownership of that layer is now a live question.

What to do before January

Five positions, in the order I would take them.

Move to 3.7 Flash if you are on 3.6. Same price, better published score, same tier. This is the rare upgrade with no cost trade to weigh.

Reprice every Gemini Flash workload at the January rate today. Not in December. The point of doing it now is that you still have time to act on the answer.

Find out whether your budget owner knows. In most teams the person who approved the model choice and the person who will see the invoice are different people, and the expiry date has not travelled between them.

Benchmark on your own traffic before standardising. A published FrontierCode number is a reason to test, not a reason to commit. Reliability over multi-step chains is the property most agentic workloads actually depend on and it is not what that score measures.

Pin the model ID and diary the deprecations. The Flash line is iterating every few weeks and retiring older IDs on a published schedule. Know which specific ID you depend on and when Google says it goes.

Gemini 3.7 Flash is a good release at a good price. It is worth being precise that the good price has a date on it, and that the date applies to the whole tier.

Frequently Asked Questions

What does Gemini 3.7 Flash cost?

Google’s published rate is $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, then $1.50 and $7.50 respectively from January 1, 2027. The lower figures are promotional and the pricing page states the change explicitly. Any comparison table quoting $0.75 without the expiry is describing a rate that ends in about four and a half months.

Is Gemini 3.6 Flash cheaper?

No. It is priced identically, at $0.75 input and $3.75 output through the end of 2026, doubling on the same date. The promotion applies to the Flash tier rather than to the newest model, so there is no cost saving in staying on the older version and no cost penalty in moving to the newer one.

How much better is it?

On FrontierCode 1.1, which measures production code quality, DeepMind publishes 43.6% for Gemini 3.7 Flash against 34.4% for 3.6 Flash. That is 9.2 points absolute, roughly 27 percent relative. It is a vendor-run benchmark reported by the vendor, so treat it as a well-documented claim rather than an independent verification, and note that a 43.6% score still means most tasks were not completed.

Should I upgrade from 3.6 Flash?

For most workloads, yes, because there is no cost difference and the published coding score is materially higher. The caveat is the same as for any model swap: run your own evaluation traffic through it before standardising, particularly if your workload depends on consistency across long multi-step chains rather than single-shot code generation, since that is not what the headline benchmark measures.

What happens on January 1, 2027?

The published rate doubles for both Flash versions. Nothing else changes automatically, there is no migration and no action required, and your integration keeps working. You simply pay twice as much per token for the same workload. That is why repricing now is worth doing while there is still time to change the decision.

Will Google extend the promotion?

Unknown, and not something to plan around. The pricing page currently states a specific date and a specific higher rate, which is the only thing on the record. Building a budget on the assumption that a published price increase will be quietly dropped is a bet against the vendor’s own documentation.

How long will this model ID last?

Unstated, but the Flash line is moving quickly. Version 3.7 arrived three weeks after 3.6, and Google’s deprecation schedule retires the Gemini 2.5 family in October 2026, with `gemini-2.5-flash-image` shutting down on October 2 and three Imagen 4 model IDs already retired on August 17. Treat any specific Flash model ID as a dependency with a shorter life than your project.

Does the price change affect batch or cached tokens?

The rates quoted here are the standard input and output prices published on the Gemini API pricing page. Google lists separate rates for other modes, and those have their own entries and their own conditions. If a meaningful share of your spend runs through batch processing or context caching, check those rows directly rather than assuming the standard-tier promotion and expiry apply identically.

Digital Matters

Artificial Intelligence (AI) Desk