Artificial Intelligence (AI)

Gemini 3.8 Flash Costs 57% More to Run Than 3.7 Flash, at Exactly the Same Price

Gemini 3.8 Flash and Gemini 3.7 Flash carry an identical published rate card of seventy five cents per million input tokens and three dollars seventy five per million output tokens, yet Artificial Analysis paid one thousand and seventy eight dollars to run its Intelligence Index on 3.8 Flash against six hundred and eighty eight dollars for the same suite on 3.7 Flash, a fifty seven percent higher bill for two index points, because 3.8 Flash emitted one hundred and forty million output tokens against seventy nine million, which is the practical demonstration that the number that lands on an invoice is cost per completed task rather than cost per million tokens, and that thinking tokens billed as output, a four thousand and ninety six token minimum before implicit caching can hit, and thought signatures replayed on every agentic turn are the mechanisms that separate the two figures.

Gemini 3.8 Flash shipped on September 2, 2026, three weeks after 3.7 Flash, and it carries an identical published rate card: $0.75 per million input tokens, $3.75 per million output.

Artificial Analysis paid $1,077.95 to run its Intelligence Index on Gemini 3.8 Flash. It paid $687.92 to run the same suite on 3.7 Flash. That is a 56.7% higher bill at the same posted price, for two points of index score.

The difference is not a pricing change. It is output volume: 140 million tokens against 79 million, a 77% increase, for the same evaluation.

This piece covers what shipped, what the two models cost on paper, what they cost in practice, the benchmark result closest to the work agencies actually do, the six mechanisms that separate the rate card from the invoice, the counterexample that nearly breaks the argument, and what to do about it.

The short version: the sticker price cannot distinguish these two models, because it is the same sticker. The variable that decides your bill is tokens burned per completed task, and that is a property of how talkative a model is, not of its rate card.

What Gemini 3.8 Flash is

The model ID is gemini-3.8-flash, released September 2, 2026. Context window is 1,048,576 tokens in, 65,536 out. It accepts text, image, video, audio, and PDF, and produces text. Knowledge cutoff is March 2026 for most domains.

Thinking levels are low, medium, and high, defaulting to medium. Notably minimal is not supported and returns a 400. Google’s model card says the model "is based on Gemini 3.7 Flash," and Google positions it as building on 3.7 rather than replacing it. Google calls it "our most intelligent workhorse model."

Nothing is being deprecated. Gemini 3.7 Flash, 3.6 Flash, 3.5 Flash, and 2.5 Flash all remain available with no announced shutdown date, and Google explicitly still recommends 3.7 Flash for efficiency-first workloads.

The release cadence is itself part of the story. This is Google’s fourth Flash model in under four months: 3.5 in May, 3.6 in July, 3.7 on August 13, and 3.8 on September 2. Gemini 3.7 Flash was the current model for twenty days.

The rate card, in full

All figures per million tokens, and identical for 3.8 Flash and 3.7 Flash.

Tier Input Output, including thinking Cached input
Standard $0.75 $3.75 $0.075
Batch $0.375 $1.875 $0.0375
Flex $0.375 $1.875 $0.0375
Priority $1.35 $6.75 $0.135

Cache storage is billed separately at $0.50 per million tokens per hour.

Two things worth flagging that are genuinely in Flash’s favor. There is no context-length tiering: Vertex lists the same rate above and below 200,000 tokens, where the Pro tier doubles. And Flash is roughly 2.7 times cheaper than Pro on input and 3.2 times cheaper on output at short context, widening above 200,000 tokens where Pro tiers up and Flash does not.

One thing worth flagging that is not in its favor and gets little coverage: Vertex charges 10% more on non-global endpoints, $0.825 in and $4.125 out. If you pin to an EU region for data residency, you pay the premium.

The bill is not the rate card

Artificial Analysis publishes both the score and the cost of producing it, which is unusual and is why this comparison is possible at all.

Index (v4.2) Cost to run the index Output tokens
Gemini 3.8 Flash (high) 47 $1,077.95 140M
Gemini 3.7 Flash (high) 45 $687.92 79M

Artificial Analysis put the per-task figure at $0.58 for 3.8 Flash at high effort against $0.40 for 3.7 Flash, and stated the cause directly: "This is up ~40% from Gemini 3.7 Flash ($0.40) despite unchanged per-token pricing." The driver it identifies is a 30% increase in output tokens per task, to roughly 48,000.

Google documents the behavior itself in the model card, which is to its credit: "At times, the model might use more tokens to maximize performance, especially at higher effort levels."

Artificial Analysis also notes that 140 million output tokens sits well above the roughly 79 million median for comparable models. The model is verbose by disposition, and verbosity is billed at the output rate.

At low effort the per-task cost drops to $0.24 and at medium to $0.41, which is the practical lever if you adopt it.

On the benchmark closest to agency work, 3.8 does not beat 3.7

Zapier’s AutomationBench runs models through end-to-end business workflows across 47 simulated SaaS tools in six domains, scored with deterministic assertions rather than a language model judge, on a held-out private set. It is the closest public evaluation to what an agency actually builds: multi-tool workflow automation.

The relevant rows from version 1.0.6:

Model Score Cost per task
GPT-6 Astra (max) 41.4% $1.77
Claude Fable 5.1 with Opus 5 fallback (max) 31.4% $2.45
Gemini 3.7 Flash (high) 30.44% $0.61
Gemini 3.8 Flash (medium) 29.68% $0.55
Gemini 3.8 Flash (high) 29.68% $0.62
GPT-5.6 Sol (max) 28.77% $0.91

Two findings sit in that table.

Newer scores slightly worse and costs slightly more. 3.7 Flash at high effort beats 3.8 Flash at high effort, by three quarters of a point, at a penny less per task.

Paying for high effort bought nothing here. 3.8 Flash scores identically at medium and high, 29.68% both times, while high costs 13% more. On this workload, the effort dial is a price control with no measurable output.

Two caveats to hold, and they are the ordinary ones for a vendor-published benchmark. AutomationBench is published by Zapier, a workflow automation vendor, benchmarking agents on simulated tooling shaped like its own product, and the page carries no conflict-of-interest disclosure. And the cost column is computed at standard list pricing rather than the promotion currently in effect, which Zapier states on the page: "Promotional pricing is available for both Gemini models; Ranking and Cost/task reflect standard list pricing." At today’s promotional rate the same run is roughly half those figures. The ranking is unaffected because both Gemini rows move together.

There is also a second, unrelated evaluation called AutomationBench, run independently by Artificial Analysis with a different task set and partial-credit scoring. Its numbers run roughly twice as high and are not comparable. Gemini 3.7 Flash leads that board too.

Six places the headline price misleads

Thinking tokens are output tokens. Google’s own documentation: "When thinking is turned on, response pricing is the sum of output tokens and thinking tokens." And further: "Pricing is based on the full thought tokens the model needs to generate, despite only the summary being output from the API." You pay $3.75 per million for reasoning you never see, on the more expensive side of the meter. The split itself, and why it flips which workloads are cheap, is covered in input vs output tokens.

Verbosity is a model property, not a setting. 77% more output tokens on the same benchmark, at the same rate, is the whole 57% cost gap.

High effort may buy nothing. Verified above on AutomationBench: identical score, 13% more cost.

The cache floor makes the discount unreachable for short prompts. Implicit caching is on by default and needs no configuration, but the minimum for a cache hit on any Gemini 3.x Flash model is 4,096 tokens. Classification and extraction prompts commonly run 800 to 2,000 tokens. If that describes your workload, the 90% cached-input discount is structurally out of reach, and no amount of prompt restructuring will change it. Check usage.total_cached_tokens rather than assuming.

Thought signatures compound agentic input cost. Gemini 3 function calling requires returning encrypted thought signatures: "When using Gemini 3 models, you must pass back thought signatures during function calling, otherwise you will get a validation error." In sequential tool use, all signatures must be passed back. Every agentic turn therefore replays a growing history plus signature payloads as fresh input. We could not find a published figure quantifying that overhead, which is itself worth knowing.

Latency is a cost that does not appear on the meter. Artificial Analysis measured time to first token at 8.29 seconds for 3.8 Flash at medium, with output around 246 tokens per second, and time per task at 2.5 minutes at high effort. For an interactive CMS feature, that is a product decision rather than a billing line.

The counterexample that nearly breaks the argument

Being honest about this: the cheapest sticker price sometimes does win.

GLM-5.3-Flash has the lowest rate in the comparison set, and it also has the lowest cost per task, $0.18 against Gemini 3.8 Flash’s $0.58, at effectively equal quality on the Artificial Analysis index (46 against 47). It generates more tokens than Gemini does, 160 million against 140 million, and still costs a fifth as much to run, because its rates are far lower.

So the argument is not that cheap sticker prices lie. It is narrower and more useful: the sticker price is not sufficient information, and in the specific case of two models with the same sticker it carries no information at all.

The reason we would still not put GLM-5.3-Flash behind an interactive feature is latency, not cost. It runs at roughly 48.5 output tokens per second against Gemini 3.8 Flash’s 246 to 300, a five- to six-fold wall-clock penalty. That disqualifies it for anything a user is waiting on, and leaves it perfectly viable for overnight batch work. Cheap, fast, and concise are three separate purchases.

Google no longer publishes per-model rate limits

Worth stating plainly because it affects scoping fixed-price work: Google’s rate limits documentation no longer gives per-model requests or tokens per minute. It says limits depend on factors including usage tier and directs you to view them in AI Studio.

What is published is tier qualification and batch enqueued-token limits. Tier 1 requires a linked billing account and carries a $10 rolling ten-minute spend limit and 3 million enqueued batch tokens. Tier 2 requires $100 paid and three days, with a $50 limit and 400 million batch tokens. Tier 3 requires $1,000 paid and thirty days, with a $200 limit and 1 billion batch tokens.

You cannot capacity-plan a production build from public documentation. That is a legitimate thing to raise before quoting a fixed price on a high-volume integration.

January 1 resets everything above

The $0.75 and $3.75 rates are promotional and double on January 1, 2027, to $1.50 and $7.50, with cached input going to $0.15 and cache storage to $1.00 per million per hour. Batch and priority tiers scale with them.

We covered this when Gemini 3.7 Flash shipped and the expiry is unchanged, which is the point: it is the same promotion, on the same date, now covering a fourth model. Any annual retainer priced off the current rate is priced off a number with roughly sixteen weeks left on it.

What to actually do

Separate your workloads before you pick a model. Classification and extraction run short prompts, cannot reach the cache floor, and rarely need reasoning. Use low effort, or stay on 3.7 Flash, and take the batch tier where latency permits, because batch halves everything. Content generation is where batch pricing is the real number rather than the standard rate. Agentic site work is where AutomationBench says 3.7 Flash is at least as good and where thought-signature replay compounds input cost.

Measure cost per completed task, not cost per million tokens. Run a representative sample of your own work through both models, count output tokens, and compare. Two models with the same rate card can differ by 57%, so the rate card is not the measurement.

Check usage.total_cached_tokens in production rather than assuming the cached rate applies.

Test whether high effort earns its 13% on your workload. On at least one relevant benchmark it does not.

Price January into anything annual. Everything in this article doubles on the first of the year, and Google has shipped four Flash models in sixteen weeks, so the model you benchmark today may not be the one you ship on.

Frequently Asked Questions

What does Gemini 3.8 Flash cost?

$0.75 per million input tokens, $3.75 per million output including thinking tokens, and $0.075 for cached input, on standard processing. Batch and Flex are half. Priority is roughly 1.8 times. Those rates double on January 1, 2027.

Is it cheaper than Gemini 3.7 Flash?

The rate card is identical. In practice it is more expensive to run, because it generates substantially more output tokens for the same work. Artificial Analysis measured a 56.7% higher bill for the same benchmark suite.

Why does it cost more at the same price?

Output volume. 140 million output tokens against 79 million on the same evaluation. Google’s own model card acknowledges the model may use more tokens to maximize performance, especially at higher effort levels.

Should I upgrade from 3.7 Flash?

Not automatically. 3.7 Flash remains available with no announced shutdown date, Google still recommends it for efficiency-first workloads, and on Zapier’s AutomationBench it scores slightly higher than 3.8 Flash at slightly lower cost per task. Benchmark your own workload before moving.

Do thinking tokens cost extra?

They are billed at the output rate. There is no separate line item. Google’s pricing table header says output price includes thinking tokens.

Why is my cached-input discount not applying?

The minimum for an implicit cache hit on Gemini 3.x Flash is 4,096 tokens. Prompts shorter than that cannot hit the cache regardless of how repetitive they are. Verify with `usage.total_cached_tokens`.

Is high thinking effort worth paying for?

On Zapier’s AutomationBench, no: medium and high both score 29.68% while high costs 13% more. On other workloads it may be. Test it rather than assuming.

What is the context window?

1,048,576 tokens in and 65,536 out, with no context-length price tiering, which is a real advantage over the Pro tier.

Is Gemini 3.7 Flash being deprecated?

No. Google has announced no shutdown date for 3.7, 3.6, 3.5, or 2.5 Flash. The only Flash-family model with a scheduled shutdown is 3.1 Flash Lite, no earlier than May 2027.

Is AutomationBench independent?

The version cited here is published by Zapier, which sells workflow automation, and the page carries no conflict-of-interest disclosure. A separate evaluation of the same name is run independently by Artificial Analysis with different tasks and partial-credit scoring. The two are not comparable.

Does an EU region cost more?

Yes. Vertex charges 10% more on non-global endpoints, roughly $0.825 in and $4.125 out at current promotional rates.

Where can I find the rate limits for my project?

Not in Google’s documentation, which no longer publishes per-model requests or tokens per minute. They are visible in AI Studio for your account.

Digital Matters

Artificial Intelligence (AI) Desk