Artificial Intelligence (AI)

Qwen3.8-Max Shipped. Three of the Five Things We Asked For Arrived.

Qwen3.8-Max reached general availability with published pricing and a benchmark table, but the activated parameter count and the promised open weights are both still missing.

Qwen3.8-Max reached general availability on 3 August 2026, two weeks after the preview we covered in our look at Qwen3.8. That post ended with a checklist: five specific things that needed to land before the model could be treated as a frontier-class option rather than a promising preview.

Three arrived. Two did not, and they are the two that matter most.

Scoring the checklist

An official announcement with a benchmark table. Arrived. Alibaba published a launch post with a table covering Terminal-Bench 2.1, OSWorld-Verified and a set of multimodal comparisons.

Published API pricing. Arrived, and it is unusually clean. Two dollars per million input tokens and six per million output on the published rate card, flat, with no context-length tiers. Implicit cache reads are $0.25 and explicit cache reads $0.17.

Independent evaluation. Arrived, in one place, and it is the most interesting result of the three.

The activated parameter count. Did not arrive. Alibaba’s model changelog states 2.4 trillion total parameters in a mixture-of-experts design and discloses no active count, no expert count and no per-token figure. There is no technical report.

A weights repository with a licence file. Did not arrive. There is no Qwen3.8 repository on Hugging Face, none in the QwenLM GitHub organisation, and none on ModelScope. Reporting suggests an open-weight release is coming, with no date and no licence named.

A correction to our own coverage. Our 26 July post is titled “Alibaba’s New Open Model.” That was wrong on the day and remains wrong: nothing about Qwen3.8 has been open at any point. Alibaba promised open weights and has not yet delivered them. The body of that post said so; the headline did not.

What "rivalling Anthropic" actually looks like on a leaderboard

The framing in circulation is that Qwen3.8-Max rivals Anthropic on benchmarks. The one independent leaderboard carrying the model says something more specific.

On Arena’s text leaderboard, Qwen3.8-Max sits at rank five with an Elo of 1496, plus or minus 10. The four models above it are Claude Fable 5 at 1509, Claude Opus 4.6 thinking at 1505, Claude Opus 4.7 thinking at 1502 and Claude Opus 4.6 at 1497. Every one of them is Anthropic’s.

So on the only independent measurement available, Qwen3.8-Max is not rivalling Anthropic. It is fifth, behind four Anthropic models, and its margin below the fourth is inside the error bar. That is a genuinely strong result for a model priced at a fraction of the field. It is not the result the coverage describes.

One caveat cuts the other way, and it is worth stating: that leaderboard snapshot predates general availability, so the votes were cast against the preview build.

The Qwen3.8-Max benchmark numbers, and who ran them

Alibaba’s headline figures are 86.6 on Terminal-Bench 2.1 and 86.1 on OSWorld-Verified. Both are self-reported.

Qwen3.8-Max does not appear on the official Terminal-Bench 2.1 leaderboard, where every listed entry is marked verified and the top score is Claude Code paired with Fable 5 at 83.8 percent. A self-reported 86.6 would lead that board by 2.8 points over anything its maintainers have confirmed. No Qwen model appears on it at all. This is the same pattern we found in poolside’s Laguna S 2.1, where a self-run number on a custom harness was widely reported as though it were a leaderboard result.

Two details in Alibaba’s own table deserve more attention than they have had.

The first is that Alibaba is not claiming the top score. Its table puts GPT-5.6 Sol at 88.8 on Terminal-Bench 2.1, ahead of Qwen3.8-Max’s 86.6. A company presenting itself as the new leader does not usually publish a competitor beating it.

The second is the harness. Most of the coding rows were run on the Claude Code scaffold, which is a competitor’s agent harness. That is a reasonable choice for comparability, and it also means the figures measure a model-plus-harness pairing rather than the model alone.

The number nobody outside Alibaba can compute

Two point four trillion parameters is a headline, not a specification, and without the activated count it cannot be turned into anything useful.

In a mixture-of-experts model, total parameters set the memory footprint and activated parameters set the arithmetic per token. Qwen’s own earlier releases show how wide that gap runs: Qwen3-235B-A22B activates 22 billion of 235 billion, and Qwen3-30B-A3B activates roughly 3 billion of 30. Apply that ratio loosely and a 2.4-trillion-parameter Max tier could be activating anywhere from 50 billion to 200 billion per token, which is the difference between two quite different products.

This is the same distinction we worked through in what self-hosting Kimi K3 actually costs: sparsity reduces the arithmetic per token, not the memory needed to hold the model. Without the active count, nobody outside Alibaba can model serving cost, compare efficiency against DeepSeek or Kimi, or judge whether $2 per million input tokens reflects genuine architectural efficiency or a price set for market share.

What the pricing actually says

The rate card repays a close read, because two things in it are not conversions of each other.

Alibaba’s international list is $2.00 input and $6.00 output per million tokens. Its China list is ¥12 and ¥36. At prevailing rates that Chinese price is roughly $1.67 and $5.00, so the international price sits about twenty percent above it. These are separately set price lists, not one number expressed in two currencies.

More useful for anyone forming an impression of cost: the preview ran at a billing coefficient of 0.05 during the day and 0.01 overnight. Nobody has yet paid list price at scale, so early reports of how cheap the model is were measuring a promotion.

The context window is a million tokens with a 131,000-token output ceiling, and unlike the previous generation’s Plus tier the pricing is not banded by context length. Whether that million tokens is natively pretrained or extended from a shorter base is not stated anywhere, and with no public weights there is no configuration file to check. Treat it as unknown rather than assuming either answer, a distinction that turned out to matter a great deal for Laguna S 2.1.

Where this leaves a buyer

The honest summary is that Qwen3.8-Max is now a real, purchasable, competitively priced frontier-adjacent model with one independent result that is genuinely good, and two disclosure gaps that should keep it off a migration plan.

If you are running evaluations, it belongs in them. A fifth-place Arena finish at this price is worth an afternoon of your own testing, and the pricing is simple enough to model without a spreadsheet.

If you are planning capacity or comparing efficiency across the Chinese open-weight cohort we mapped in open versus closed AI models, you cannot, because the activated count is missing and it is the input every such comparison needs.

And if the open weights are the reason you are interested, wait for the repository and read the licence file rather than the announcement. Qwen’s smaller dense models have shipped under Apache 2.0 for years, but the Max-tier flagships have historically shipped closed, so an open Qwen3.8-Max would be a break in the pattern rather than a continuation. Until a repository exists with a licence in it, the commitment is a sentence.

There is one more thing worth watching. As of this writing the English-language Model Studio documentation still lists only the previous generation, and the model has not appeared on OpenRouter. The Chinese documentation carries it with Singapore, Tokyo, Frankfurt and Virginia regions. The international rollout is running behind the announcement, which matters if your procurement depends on a specific region.

Frequently Asked Questions

What is Qwen3.8-Max?

It is the flagship tier of Alibaba’s Qwen3.8 family, generally available since 3 August 2026 after a preview that began 19 July. Alibaba describes it as a 2.4-trillion-parameter mixture-of-experts model with a one-million-token context window and native vision-language support. It is API-only today, served through Alibaba’s own platforms.

How much does Qwen3.8-Max cost?

Two dollars per million input tokens and six dollars per million output on the international price list, flat with no context-length tiers. Implicit cache reads are $0.25 per million and explicit cache reads $0.17. The China price list is separately set at ¥12 and ¥36, roughly twenty percent below the international rate rather than a direct conversion.

Can I download the Qwen3.8-Max weights?

Not yet. There is no repository on Hugging Face, ModelScope or the QwenLM GitHub organisation, and no licence has been published. Alibaba has said weights are coming without naming a date or a licence. Smaller Qwen models have long shipped under Apache 2.0, but the Max-tier flagships have historically shipped closed, so this would be a change rather than a continuation.

Is Qwen3.8-Max better than Claude or GPT?

On the only independent leaderboard carrying it, no. Arena places it fifth on text with an Elo of 1496, behind Claude Fable 5 and three Claude Opus variants. Alibaba’s own benchmark table also shows GPT-5.6 Sol ahead of it on Terminal-Bench 2.1. What it does offer is a competitive result at a substantially lower price.

Are the benchmark scores independently verified?

No. The 86.6 on Terminal-Bench 2.1 and 86.1 on OSWorld-Verified are vendor-reported. Qwen3.8-Max does not appear on the official Terminal-Bench 2.1 leaderboard, where every entry is marked verified and the top confirmed score is 83.8 percent. Most of the coding rows were also run on the Claude Code harness, so they measure a model-and-scaffold pairing rather than the model alone.

Why does the activated parameter count matter so much?

Because in a mixture-of-experts model the total parameter count sets memory requirements while the activated count sets the arithmetic performed per token, and only the second predicts serving cost. Qwen’s own earlier models activated 22 billion of 235 billion, and roughly 3 billion of 30 billion. Without that figure for Qwen3.8-Max, 2.4 trillion total is an unfalsifiable headline rather than a specification.

Is the one-million-token context window real?

It is genuinely served rather than a paper figure, with a maximum output of 131,000 tokens. What is not stated anywhere is whether the window was natively pretrained or extended from a shorter base using a technique such as YaRN, and with no public weights there is no configuration file for anyone outside Alibaba to check. Treat the provenance as unknown.

Should we move production workloads to it?

Not yet, though it earns a place in your evaluation set. The pricing is simple, the independent result is real, and testing it on your own workload costs an afternoon. Two things argue against migration: the missing activated count means you cannot model what it will cost you at scale, and the international rollout is currently behind the announcement, with the English documentation and major aggregators not yet carrying it.

Digital Matters

Artificial Intelligence (AI) Desk