Between July 21 and August 5, 2026, four separate labs published a coding benchmark figure that the technology press reprinted as though it were a position on a leaderboard. None of those four numbers appears on the leaderboard in question. This is not a story about dishonesty, and the accusation framing gets it wrong. It is a story about a measurement pipeline with no verification step at the vendor end, and vendor benchmark scores are what comes out of that pipeline: real numbers, produced in good faith, that do not carry the meaning the coverage assigns them.
The fortnight in question
Four releases, four self-run figures on the same benchmark, Terminal-Bench 2.1.
Poolside published Laguna S 2.1 on July 21 with a Terminal-Bench 2.1 result of 70.2% with thinking mode enabled. DeepSeek’s API changelog records V4-Flash-0731 arriving on July 31, 2026 with a Terminal Bench 2.1 figure of 82.7. Alibaba’s model changelog dates Qwen3.8-Max to August 3, 2026, and its launch table reported 86.6 on the same benchmark, with GPT-5.6 Sol listed above it at 88.8, as reported at launch. Meta shipped Muse Code and Muse Spark 1.2 on August 5, 2026, claiming 82.9%.
We covered each of these individually: poolside’s Laguna S 2.1, Qwen3.8-Max, DeepSeek V4-Flash-0731 and Meta’s Muse Code. Taken one at a time, each looked like a footnote. Four in fifteen days is a pattern worth naming.
What the leaderboard actually contains
The official Terminal-Bench 2.1 leaderboard lists 17 entries. Every one carries the same note: "A Terminal-Bench team member ran the evaluation and verified the results." Every entry ran five trials. Every entry carries a confidence interval, ranging from plus or minus 1.1 to plus or minus 1.7 points. The submission rule is short: "submissions may not modify timeouts or resources."
The top of the board is Claude Code paired with Fable 5 at 83.8%, then Codex with GPT-5.5 at 83.1%, then Terminus 2 with Fable 5 at 80.4%. Terminal-Bench itself is described as "a stanford x laude collaboration."
Now place the four claims against it. Poolside has no entry. DeepSeek has no entry, and no DeepSeek model appears anywhere on the board. No Qwen model appears either. Meta appears once, but not with the model it was promoting: Muse Spark 1.1 sits eighth at 76.2%, plus or minus 1.2, run on the mini-SWE-agent harness. Muse Spark 1.2 is absent.
That last row is the most useful fact in this entire story, because it is the one case where a verified number and a vendor number describe the same model. Meta’s model page lists Muse Spark 1.1 at 80.0. The verified figure is 76.2. The gap is 3.8 points, and 80.0 sits 2.6 points above the top of the verified confidence interval.
The harness moves the score more than the model does
This is the part that makes vendor benchmark scores non-comparable, and the leaderboard proves it without needing any outside evidence. The board lists several models twice, under different agent harnesses. Same model weights, same benchmark, same trial count, same verification. Only the scaffold differs.
Gemini 3 Pro scores 73.9% under Terminus 2 and 65.8% under Gemini CLI. That is a spread of 8.1 points from harness choice alone. GPT-5.5 scores 83.1% under Codex and 78.0% under Terminus 2, a spread of 5.1 points. Fable 5 scores 83.8% under Claude Code and 80.4% under Terminus 2, a spread of 3.4 points. Opus 4.7 scores 68.9% under Claude Code and 66.1% under Terminus 2, 2.8 points apart. Gemini 3.1 Pro is the outlier in the other direction, at 65.8% and 65.6% under two harnesses, a spread of 0.2 points.
Hold that 8.1 against the shape of the board. The distance from the top verified entry, 83.8%, to the eighth entry, 76.2%, is 7.6 points. Harness choice alone moved one model further than the distance from first place to eighth. A vendor that picks its own scaffold is not reporting a model score. It is reporting a model-and-scaffold score, and the scaffold is a free variable worth more than most model generations.
If you are choosing between coding agents rather than raw models, the pairing is the product, which is why comparing Claude Code against OpenAI Codex is a different exercise from comparing the models underneath.
Nobody is lying, and their own footnotes prove it
Read the vendor disclosures and the accusation collapses. Poolside states plainly: "All Laguna S 2.1 agentic benchmarking was completed using our internal fork of the Laude Institute’s Harbor Framework with our agent harness, a maximum of 500 steps and sandboxed execution via our internal sandbox service." On a separate benchmark it goes further, noting that Laguna S 2.1 "scored 40.4% in our agent harness, pool, not mini-swe-agent which DeepSWE’s leaderboard uses." That is a company telling you, unprompted, that its harness is not the leaderboard’s harness.
DeepSeek is equally explicit. Its changelog records that "the official DeepSeek-V4-Flash was tested using the DeepSeek Harness minimal mode (to be released soon) as the framework, with the max effort level, topp=0.95, and temperature=1.0." Note the parenthesis. The harness that produced 82.7 was not public at the time the figure was.
Meta’s methodology page is the most detailed of the four. It names the harnesses directly, "Muse Code for Muse Spark 1.2" and "mini-swe-agent for Muse Spark 1.1", confirms the evaluation "uses all 89 tasks in the official 2.1 release", reports "the average task success rate (pass@1) across five attempts", runs "the maximum available reasoning strength for each model", and places each attempt "in an isolated Daytona sandbox". It also carries this warning: "Our evaluation setup (e.g., agent tools and system prompts) may not be specifically tuned for proprietary third-party models. Therefore, the results may not reflect these models’ best performance when used in environments tailored to their specific strengths." On a different benchmark it states outright that its evaluation "is not harness-identical to the leaderboard."
These are not the disclosures of companies trying to deceive anyone. They are the disclosures of companies operating in a system where nothing requires them to submit for verification and nothing stops a footnote from being dropped in transit.
What goes missing from vendor benchmark scores
The variables that move a score by several points are the ones that survive least well when a figure is lifted into a headline.
Trial count. The leaderboard runs five trials for every entry. Poolside ran four, listing "Terminal-Bench 2.1: mean pass@1 averaged over 4 attempts per task" and three attempts on several other benchmarks. DeepSeek’s changelog does not state a trial count at all. Meta ran five.
Harness. Four vendors, four harnesses, one of them unreleased.
Reasoning effort. Poolside’s 70.2 is a thinking-mode figure. DeepSeek used "the max effort level". Meta used "xhigh reasoning effort". These are not free at inference time, and the score you read is the score at the most expensive setting.
Sampling parameters. DeepSeek is the only one of the four to publish them.
Confidence intervals. Every verified entry has one. None of the four vendor figures does. When the intervals on the board run to plus or minus 1.7 points, a vendor claim quoted to one decimal place is projecting a precision nobody has demonstrated. The top two verified entries, 83.8% and 83.1%, are separated by 0.7 points against intervals of plus or minus 1.2 and plus or minus 1.1. They overlap. They are, on the published evidence, tied.
Why the press turns a figure into a rank
The mechanism is mundane. A number looks like a number. Terminal-Bench 2.1 is a named, versioned benchmark with a public leaderboard, so "82.7 on Terminal-Bench 2.1" reads as a position on that board rather than as an unrelated measurement that happens to share a name. Checking requires opening the leaderboard, scanning 17 rows for a model that is not there, and then explaining an absence, which is harder to write than a ranking.
There is also a supply problem. The Terminal-Bench documentation notes that "A new leaderboard submission process is coming soon." Verification takes a maintainer’s time, and launch days do not wait for it. The result is a permanent lag in which the only available figure is the vendor’s, and the one case we can check came in 3.8 points high.
A checklist for reading any vendor benchmark claim
Five questions, each answerable in a minute.
- Is the model on the leaderboard? Not the vendor, the specific model and version. Muse Spark 1.1 being listed tells you nothing about 1.2.
- What harness produced the number? If it is the vendor’s own, or an internal fork, the figure measures a pairing. If the harness is unreleased, the figure is unreproducible by anyone outside the company.
- How many trials, and is there a confidence interval? A single decimal with no interval and no trial count is a point estimate dressed as a measurement.
- What reasoning effort and sampling settings? Maximum-effort figures are the ceiling, not the default you will be billed for.
- Does the vendor’s own table show a competitor ahead? Alibaba’s did, listing GPT-5.6 Sol at 88.8 above its own 86.6. Publishing a loss is a signal of good faith, and a reminder that you are reading a marketing artifact, not a referee’s scorecard.
Treat unverified vendor benchmark scores as what they are: self-reported measurements taken under conditions the vendor chose, not reproduced by anyone else. That is the correct default until a maintainer runs the evaluation.
Frequently Asked Questions
Are these four vendors accused of cheating?
No. All four published methodology notes disclosing their setup, and poolside and Meta both stated explicitly that their harness differs from the relevant leaderboard’s. The problem is structural: there is no requirement to submit for verification, and the disclosures rarely survive being reprinted.
How much can an agent harness change a benchmark score?
On the verified Terminal-Bench 2.1 leaderboard, Gemini 3 Pro scores 73.9% under Terminus 2 and 65.8% under Gemini CLI, a spread of 8.1 points on identical model weights. GPT-5.5 spans 5.1 points across two harnesses. Some pairings barely move: Gemini 3.1 Pro differs by 0.2 points.
Which of the four models appears on the official leaderboard?
Only Meta, and only with the previous version. Muse Spark 1.1 sits eighth at 76.2% under mini-SWE-agent. Poolside, DeepSeek and Alibaba have no model on the board in any version.
What was the gap between Meta’s claim and the verified number?
Meta published 80.0 for Muse Spark 1.1. The verified figure for the same model is 76.2%, plus or minus 1.2: a 3.8-point gap, with the vendor claim 2.6 points above the top of the interval.
Does a self-run figure make a benchmark claim worthless?
No, it makes it non-comparable. Self-run vendor benchmark scores are still evidence about the model under the vendor’s own conditions, which may be the conditions you will actually use. They just cannot be placed on the same axis as a verified entry produced under a different harness and trial count.
Why do confidence intervals matter so much here?
Because the verified intervals run from plus or minus 1.1 to plus or minus 1.7 points, differences smaller than about three points are frequently not differences at all. The top two verified entries, at 83.8% and 83.1%, have overlapping intervals and are effectively tied.
Can I reproduce a vendor number myself?
Sometimes. Terminal-Bench runs through the Harbor framework, and the documentation covers running the dataset locally with a built-in or custom agent. But DeepSeek’s 82.7 was produced with a harness described in its own changelog as “to be released soon”, so that specific figure was not reproducible by outsiders when it was published.
What single check catches most of these claims?
Open the leaderboard and search for the exact model version. In this fortnight that step would have caught all four, because three models were absent entirely and the fourth was listed only in a previous version, at a lower score than the vendor published.