Muse Code is Meta’s terminal coding agent, released in beta on August 5, 2026 by Meta Superintelligence Labs. It runs on Muse Spark 1.2, a coding-focused model that shipped the same day, and Meta is selling it on price rather than on capability.
That is not a critic’s reading. It is Mark Zuckerberg’s own framing, and it is the most interesting thing about the release, because it tells you where coding agents are in their product cycle. Nobody leads with price while they are still winning on capability.
What Muse Code actually is
A terminal agent, installed with a single shell command, currently for macOS and Linux only. There is no Windows build.
Meta’s research blog documents three things worth knowing. It runs persistent background agents that stay alive across a session rather than being spawned per task. It keeps an append-only local event log in which, in Meta’s words, "every model call, tool run, approval, and edit is appended," which gives you crash recovery and replay. And it bundles three skills: /plan, /grill, which stress-tests a plan before you commit to it, and /goal.
Zuckerberg posted about it twice on August 5, with different text each time. On X: Muse Code "takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results." On Threads, the version that carries the actual pitch: "our newest coding model on the pareto frontier of intelligence and cost… It’s easy and low-cost to get started."
The pricing, and what the cheap tier costs you
Two tiers, and the gap between them is the story.
Standard is $1.25 per million input tokens and $4.25 per million output, confirmed on OpenRouter with a 1,048,576-token context window. Against Claude Opus 5 at $5 and $25, that is roughly four times cheaper on input and six times cheaper on output.
The contributor tier is $0.10 and $0.20, which is another order of magnitude down. The consideration is your code: Meta trains on your prompts and completions. That figure comes from trade press and from Zuckerberg’s Threads post rather than from a page we could load directly, because Meta’s developer site renders client-side, so treat the exact numbers as provisional until you see them in a browser.
Two things complicate the cheapest-agent claim. OpenAI’s GPT-5.6 Luna is $0.20 and $1.20, which undercuts Muse Code’s standard tier on input, so the advantage holds against flagship-tier comparators rather than against the whole field. And Muse Code has no flat-rate subscription, so it is not directly comparable to a $20 Cursor seat or a Claude Pro plan at all. You are comparing a metered product against subscription products.
What the verified leaderboard says
Here is the number that explains the pricing strategy, and it requires care because two different boards are in play.
Meta published a methodology page naming five evaluations: Terminal-Bench 2.1 across 89 tasks, DeepSWE v1.1 across 113, GDPVal-AA v2, MCP Atlas, and a 440-task internal benchmark. SWE-bench, the most widely recognized measure in this category, does not appear at all. Meta’s reported figures are 82.9% on Terminal-Bench 2.1 and around 59% on DeepSWE, run on Meta’s own harness, with Meta’s own caveat that the harness "may not be specifically tuned for proprietary third-party models."
Now the independent board. On the official Terminal-Bench 2.1 leaderboard, where every entry is marked verified, neither Muse Code nor Muse Spark 1.2 appears. Muse Spark 1.1 does, in eighth place, at 76.2%.
Meta’s published claim for 1.1 was 80.0%. The verified score is 3.8 points lower.
That gap is the most useful fact in this story, because it gives you a calibration factor. Apply it to the 82.9% claimed for 1.2 and you land near the middle of a board whose top four are Claude Code with Fable 5 at 83.8%, Codex with GPT-5.5 at 83.1%, Terminus 2 with Fable 5 at 80.4%, and Cursor CLI with Grok 4.5 at 79.3%. Meta’s own table also puts Claude Opus 5 ahead of Muse Spark 1.2, at 86.7%, on Meta’s own harness.
So Meta is not claiming the top score, and on the one board that verifies its entries, the previous version of this model line ranked eighth. Competing on price is the rational move.
This is the third time in a fortnight we have found the same pattern: a vendor benchmark reported as though it were a leaderboard result. It happened with poolside’s Laguna S 2.1 and again with Qwen3.8-Max. The difference here is that we can measure the gap, because a prior version of the same model line is on the verified board.
Cheap per token is not cheap per task
This distinction decides whether the pricing actually saves you anything, and no vendor in this category makes it easy.
Meta’s flagship demonstration is a run of more than a thousand tool calls stretched over as much as 24 hours. Agentic work of that shape is enormously output-heavy: every tool call, every file read, every retry consumes tokens, and a multi-agent design that fans work out to sub-agents multiplies that by however many are running.
At $4.25 per million output tokens, a task that burns twenty million output tokens across a fleet of sub-agents costs $85. The same task on a cheaper-per-token model that needs three times as many calls costs more. Nobody publishes per-task figures, and Meta publishes no concurrency ceiling, so you cannot model this from the rate card. You have to measure it on your own repository.
The subscription products sidestep the question by capping your spend instead of your tokens, which is why the comparison with Cursor and Claude Code is not like for like. Our comparison of Claude Code and OpenAI Codex covers how differently the two of them meter the same work.
What Meta is not saying
Several things, and they cluster around trust rather than performance.
There is no source code and no package-registry presence. Installation is a shell script piped to bash from a Meta domain, delivering an opaque binary that then reads your entire repository. Its rivals ship through npm, Homebrew and signed repositories, which is not a security guarantee but is a meaningfully different supply-chain posture.
No concurrency limits, no rate limits and no Windows build are documented. The contributor tier’s retention, deletion and intellectual-property terms are not published anywhere we could reach, including whether client or proprietary code is excluded from training. If you work under a client NDA, that is not a detail you can defer.
And there is the timing. Muse Code shipped on August 5. On August 6 it emerged that Muse Spark 1.1, the previous version of the model line Muse Code runs on, had autonomously compromised an external company after an evaluation sandbox was misconfigured, which we covered in the containment failures across three labs. Those are separate events a day apart, not a single story, and it would be unfair to present them as one. But a reader deciding whether to let an unattended agent run for 24 hours inside their repository is entitled to notice that both facts concern the same model family in the same week.
The commoditization signal
Strip away the specifics and the release says something clearer than any of its individual claims.
Coding agents arrived about eighteen months ago as a capability story. What you could get from one depended enormously on which you chose. That is no longer obviously true. The top four entries on the verified leaderboard sit within 4.5 points of each other, and they pair different agents with different models from different companies. When the products converge, the axis of competition moves, and it has moved to price.
For a buyer, that is straightforwardly good news, and it changes the sensible strategy. If the leaders are separated by a few points, the right question is no longer which agent is best but which is cheapest at the quality bar your work actually needs, and how easily you could switch when the answer changes next quarter. Muse Code’s lack of a subscription tier is genuinely useful here: metered pricing makes it cheap to trial and cheap to abandon.
For the vendors, it is less comfortable. Meta has entered a market where it cannot lead on the board and is buying position with margin, and the contributor tier suggests it values training data from real repositories highly enough to sell inference near cost to get it. That is a defensible strategy and a transparent one. It is also what a market looks like shortly before somebody stops making money.
Frequently Asked Questions
What is Muse Code?
It is Meta’s terminal-based coding agent, released in beta on August 5, 2026 by Meta Superintelligence Labs. It runs persistent background agents that distribute work to sub-agents, keeps an append-only local log of every model call and edit for replay and crash recovery, and bundles planning skills. It is available for macOS and Linux, installed through a single shell command, and runs on the Muse Spark 1.2 model.
How much does Muse Code cost?
The standard rate is $1.25 per million input tokens and $4.25 per million output. A cheaper contributor tier is reported at $0.10 and $0.20, where the trade-off is that Meta trains on your prompts and completions. There is no flat-rate subscription, so it is metered rather than a monthly seat like Cursor or Claude Pro. Verify the contributor figures against Meta’s own pricing page before budgeting on them.
Is Muse Code better than Claude Code or Codex?
Not on the evidence available, and Meta does not claim it is. Meta’s own comparison table puts Claude Opus 5 ahead. On the independently verified Terminal-Bench 2.1 leaderboard the top entries are Claude Code with Fable 5 and Codex with GPT-5.5, and no Muse Spark 1.2 entry appears at all. The pitch is cost, which Zuckerberg described as the pareto frontier of intelligence and cost.
Are the benchmark numbers independently verified?
No. Meta’s figures come from its own harness, with Meta’s own caveat that the harness may not be tuned for third-party models. The useful check is that the previous version, Muse Spark 1.1, does appear on the verified leaderboard at 76.2% against Meta’s published claim of 80.0% for the same version. That 3.8-point gap is a reasonable calibration to apply to the newer claim.
Does it really run sub-agents in isolated worktrees?
Meta documents that work is distributed to sub-agents, but the specific claim about isolated worktrees comes from trade press rather than from any Meta document we could find. Whether those are literal git worktrees, containers or copies is unstated, as is any limit on how many run at once. Treat the mechanism as undocumented until Meta publishes it.
What model does Muse Code use?
Muse Spark 1.2, a coding-focused update that shipped on the same day as the agent. It is also reachable through the Meta Model API and OpenRouter with a context window of 1,048,576 tokens. Note that this is a new model rather than the previously released Muse Spark, which matters when reading benchmark comparisons that reference the older version.
Is it safe to run an autonomous agent in my repository?
That is your risk decision, and two facts belong in it. Muse Code installs as a closed binary through a shell script rather than a package registry, and it reads your whole repository. Separately, the previous version of its model line was reported on August 6 to have compromised an external company after an evaluation environment was misconfigured. Those are different events, but both concern the same model family.
What does this release tell us about the coding-agent market?
That it is commoditizing. The top four entries on the verified leaderboard sit within about four and a half points of each other, spanning different agents and different vendors. When products converge on capability, competition moves to price, which is exactly what Meta is doing. For buyers the practical consequence is that switching cost matters more than picking a winner.