IT Infrastructure

Kimi K3 Hosting: What Renting the Open Frontier Model Actually Costs

Kimi K3 hosting compared across day-zero providers, where the per-token price matches the closed APIs it competes with and the real saving appears in cost per completed task rather than cost per token.

Kimi K3 hosting became a real option the moment Moonshot published the weights on 26 July, because four infrastructure providers stood the model up the same day. Together AI, Modal, Fireworks, and Runpod all shipped day-zero endpoints. For most teams that is the question that actually matters, because self-hosting the model means holding roughly 1.56TB of weights on Blackwell-class hardware, which is not a purchase most companies are going to make. Renting is the realistic path.

So here is the reported comparison: what the day-zero providers charge, how that lands against the closed API you are probably already paying for, and where the saving is real rather than rhetorical. The short version is that the answer surprised us.

The price that did not drop

The expectation baked into most coverage of open-weight releases is that they crash the price. Weights are free, therefore tokens get cheap, therefore the closed labs have a problem.

That is not what happened here. Moonshot’s first-party API prices K3 at $3.00 per million input tokens, $0.30 on a cache hit, and $15.00 per million output tokens. Together AI lists exactly the same numbers. The day-zero hosts did not undercut the model’s originator, they matched it.

More striking is the comparison to Moonshot’s own previous generation. K2.6 sold output at $4 per million. K3 sells it at $15, a near quadrupling. Moonshot priced this model like a frontier product, because on the independent leaderboards it is one, and the hosting partners took the same view. If you arrived expecting an open-weight discount at the sticker level, there isn’t one.

There is a second, less comfortable comparison. Against other open-weight models, K3 is the expensive option. Artificial Analysis measures an average cost per task on its Intelligence Index of $0.94 for K3, against $0.32 for GLM-5.2 and $0.04 for DeepSeek V4 Pro. If your goal is simply to spend less by moving to an open model, K3 is not the model that does it. It competes with the frontier, and it is priced accordingly.

Where the saving actually shows up

The saving is real, but it lives one level up from the token price.

On that same Artificial Analysis measure, K3 costs $0.94 per task against $1.04 for GPT-5.6 Sol and $1.80 for Claude Opus 4.8. Roughly half the cost of Opus for comparable measured intelligence, at an index score of 57 that places it third overall behind Claude Fable 5 and GPT-5.6 Sol.

Fireworks published a head-to-head against Opus 5 on its own infrastructure, and it is worth reading carefully because it is a vendor benchmark with a vendor’s incentives. On a 480-task software engineering set, K3 scored 92.7 percent accuracy at $0.52 per task against Opus 5 at 94.8 percent and $1.05. On an 83-task terminal set, K3 managed 81.9 percent at $0.35 against 85.5 percent at $1.61. On algorithmic problems both models scored 88.0 percent, at $0.064 and $0.176 respectively.

Read those numbers honestly and they say something more specific than "cheaper." K3 is slightly less accurate and substantially less expensive. On the software engineering set it also took considerably more turns to get there, 55.6 against 37.9, and still cost half as much. That is the actual shape of the trade: you are buying a small accuracy concession in exchange for a large cost reduction, and whether that is a good deal depends entirely on whether your workload tolerates a retry.

Part of what makes this work is token efficiency. Artificial Analysis found K3 used 21 percent fewer output tokens than K2.6 to complete the same nine evaluations, roughly 132 million against 166 million, while scoring higher. Since output is the expensive side of every bill, a model that reaches the answer in fewer tokens can cost less per finished job even at a higher price per token. Our explainer on input versus output tokens covers why that asymmetry drives most real invoices.

What each provider is actually selling

The four day-zero hosts converged on price and diverged on everything else, which is where the decision actually gets made.

Together AI offers the most conventional package: serverless, dedicated, and provisioned-throughput deployment against a 99.9 percent SLA, on US-based infrastructure with zero data retention. It is the option that looks most like the enterprise API you are already buying.

Modal did the most interesting engineering. It trained a custom DFlash speculator tuned to K3’s architecture, which lifted interactivity from roughly 100 tokens per second to 460, and per-GPU throughput from 800,000 to 1.5 million tokens per minute. Speculative decoding predicts several tokens ahead and verifies them in a batch, which helps most on exactly the long generation-heavy agentic work K3 is built for. Modal also includes $30 of free compute monthly, which makes evaluation genuinely free.

Fireworks leaned on compliance and customisation: US-only serverless endpoints with zero data retention, a batch mode at 50 percent off for work that can wait, and a private preview of serverless fine-tuning billed per token rather than per reserved GPU hour.

Runpod covers both ends, with a per-token public endpoint for spiky traffic and Blackwell capacity if you later decide to run the weights yourself.

The jurisdiction question, answered properly

There is an objection that comes up in every conversation about Chinese models, and it deserves a direct answer rather than a shrug: sending your data to an API operated in China is a governance problem for a lot of organizations, and for some it is disqualifying.

Open weights dissolve that specific problem in a way a closed model never can. K3 is a Chinese model, but Together AI, Fireworks, and Runpod are American companies serving it from US infrastructure, and both Together and Fireworks offer zero data retention as a default rather than an upgrade. You can run the model without your prompts crossing a jurisdiction you are uncomfortable with, because the artefact is portable in a way an API endpoint is not.

That portability is the structural benefit people usually miss when they focus on price. Four vendors serve identical weights, so switching providers is a change of base URL rather than a migration. If one host raises prices, degrades, or disappears, the model does not go with it. No closed API offers that, at any price, because the model and the vendor are the same thing.

The reasoning-token trap

One practical warning, because it will show up on your first invoice rather than in the documentation.

K3 runs with reasoning always on, and those reasoning tokens count as output. They count against your max_tokens limit and they count against your bill. Teams porting a prompt from a non-reasoning model routinely set a modest token ceiling, watch the model consume the entire budget thinking, and get back an empty answer. At $15 per million output tokens, an always-on reasoning stream is not a rounding error either.

Set a generous ceiling, read the reasoning content separately from the answer, and measure your real output volume on representative traffic before you extrapolate a monthly figure from a handful of test calls.

What the leaderboards are not telling you

Two caveats belong on any honest recommendation here.

The first is hallucination. Artificial Analysis found K3’s accuracy improved substantially over K2.6, from 33 to 46 percent on its knowledge evaluation, while its hallucination rate went the wrong way, rising from 39 to 51 percent. The model knows more and also asserts more confidently when it does not know. For agentic work that touches production systems, that combination deserves real scrutiny and a verification layer.

The second is simply time. These weights are two days old. Every figure above comes from vendor benchmarks or from evaluation harnesses run in the first hours after release, and none of it constitutes production experience. Moonshot itself had to pause new subscriptions when demand outran capacity. Leaderboard position is not the same as reliability under sustained load, and nobody has that data yet.

Worth noting too, if you are considering building on this rather than just calling it: the Kimi K3 License is derived from MIT and permissive for almost everyone, but it requires a separate agreement if you run a model-as-a-service business above roughly $20 million in annual revenue, and attribution in your interface above 100 million monthly active users.

How to decide

The framing that survives contact with a spreadsheet is this. Renting K3 makes sense if you are currently paying frontier rates for work that does not need frontier accuracy, and your workload can absorb an occasional retry. That is a genuinely large category, and it is where the money is.

It does not make sense if you simply want a cheaper model, because GLM-5.2 and DeepSeek V4 Pro are dramatically cheaper per task. It does not make sense for work where a two-point accuracy gap carries real cost. And self-hosting remains a different conversation entirely, justified by data sovereignty or fine-tuning rather than by price.

The broader consequence is the one Nathan Lambert has been making at Interconnects, and it is about margins rather than capability. Every closed lab now has a credible free substitute one rung below it on the same public leaderboards. That compresses what the frontier labs can charge for mid-tier work, which reduces the profit available to reinvest and the terminal value the market assigns them. Our piece on managing AI investments in the agentic era covers why that pricing pressure matters more to most buyers than any individual benchmark result.

Frequently Asked Questions

How much does Kimi K3 hosting cost?

Moonshot’s first-party API charges $3.00 per million input tokens, $0.30 on a cache hit, and $15.00 per million output tokens. Together AI lists identical pricing, and the other day-zero hosts are broadly comparable. Artificial Analysis calculates a blended rate of $2.31 per million tokens at a typical agentic cache-input-output ratio. The day-zero providers matched Moonshot’s pricing rather than undercutting it.

Is Kimi K3 cheaper than Claude or GPT?

Not per token, but yes per completed task. Artificial Analysis measures $0.94 per task for K3 against $1.04 for GPT-5.6 Sol and $1.80 for Claude Opus 4.8. Fireworks’ own head-to-head against Opus 5 showed roughly half the cost at slightly lower accuracy across software engineering, terminal, and algorithmic tasks. The saving comes from token efficiency and task pricing, not from a lower sticker rate.

Which providers host Kimi K3?

Together AI, Modal, Fireworks, and Runpod all launched day-zero endpoints alongside the 26 July weight release. Together offers serverless, dedicated, and provisioned throughput on a 99.9 percent SLA. Modal runs a custom speculator that reaches 460 tokens per second. Fireworks adds US-only endpoints, a discounted batch mode, and serverless fine-tuning in preview. Runpod provides both a per-token endpoint and Blackwell capacity for self-hosting.

Can I use a Chinese model without sending data to China?

Yes, and this is one of the strongest practical arguments for open weights. Together AI, Fireworks, and Runpod are American companies serving K3 from US infrastructure, and Together and Fireworks both offer zero data retention by default. Because the weights are portable, the model’s origin and the jurisdiction of the infrastructure serving it are separate questions. That separation is impossible with a closed API.

Why did my Kimi K3 bill come in higher than expected?

Most likely the reasoning tokens. K3 runs with reasoning always on, and those tokens are billed as output at $15 per million as well as counting against your `max_tokens` limit. Teams that port a token ceiling over from a non-reasoning model frequently find the budget consumed by thinking, returning an empty answer and a real charge. Measure output volume on representative traffic before projecting monthly spend.

Should I rent Kimi K3 or self-host it?

Rent, unless you have a specific reason not to. The published weights are roughly 1.56TB and require Blackwell-class hardware, which puts a realistic single-node deployment out of reach for most organizations on cost alone. Self-hosting is justified by data-sovereignty requirements that rule out any third party, or by fine-tuning that genuinely differentiates your product. Price alone almost never justifies it.

Is Kimi K3 accurate enough for production work?

It depends on your tolerance for retries and your verification layer. It ranks third on the Artificial Analysis Intelligence Index at a score of 57, behind Claude Fable 5 and GPT-5.6 Sol, and slightly behind Opus 5 on vendor-run coding benchmarks. The caveat that deserves the most weight is its hallucination rate, which rose from 39 to 51 percent versus K2.6 even as accuracy improved. Confident wrong answers are the failure mode to design around.

Are there licensing restrictions on Kimi K3?

For most users, effectively none. The Kimi K3 License is derived from MIT and permits commercial use, modification, and redistribution at no license cost. Two conditions apply at scale: a model-as-a-service business above roughly $20 million in revenue over any twelve-month period needs a separate agreement with Moonshot, and any product above 100 million monthly active users or $20 million in monthly revenue must display “Kimi K3” in its interface.

Digital Matters

IT Infrastructure Desk