Two comparison cards side by side: Kimi K3 from Moonshot AI at $3/$15 per 1M tokens, a 1M-token context, open weights under a custom license and an Artificial Analysis index of 57, against Claude Opus 4.8 from Anthropic at $5/$25 per 1M tokens, a 1M-token context, five effort levels and an index of 56

Last updated on

Kimi K3 vs Claude: the price gap is smaller than it looks


Moonshot’s Kimi K3 landed on July 16, 2026 benchmarking within a point of Claude Opus 4.8 while charging 40% less per token, and the internet did what it always does with that setup: declared the expensive model obsolete. The rate card is real. The conclusion isn’t, and the reason has less to do with model quality than with which Claude you’re actually comparing against and how many tokens each model burns getting to an answer.

TL;DR

  • On capability, it’s a coin flip. Artificial Analysis scores K3 at 57 and Opus 4.8 at 56 on its Intelligence Index. Moonshot itself says K3 trails Claude Fable 5.
  • K3 is 40% cheaper than Opus 4.8 per token ($3/$15 vs $5/$25) — and priced identically to Claude Sonnet 5’s $3/$15, which is currently discounted to $2/$10.
  • Every K3 benchmark Moonshot published was run at max reasoning effort, which is also K3’s launch default. Opus 4.8 defaults to high and gives you four cheaper rungs.
  • Claude has verified coding scores and a track record. K3 has launch-week leaderboard placements and no independent SWE-bench results yet.
  • Open weights are the one axis where K3 wins outright, and they hadn’t shipped as of this writing.
  • We haven’t bench-tested the two head to head. This is spec, pricing, and published-benchmark analysis; the hands-on run is still to come.

Which Claude are we comparing?

“Kimi K3 vs Claude” is an underspecified question, because Claude is four models at four price points. Picking the wrong one is how most comparisons on this query go wrong.

Anthropic’s lineup runs Haiku 4.5 ($1/$5 per million tokens in/out), Sonnet 5 ($3/$15), Opus 4.8 ($5/$25), and Fable 5 ($10/$50) — we break the tiers down in our Claude models comparison. K3 benchmarks near Opus 4.8, so that’s the capability matchup. But it’s priced exactly at Sonnet 5. Both comparisons are legitimate, and they point in opposite directions.

Fable 5 is out of scope by Moonshot’s own admission: its launch post says K3 trails Claude Fable 5 and GPT-5.6 Sol on overall performance. When a vendor concedes a loss in its own announcement, take it at face value.

How do the specs compare?

Opus 4.8 and K3 are closer on paper than the price difference suggests. The gaps that matter are in output ceiling and effort control.

Kimi K3Claude Opus 4.8
ReleasedJuly 16, 2026May 28, 2026
Architecture2.8T params, MoE (16 of 896 experts active)Not disclosed
Context window1M tokens1M tokens
Max output131,072 default, up to 1,048,576128,000
Input / output per 1M$3.00 / $15.00$5.00 / $25.00
Cache-hit input$0.30$0.50
Reasoning effortmax at launch; low / high announced for laterlow / medium / high / xhigh / max
WeightsShipped July 26, 2026 (custom Kimi K3 License)Closed, API only

The output ceiling is K3’s one clean spec win, and it’s a real one for long-form generation: 1,048,576 tokens configurable against Anthropic’s fixed 128K (Kimi platform docs). Whether you want a single response that long is another question.

Is Kimi K3 actually cheaper?

Per token, yes, by 40% against Opus 4.8. Per completed task the gap narrows, and three things distort the straight rate-card read.

The rate card isn’t priced against Opus. K3’s $3/$15 is, to the cent, Claude Sonnet 5’s standard pricing. Sonnet 5 is also running introductory rates of $2 input and $10 output through August 31, 2026, which makes K3 roughly 50% more expensive than the Claude tier it’s price-matched to until that promo ends. Moonshot priced K3 at the mid tier and benchmarked it against the top one. That’s good positioning, not a discount.

Prompt caching doesn’t move the needle either way. Both vendors charge 0.1× base input for a cache hit — $0.30 for K3, $0.50 for Opus 4.8 — so the ratio between them is unchanged. Moonshot reports its stack clears a 90% cache-hit rate on coding workloads, which is good news for K3’s effective input cost and equally good news for anyone using Claude’s caching well.

Effort is where the real cost lives. Buried in Moonshot’s launch post is a line that reframes every K3 benchmark you’ve seen: “All Kimi K3 results reported below are obtained with the reasoning effort set to max…” The same post says K3 ships with max effort as the launch default, with low- and high-effort modes coming in later updates. Kimi’s platform docs now list low, high, and max, so the lower rungs appear to be arriving, but the published scores are all max-effort scores.

Comparison card: Kimi K3 at $3 / $15 per million tokens against Claude Opus 4.8 at $5 / $25, with Artificial Analysis scoring them 57 versus 56 on its Intelligence Index with both models at max effort, and a note that every published Kimi K3 benchmark runs at max reasoning effort — its launch default

Opus 4.8 defaults to high and exposes five levels. That difference is not cosmetic. As we found testing the effort dial in our Opus 4.8 review, the setting is the single biggest lever on the model’s token spend, and cranking it to max reflexively is the most common way to waste money on Claude. A model whose only published configuration is the most expensive one is a model you cannot yet tune down. The early reports match: in launch-week Hacker News threads, one developer found a single agentic task nearly exhausted a five-hour window on Kimi’s $19 plan.

A token isn’t a token across vendors. Comparing rate cards assumes both companies count tokens the same way. They don’t. Anthropic notes that Opus 4.7 and later, Sonnet 5, and Fable 5 run a newer tokenizer that produces roughly 30% more tokens for the same text (pricing docs). Read that carefully before drawing a conclusion: it’s a comparison against Anthropic’s own earlier models, not against Moonshot’s. There’s no published figure for how K3 tokenizes the same input, so this doesn’t tell you which way the cross-vendor ratio moves. What it does tell you is that a 40% gap between two rate cards is a softer number than it looks.

Which is why the only figure that settles this is your own. Artificial Analysis publishes cost-per-task and output-tokens-per-task for both models, and K3’s verbosity visibly eats into its rate-card lead there. If you’re evaluating K3 on a metered API, set spending caps before you start and measure cost per finished task, not per million tokens. Our guide to cutting token costs applies to both models equally.

Which one is better at coding?

Nobody outside the two vendors can answer this properly yet, and any article that gives you a confident winner is guessing.

The problem is that the two companies publish different benchmarks. Anthropic’s Opus 4.8 system card reports 88.6% on SWE-bench Verified, 69.2% on the harder SWE-bench Pro, and 74.6% on Terminal-Bench 2.1. K3 scores 67.3 on the official DeepSWE leaderboard under the standard mini-SWE-agent harness, and 90.4 on BrowseComp with the full 1M context and no context management. Those measure different things. There is no shared, published, apples-to-apples coding number.

One detail in K3’s favor: that 67.3 comes from the neutral leaderboard harness, and Moonshot’s own KimiCode harness reports 67.5 on the same benchmark. When a vendor’s in-house harness lands within two-tenths of the standardized one, the in-house numbers are more trustworthy than they’d otherwise be.

The closest thing to a neutral referee is Artificial Analysis, which runs both through the same Intelligence Index: K3 scores 57, Opus 4.8 scores 56. Worth knowing what that second number is, though. AA scores the model as “Claude Opus 4.8 (Adaptive Reasoning, Max Effort)”, so this is an effort-matched comparison, not a default-vs-default one. A single point is noise either way. Treat it as parity, not a K3 win.

What still tilts the practical read toward Claude is verification, not capability. Independent SWE-bench results for K3 haven’t been published, so its SWE-side claims rest on Moonshot’s reporting. Terminal-Bench is the partial exception: AA runs v2.1 on both models as one of the nine evaluations inside that Intelligence Index, which makes it the first genuinely independent coding datapoint K3 has. And launch-week leaderboard positions churn. K3 took the #1 spot on Arena’s frontend coding board at launch, which is genuinely impressive and also the kind of placement that moves as votes accumulate. If K3 still holds it in September, that means considerably more than the July snapshot.

Where K3 looks strongest on the evidence available is long-context and browsing-style agentic work, which is what that 90.4 BrowseComp score with no context management is actually measuring.

What about open weights?

This is K3’s one unambiguous advantage over Claude, and it stopped being a promise on July 26, 2026, when Moonshot published the weights on Hugging Face a day ahead of its own deadline.

Claude is closed and API-only, with no self-hosting path at any price. That’s the differentiator, and it isn’t going to close: if your requirement is “the weights run on infrastructure we control,” Claude cannot meet it and K3 now does.

Read the license before you count on it, though. K2 was Modified MIT; K3 is a bespoke Kimi K3 License that requires a separate agreement to offer the model as a service once aggregate revenue passes $20 million over any 12 months, and mandates displaying “Kimi K3” in your interface above 100 million monthly active users or $20 million in monthly revenue. For internal engineering that’s a non-issue. For a product you sell, it’s a clause your legal team will want to see.

The practical caveat hasn’t changed: a 2.8T-parameter model is multi-GPU server territory even quantized, so this buys you cheaper third-party hosting and a compliance story, not a model on your laptop. The clearest proof is GitHub — since August 6 it serves K3 inside Copilot, hosted on Fireworks AI rather than proxied to Moonshot, which is the data-path argument for open weights made concrete. Our Kimi K3 explainer covers the license in more detail, and our security and privacy guide covers how we assess where code and prompts travel.

Kimi Code vs Claude Code: the subscription math

If you’re buying a coding agent rather than metering an API, the comparison shifts to products, and the entry prices are nearly identical.

Kimi Code has no standalone price — it’s bundled into Kimi’s membership tiers, starting at $19/mo:

Claude Code is bundled into Claude subscriptions from $20/mo:

A dollar apart, and the difference is in what you can predict. Anthropic publishes usage as multiples of the Pro tier; Moonshot publishes credit multipliers (1× through 30×) that don’t map to any disclosed token count. On a max-effort default, opaque credits are a harder thing to budget against. Full plan detail for both sits in our pricing tracker.

So which should you use?

Use Claude Opus 4.8 when the work is expensive to get wrong or runs unattended. The effort dial is worth real money on long-running agents, and the coding scores have been checked by people other than the vendor.

Try Kimi K3 for long-context and browsing-heavy work, where its strongest published evidence actually sits. The open weights are the other reason, if hosting location is a hard requirement for you — and since August they are a real option rather than a roadmap item.

And if what grabbed you was the price, go compare Sonnet 5 instead. Same rate card as K3, currently cheaper than both, and it’s the tier most developers should be defaulting to anyway.

We flagged this at the top and it bears repeating at the bottom: this is spec, pricing, and published-benchmark analysis, all verified against primary sources, not a scored head-to-head on our own bench. We run Opus 4.8 in daily work. We haven’t put K3 through the same, and until we do, the cost-per-task question stays genuinely open. That run is next. The rest of our comparisons cover tools we have tested.

  • kimi-k3
  • moonshot-ai
  • anthropic
  • claude-opus-4-8
  • ai-models-apis

Frequently asked questions

Is Kimi K3 better than Claude?

On the one third-party benchmark that scores both the same way, Artificial Analysis's Intelligence Index, Kimi K3 sits at 57 and Claude Opus 4.8 at 56 — close enough to call a tie, and note that AA scores Opus 4.8 at max effort, so it's effort-matched. Moonshot's own launch post says K3 trails Claude Fable 5 on overall performance. The gap is in verification rather than capability: Opus 4.8 has published SWE-bench numbers (88.6% Verified, 69.2% Pro) while K3's independent SWE-bench results haven't landed yet.

Is Kimi K3 cheaper than Claude?

Cheaper than Claude Opus 4.8, yes: $3 per million input tokens and $15 per million output, against $5 and $25. But K3's rate card is identical to Claude Sonnet 5's standard $3/$15, and Sonnet 5 is running introductory pricing of $2/$10 through August 31, 2026 — so against that tier, K3 is currently the more expensive option.

Which Claude model should I compare Kimi K3 to?

Opus 4.8 for capability, Sonnet 5 for price. K3 benchmarks near Opus 4.8 while charging Sonnet 5 rates, which is exactly the pitch. Claude Fable 5 ($10/$50) sits above both and is the model Moonshot concedes K3 trails.

Can I self-host Kimi K3 instead of using Claude's API?

You can now, though probably not on your own hardware. Moonshot published the weights on July 26, 2026 under a custom Kimi K3 License — open weights, but not OSI open source: it requires a separate agreement to offer the model as a service above $20 million in aggregate 12-month revenue. A 2.8 trillion-parameter model needs multi-GPU server hardware regardless of quantization, so the practical beneficiaries are hosting providers and teams with data-residency requirements, not individual developers. Claude offers no self-hosting path at any price, so this remains K3's one structural advantage.