Kimi K3 Token Hub

Kimi K3 pricing

Last verified: · source

Kimi K3 costs $3.00 per million input tokens and $15.00 per million output tokens on the official Moonshot API. Cached input is discounted (official cached rate not yet published). That puts K3 at roughly 3× the price of Kimi K2.6, level with Claude Sonnet's tier, and well above DeepSeek V4 Pro. Below: what real requests cost, how caching changes the math, and how every channel compares.

Official API rates

MeterPrice / 1M tokensNotes
Input (cache miss)$3.00Standard prompt tokens
Input (cache hit)discounted — official rate TBAReported 60–80% effective saving on cache-heavy loads
Output$15.00Completion + reasoning tokens

What does one request cost?

Take a typical agent step: 10,000 input tokens, 1,000 output tokens. Without caching that is $0.045 — $0.030 of input plus $0.015 of output. With a 70% cache hit rate — normal for coding agents that resend the same repo context — that request drops to roughly $0.029 using our estimated cache discount. Notice what that implies:

  • Output dominates small-prompt workloads. At $15/M, one output token costs as much as five input tokens.
  • Cache hit rate moves the bill more than anything else you control — more than switching providers within the same tier.

Run your own workload through the cost calculator — it models cache rates, retries and monthly volume, and compares the same load across five models.

Kimi K3 vs other models

Kimi K3 API pricing comparison 2026: input and output cost per million tokens vs Claude Sonnet, GPT-5.5, DeepSeek V4 Pro and Kimi K2.6
ModelInput / 1M tokensOutput / 1M tokensContextLicense
Kimi K3$3.00$15.001,048.576Kopen-weight
Kimi K2.6$0.95$4.00262.144Kopen-weight
DeepSeek V4 Pro$1.10*$4.40*262.144Kopen-weight
GPT-5.5$2.50*$20.00*400Kproprietary
Claude Sonnet$3.00*$15.00*1,000Kproprietary

* listed price at time of verification; check the linked source for the vendor's current rate.

Context for that table: K2.6 remains the budget pick of the Kimi family at $0.95/$4 — if your workload doesn't need K3's reasoning depth or 1M context, the older model is close to 3.5× cheaper per token. Moonshot's own data says K3 finishes tasks with about 21% fewer output tokens than K2.6, so the effective multiplier on real bills is smaller, but K2.6 still wins on pure economics. Against closed models, K3 matches Claude Sonnet's $3/$15 sticker while undercutting flagship tiers, and its open-weights release (due July 27, 2026) should bring cheaper third-party hosting soon after — we'll track those prices here.

Where to buy Kimi K3 tokens

  • Moonshot / Kimi platform — first-party, full feature set including context caching. See the API guide.
  • OpenRouter — same $3/$15 pass-through, unified billing across models.
  • Third-party GPU clouds — expected after the weights release; prices TBD.

How we verify these numbers

Every price on this site lives in a single registry file with a source URL and a verification date, rendered as the "Last verified" stamp above. When Moonshot changes a rate, we update the registry and every page, table and calculator updates together. If you spot a discrepancy, tell us — the correction takes minutes.

Frequently asked questions

How much does the Kimi K3 API cost?

The Kimi K3 API costs $3.00 per million input tokens and $15.00 per million output tokens (verified July 21, 2026 against the OpenRouter listing, which passes through Moonshot's official rate).

Does Kimi K3 have a cached-input discount?

Prompt caching is supported and cached input is billed at a discount, but Moonshot has not published the exact K3 cached rate yet. Reports suggest effective input costs drop 60–80% on cache-heavy workloads. We flag all cache math on this site as an estimate until the official rate lands.

Is Kimi K3 more expensive than Kimi K2.6?

Yes — about 3× on input ($3 vs $0.95) and 3.75× on output ($15 vs $4). Moonshot reports K3 uses roughly 21% fewer output tokens for the same tasks, which narrows the real-world gap somewhat.

How much does a typical Kimi K3 request cost?

A request with 10,000 input tokens and 1,000 output tokens costs about $0.045 without caching: $0.03 for input plus $0.015 for output. A million such requests would be roughly $45,000 per month — which is why cache hit rate matters so much.

Sources