Kimi K3 pricing
Last verified: · source
Kimi K3 costs $3.00 per million input tokens and $15.00 per million output tokens on the official Moonshot API. Cached input is discounted (official cached rate not yet published). That puts K3 at roughly 3× the price of Kimi K2.6, level with Claude Sonnet's tier, and well above DeepSeek V4 Pro. Below: what real requests cost, how caching changes the math, and how every channel compares.
Official API rates
| Meter | Price / 1M tokens | Notes |
|---|---|---|
| Input (cache miss) | $3.00 | Standard prompt tokens |
| Input (cache hit) | discounted — official rate TBA | Reported 60–80% effective saving on cache-heavy loads |
| Output | $15.00 | Completion + reasoning tokens |
What does one request cost?
Take a typical agent step: 10,000 input tokens, 1,000 output tokens. Without caching that is $0.045 — $0.030 of input plus $0.015 of output. With a 70% cache hit rate — normal for coding agents that resend the same repo context — that request drops to roughly $0.029 using our estimated cache discount. Notice what that implies:
- Output dominates small-prompt workloads. At $15/M, one output token costs as much as five input tokens.
- Cache hit rate moves the bill more than anything else you control — more than switching providers within the same tier.
Run your own workload through the cost calculator — it models cache rates, retries and monthly volume, and compares the same load across five models.
Kimi K3 vs other models

| Model | Input / 1M tokens | Output / 1M tokens | Context | License |
|---|---|---|---|---|
| Kimi K3 | $3.00 | $15.00 | 1,048.576K | open-weight |
| Kimi K2.6 | $0.95 | $4.00 | 262.144K | open-weight |
| DeepSeek V4 Pro | $1.10* | $4.40* | 262.144K | open-weight |
| GPT-5.5 | $2.50* | $20.00* | 400K | proprietary |
| Claude Sonnet | $3.00* | $15.00* | 1,000K | proprietary |
* listed price at time of verification; check the linked source for the vendor's current rate.
Context for that table: K2.6 remains the budget pick of the Kimi family at $0.95/$4 — if your workload doesn't need K3's reasoning depth or 1M context, the older model is close to 3.5× cheaper per token. Moonshot's own data says K3 finishes tasks with about 21% fewer output tokens than K2.6, so the effective multiplier on real bills is smaller, but K2.6 still wins on pure economics. Against closed models, K3 matches Claude Sonnet's $3/$15 sticker while undercutting flagship tiers, and its open-weights release (due July 27, 2026) should bring cheaper third-party hosting soon after — we'll track those prices here.
Where to buy Kimi K3 tokens
- Moonshot / Kimi platform — first-party, full feature set including context caching. See the API guide.
- OpenRouter — same $3/$15 pass-through, unified billing across models.
- Third-party GPU clouds — expected after the weights release; prices TBD.
How we verify these numbers
Every price on this site lives in a single registry file with a source URL and a verification date, rendered as the "Last verified" stamp above. When Moonshot changes a rate, we update the registry and every page, table and calculator updates together. If you spot a discrepancy, tell us — the correction takes minutes.
Frequently asked questions
How much does the Kimi K3 API cost?
The Kimi K3 API costs $3.00 per million input tokens and $15.00 per million output tokens (verified July 21, 2026 against the OpenRouter listing, which passes through Moonshot's official rate).
Does Kimi K3 have a cached-input discount?
Prompt caching is supported and cached input is billed at a discount, but Moonshot has not published the exact K3 cached rate yet. Reports suggest effective input costs drop 60–80% on cache-heavy workloads. We flag all cache math on this site as an estimate until the official rate lands.
Is Kimi K3 more expensive than Kimi K2.6?
Yes — about 3× on input ($3 vs $0.95) and 3.75× on output ($15 vs $4). Moonshot reports K3 uses roughly 21% fewer output tokens for the same tasks, which narrows the real-world gap somewhat.
How much does a typical Kimi K3 request cost?
A request with 10,000 input tokens and 1,000 output tokens costs about $0.045 without caching: $0.03 for input plus $0.015 for output. A million such requests would be roughly $45,000 per month — which is why cache hit rate matters so much.