Kimi K3 Token Hub

Kimi K3 benchmarks — a living tracker

Last updated:

Five days after launch, the headline numbers are: Artificial Analysis Elo 1547 (private evals, +732 over Kimi K2.6), self-reported wins against Claude Opus 4.8 (max) and GPT-5.5 (high) on most of Moonshot's published table, first place on Arena.ai's Frontend Code arena, and ~21% fewer output tokens than K2.6 on matched tasks. Independent reproduction is still thin — this page tracks numbers as they firm up.

Scoreboard (as of July 21, 2026)

BenchmarkKimi K3 resultStatus
Artificial Analysis (private evals)Elo 1547 (+732 vs K2.6)independent
Moonshot launch table vs Opus 4.8 maxmostly aheadself-reported
Moonshot launch table vs GPT-5.5 highmostly aheadself-reported
vs Claude Fable 5 / GPT-5.6 Soltrailingself-reported
Arena.ai Frontend Code arena#1 at launch weekcommunity voting
Output-token efficiency vs K2.6−21% tokens on matched tasksself-reported

How to read these numbers

Keep three caveats taped to your monitor. Self-reported tables are marketing artifacts — every lab picks favorable benchmarks; the pattern to trust is direction, not decimals. Elo systems measure preference, not correctness — a model can win Elo by writing pleasant answers while failing hard tasks. And open-weight verification hasn't happened yet: until weights land (due July 27, 2026) nobody outside Moonshot can rerun the evals. The strongest launch-week signal is arguably the output-token efficiency claim, because it is directly measurable from API usage logs — and it compounds with the $15/M output price into real cost differences, which you can model in the cost calculator.

What we're waiting for

  • Weights release → independent benchmark reruns (SWE-bench-style agentic suites first, typically).
  • Artificial Analysis public-suite scores to complement the private-eval Elo.
  • Long-context stress tests against the 1M window — see what fits in 1M tokens.
  • Third-party hosting latency/throughput data once GPU clouds pick the weights up.

Each of those gets added here with a dated entry when it happens; the update stamp at the top is the audit trail. For spec-level context see what is Kimi K3; for what the scores cost to use, see pricing.

Frequently asked questions

Is Kimi K3 better than GPT-5.5?

On Moonshot's self-reported benchmark table K3 mostly outperforms GPT-5.5 at 'high' effort, and Artificial Analysis scored K3 at Elo 1547 on private evals. But it still trails the newest closed flagships, and self-reported numbers routinely shrink under independent reproduction — treat 'better' as workload-specific.

What is Kimi K3's Artificial Analysis score?

Artificial Analysis reported an Elo of 1547 on its private evaluations at launch — roughly +732 points over Kimi K2.6, one of the largest generation-over-generation jumps recorded for an open-weight family.

Where does Kimi K3 rank for coding?

K3 led Arena.ai's Frontend Code arena at launch week and is reported to use ~21% fewer output tokens than K2.6 on identical tasks, which matters for both quality and cost in agentic coding loops.

Sources