← Back

Moonshot AI

companyCredibility: 57%

Why this score? Chinese AI lab (Kimi models); primary for its own releases, motivated on capability claims.

Tracked Statements (2)

It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.

Context: The checkable core holds. The published weights are for a 2.8-trillion-parameter model, larger than any previously released open-weight model — corroborated independently by Tom's Hardware and, before the release, by Rahul Shome of the Australian National University in Nature, who said K3 "would be the largest open-weight model". Two caveats that do not falsify it: "3T-class" is a generous rounding of 2.8 trillion, and only 104B of those parameters activate per token. The trailing capability claim is separately corroborated — Artificial Analysis scores K3 at 57 on its Intelligence Index v4.1, four points off the highest score on that index — but the benchmark table Moonshot published with the weights is not: on it, K3 is run with Moonshot's own Kimi Code harness while rivals use Claude Code or Codex, and Moonshot discloses that Claude Fable 5 hit fallbacks on 35% of its SWE-Marathon tasks.

Moonshot AI claims Kimi K3 beat Claude Opus 4.8 and GPT 5.5 — models CNBC describes as sitting just behind Anthropic's and OpenAI's leading-edge systems — on benchmarks including coding and general agents.±

Context: Both halves are now measurable, and they part company. On Artificial Analysis's Intelligence Index v4.1 — the named independent index this record uses — Kimi K3 scores 57 against GPT-5.5 (xhigh) at 55, so that half stands. The Claude Opus 4.8 half, unscorable in this record when the claim was first checked, now resolves to 56 at max effort: K3 leads by one point, and one point is a margin Artificial Analysis itself declines to treat as a lead — it called Claude Opus 5's identical one-point margin over Claude Fable 5 on 24 July an effective tie. By AA's own yardstick K3 ties Opus 4.8 rather than beating it. The framing caveat holds and has widened: neither model named was its lab's flagship, and since the claim was made Claude Opus 5 has taken the top of the index at 61, four points above K3. The domain-level assertion — that K3 won specifically on coding and general-agent benchmarks — still rests on Moonshot's own benchmark table, which runs K3 under Moonshot's Kimi Code harness while rivals use theirs, and is not independently replicated in this record.