Kimi K3 · Reference sheet
Every number worth remembering, and where in the report to find it. Built to print.
┌──────────────────────────────────────────────────────────┐
│ 1 × Gated MLA + Stable LatentMoE global, NoPE │
│ 3 × KDA + Stable LatentMoE linear, O(n) │ ◀── × 23 blocks
└──────────────────────────────────────────────────────────┘
+ 1 trailing Gated MLA
93 layers = 69 KDA + 24 Gated MLA 3:1 hybrid ratio
AttnRes: 8 blocks × 12 layers (final block partial), N=8
Input: MoonViT-V2 (401M, 27 layers) ─▶ MLP projector ─▶ backbone
| Total parameters | 2.78 T |
| Activated / token | 104.2 B |
| Layers | 93 |
| Hidden dimension | 7,168 |
| Attention heads | 96 |
| Vocabulary | 160 K |
| Context window | 1 M |
| Dense layers | 1 |
| MTP layers | 1 |
| Routed experts / layer | 896 |
| Active per token | 16 |
| Sparsity ratio | 56 |
| Shared experts | 2 |
| Latent MoE width ℓ | 3,584 |
| as fraction of d | 0.5 × |
| MoE hidden / expert | 3,072 |
| Parameters | 401 M |
| Layers | 27 |
| Attention heads | 12 |
| Patch size | 14 |
| Pixel-shuffle downsample | 2 × 2 |
| Max input resolution | 3584² |
| Name | Job | § |
|---|---|---|
| KDA | Linear-cost token mixing | 2.1.1 |
| Gated MLA | Global attention, gated output | 2.1.2 |
| AttnRes | Retrieve across depth | 2.2 |
| LatentMoE | Experts at half width | 2.3 |
| SiTU-GLU | Bounded activations | 2.3.2 |
| Quantile Bal. | Balance 896 experts | 2.3.3 |
| MoonViT-V2 | Vision, from scratch | 2.4 |
| Per-Head Muon | Per-head update scale | 2.5 |
| KDA log-decay floor gmin | −5 |
| → min retention α | 6.7e−3 |
| SiTU gate cap β₁ | 4 |
| SiTU up cap β₂ | 25 |
| → output bound | 100 |
| AttnRes blocks N | 8 |
| Weight decay | 0.1 |
| LR warmup | 1% |
| LR schedule | cosine |
| Pre-training start | 8 K |
| Pre-training extended | 64 K |
| Cooldown start | 256 K |
| Cooldown final | 1 M |
| Input, cache hit | $0.30 |
| Input, cache miss | $3.00 |
| Output | $15.00 |
| Recommended accelerators | 64+ |
| Quantization (QAT) | MXFP4 / MXFP8 |
| Coding cache-hit rate | >90% |
| Default effort | max |
| Benchmark | K3 | best rival |
|---|---|---|
| BrowseComp | 91.2 | 90.4 |
| Terminal-Bench 2.1 | 88.3 | 88.8 |
| ProgramBench | 77.8 | 77.6 |
| SWE-Marathon | 42.0 | 40.0 |
| FrontierSWE | 81.2 | 86.6 |
| DeepSWE | 67.5 | 73.0 |
| WebDev Arena (Elo) | 1,678 | 1,634 |
| AA Intelligence v4.1 | 57.1 | 59.9 |
All Claude Fable 5 results include potential fallbacks; all GPT-5.6 Sol results include potential cyberguards. K3 leads on BrowseComp, ProgramBench, SWE-Marathon and WebDev Arena, and trails Fable 5 and GPT-5.6 Sol overall. The 2.5× scaling-efficiency claim is the authors' own fit on held-out validation data, not independently replicated.
| Limitation | Practical consequence |
|---|---|
| Sensitivity to thinking history | Trained with preserved thinking history. Quality can become highly unstable if you don't pass reasoning back, or switch sessions mid-stream. |
| Excessive proactiveness | May make unexpected decisions on your behalf. Constrain via system prompt or AGENTS.md. |
| User-experience gap | Noticeable gap versus Claude Fable 5 and GPT-5.6 Sol, beyond what benchmarks capture. |