Kimi K3 · Explorer

Point at the architecture

The same explanations as the course, reached by pointing at a mechanism instead of scrolling to it. Ten parts, grouped the way the report groups them.

Each box below opens that mechanism's explanation in place. The layout is the argument: the three axes are how the report itself organises K3, and the two boxes tucked under Stable LatentMoE are there because they fix problems its extreme sparsity creates rather than standing on their own.

Map of the Kimi K3 architecture Ten clickable parts of K3, grouped by the report's own three axes of information flow. Sequence carries KDA and the chunkwise form with g_min and NoPE. Depth carries Attention Residuals. Width carries Stable LatentMoE, with SiTU-GLU and Quantile Balancing nested beneath it because both are fixes for problems extreme sparsity creates. Outside the backbone sit MoonViT-V2 for input, Per-Head Muon for optimisation, serving, and post-training. THE BACKBONE · three axes of information flow SEQUENCE how tokens mix KDA chunkwise · g_min · NoPE DEPTH how layers mix Attention Residuals WIDTH how channels mix Stable LatentMoE SiTU-GLU bounds activations for 4-bit Quantile Balancing balances 896 experts in one step OUTSIDE THE BACKBONE · everything else K3 needs MoonViT-V2 input Per-Head Muon optimisation Serving K3 systems Post-training behaviour

Nothing here is new material. If you are reading K3 for the first time, start at lesson 1 and let the sequence do its work; this page is built for coming back.

Choose a mechanism above.