1、KIMIK3:OPENFRONTIERINTELLIGENCETECHNICALREPORT OFKIMIK3Kimi TeamABSTRACTWe introduce Kimi K3,a 2.8T parameter Mixture-of-Experts model with 104 billion activatedparameters,native vision capabilities,and a 1-million-token context window.Kimi K3 is built onKimi Delta Attention 64 and Attention Residua
2、ls 58,which improve information flow acrosssequence length and model depth.Together with Stable LatentMoE,which effectively activates16 of 896 routed experts per token,and refi ned training and data recipes,these advances yield anapproximately2.5improvement in overall scaling effi ciency over Kimi K
3、2 59.Post-traininghighlights reinforcement learning across general,agentic,and coding domains and multiple reasoning-effort levels,enabling compositional generalization and robust long-horizon execution.At 2.8T scale,Kimi K3 is supported by infrastructure advances in multiple areas:algorithmsystem c
4、o-design forKDA,perfectly balanced expert-parallel training with effi cient memory management,million-tokenagentic RL with persistent rollout and sandbox states,and deployment innovations.Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizoncoding,agentic,kn
5、owledge,reasoning,and vision tasks.While its overall performance still trails themost powerful proprietary models,namely Claude Fable 5 and GPT-5.6 Sol,Kimi K3 consistentlyoutperforms other open and proprietary models evaluated in our suite.We release the full Kimi K3model weights to facilitate futu
6、re research and accelerate the broader deployment and adoption offrontier intelligence.1CodingAll maxed out on thinking effort:max or xhigh.DeepSWEGPT-5.6 SolFable 5Kimi K3GPT-5.5Opus 4.8GLM-5.2Terminal-Bench 2.1GPT-5.6 SolKimi K3Fable 5Opus 4.8GPT-5.5GLM-5.273.070.067.567.059.046.2Kimi Code Bench 2