Moonshot's Kimi K3 scores 1,679 on the Code Arena: Frontend benchmark, surpassing Claude Fable 5 and GPT-5.6 Sol to become the first Chinese model to claim the top spot. However, on FrontierMath Tier 4, Kimi K3 reaches only about 39 percent accuracy, while models from OpenAI and Anthropic approach 90 percent. The results show Kimi K3 excelling in frontend code generation but trailing significantly in complex mathematical reasoning.
No score is assigned. Sources and their independence are shown in the citation chain below.