JetBrains released Mellum2, a 12B Mixture-of-Experts model optimized for low-latency text and code workloads. The model activates only a subset of parameters per token, achieving competitive performance with similarly sized open models while delivering more than 2x faster inference. Use cases include routing and orchestration, RAG pipelines, sub-agent tasks, and private deployment. Nikita Pavlichenko announced the release, describing Mellum2 as a well-scoped model for high-frequency tasks within larger AI systems.
No score is assigned. Sources and their independence are shown in the citation chain below.