DeepSeek released two Mixture-of-Experts checkpoints, DeepSeek-V4-Pro and DeepSeek-V4-Flash, both featuring a 1M-token context window. The models use hybrid attention mechanisms—Compressed Sparse Attention and Heavily Compressed Attention—to reduce KV cache memory to roughly 2% of standard architectures. DeepSeek-V4-Pro requires 27% of single-token inference FLOPs compared to its predecessor. Writing for Hugging Face, ben burtenshaw notes the design targets long-running agentic workloads, preserving reasoning traces across tool calls and user message boundaries.
No score is assigned. Sources and their independence are shown in the citation chain below.