← Back to the wire

DeepSeek-V4: a million-token context that agents can actually use

AnnouncementModelApr 24, 2026

DeepSeek released two Mixture-of-Experts checkpoints, DeepSeek-V4-Pro and DeepSeek-V4-Flash, both featuring a 1M-token context window. The models use hybrid attention mechanisms—Compressed Sparse Attention and Heavily Compressed Attention—to reduce KV cache memory to roughly 2% of standard architectures. DeepSeek-V4-Pro requires 27% of single-token inference FLOPs compared to its predecessor. Writing for Hugging Face, ben burtenshaw notes the design targets long-running agentic workloads, preserving reasoning traces across tool calls and user message boundaries.

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01highPRIMARY
Hugging FaceCompanyDeepSeekCompanyben burtenshawPersonDeepSeek-V4-ProModelDeepSeek-V4-FlashModel
Canonical: https://huggingface.co/blog/deepseekv4