← Back to the wire

Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents

AnnouncementModelApr 28, 2026

NVIDIA announced NVIDIA Nemotron 3 Nano Omni, an omni-modal understanding model for document analysis, speech recognition, audio-video understanding, agentic computer use, and general reasoning. It combines a hybrid Mamba-Transformer-MoE backbone with C-RADIOv4-H vision and Parakeet-TDT-0.6B-v2 audio encoders. The model leads benchmarks including MMlongbench-Doc, OCRBenchV2, WorldSense, and VoiceBench while delivering up to 9x higher throughput than alternatives. Checkpoints are available on HuggingFace.

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

NVIDIACompanyNVIDIA Nemotron 3 Nano OmniModelTuomas RintamakiPerson
Canonical: https://huggingface.co/blog/nvidia/nemotron-3-nano-omni-multimodal-intelligence