NVIDIA announced NVIDIA Nemotron 3 Nano Omni, an omni-modal understanding model for document analysis, speech recognition, audio-video understanding, agentic computer use, and general reasoning. It combines a hybrid Mamba-Transformer-MoE backbone with C-RADIOv4-H vision and Parakeet-TDT-0.6B-v2 audio encoders. The model leads benchmarks including MMlongbench-Doc, OCRBenchV2, WorldSense, and VoiceBench while delivering up to 9x higher throughput than alternatives. Checkpoints are available on HuggingFace.
No score is assigned. Sources and their independence are shown in the citation chain below.