nemotron3

NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows.

Görüntü Araç Kullanımı Akıl Yürütme 33b
Hızlı Kurulum (Ollama kuruluysa)
ollama run nemotron3

Ollama kurulu değil mi? ollama.com/download — Windows, macOS ve Linux için ücretsiz. İlk çalıştırmada model indirilir, sonrası tamamen çevrimdışıdır.

Varyantlar

Boyut büyüdükçe kalite artar, donanım ihtiyacı yükselir. Başlangıç için küçük varyantı deneyin.

EtiketBoyutBağlamGirdiKomut
33b 28GB 128K Text, Image ollama run nemotron3:33b
33b-q8 36GB 128K Text, Image ollama run nemotron3:33b-q8
33b-q4_K_M 28GB 128K Text, Image ollama run nemotron3:33b-q4_K_M
33b-bf16 66GB 128K Text, Image ollama run nemotron3:33b-bf16

Model Detayları ve Benchmarklar (kaynak: ollama.com)

NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (OCR), and speech transcription capabilities, enabling end-to-end processing of rich enterprise content such as meeting recordings, M&E assets, training videos, and complex business documents. NVIDIA Nemotron 3 Nano Omni was developed by NVIDIA as part of the Nemotron model family.

This model is available for commercial use.

This model was improved using Qwen3-VL-30B-A3B-Instruct, Qwen3.5-122B-A10B, Qwen3.5-397B-A17B, Qwen2.5-VL-72B-Instruct, and gpt-oss-120b.

License/Terms of Use

Governing Terms: Use of this model is governed by the NVIDIA Open Model Agreement

Use Case

This model is designed for enterprise customers requiring multimodal understanding capabilities. Expected users include: - Customer service applications (e.g., Doordash video of drop-off at a given address via OCR, drive-thru order verification) - Media and Entertainment (M&E) — video and speech analysis, dense captions, video search and summarization - Document intelligence for AI assistants (contracts, SOW/MSA, scientific discovery, financial documents) - GUI automation for AI agentic applications (incident management, agentic search, browser agents, email agents)