Together AI Brings NVIDIA Nemotron 3 Nano Omni to Developers on Day 0
Quick Answer
Together AI has launched the NVIDIA Nemotron 3 Nano Omni, a multimodal AI model that combines video, audio, images, and language reasoning.
Quick Take
This model, utilizing a hybrid Mamba-Transformer architecture, allows developers to build agentic applications with high efficiency and low latency, streamlining deployment from prototype to production without infrastructure management.
Key Points
- Nemotron 3 Nano Omni supports 256K tokens of shared multimodal input context.
- The model's architecture activates ~3B parameters per token from a total of 30B.
- Together AI ensures reliable performance for agent applications during traffic spikes.
- Developers can deploy the model without managing infrastructure, focusing on building.
- Open weights and data control allow deployment across various environments without lock-in.
Source Excerpt
NVIDIA Nemotron 3 Nano Omni is now on Together AI: a single open model that reasons across video, images, audio, and text, built for agentic workloads at scale.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Together AI
See more →
Open, convenient and predictable: Introducing Provisioned Throughput
Together AI introduces Provisioned Throughput, offering guaranteed inference capacity for MiniMax M3 and GLM-5.2 at $0.05 per PTU per minute, achieving costs up to 90% lower than Claude Opus 4.8. This new model provides predictable pricing and a 99% uptime SLA, catering to companies transitioning to open weight models for production workloads.

