Nemotron Labs Audex 30B A3B
nvidia/Nemotron-Labs-Audex-30B-A3B
Open Source · chat · open-weights
Open
Alert me on changes
Context
1M
Max output
—
Weights
Open
API $/1M
—
Modalities
text
Released
06 Jul 2026
License: other · nvidia/Nemotron-Labs-Audex-30B-A3B
AI summary
● machine-written
NVIDIA Nemotron-Labs-Audex-30B-A3B: Unified Audio-Text LLM
Nemotron-Labs-Audex-30B-A3B is a unified audio-text language model built on the Nemotron-Cascade-2 backbone, combining 30B total parameters with 3B activated per token in a Mixture-of-Experts architecture. It supports audio understanding, speech recognition, text-to-speech, and audio generation while maintaining text reasoning and long-context capabilities. The model operates in both thinking and instruct modes with up to 1M token context length.
What's new
- Extends vocabulary for discrete audio tokens for speech and general audio outputs
- Adds audio encoder for speech and general audio inputs
- Supports both thinking and instruct (non-thinking) modes with reasoning tags
- Maintains text reasoning capabilities from Nemotron-Cascade-2 backbone
- Hybrid Mamba-Transformer architecture with 52 layers
Best for
Audio understanding and speech recognitionText-to-speech and audio generationMulti-modal reasoning tasks requiring audio-text understandingAgentic AI systems with audio capabilities
Source: https://huggingface.co/nvidia/Nemotron-Labs-Audex-30B-A3B