Nemotron Labs TwoTower 30B A3B Base BF16
nvidia/Nemotron-Labs-TwoTower-30B-A3B-Base-BF16
Open Source · chat · open-weights
Open
Alert me on changes
Context
—
Max output
—
Weights
Open
API $/1M
—
Modalities
text
Released
11 Apr 2026
License: other · nvidia/Nemotron-Labs-TwoTower-30B-A3B-Base-BF16
AI summary
● machine-written
NVIDIA Nemotron-Labs TwoTower 30B: Open-weight diffusion model with 2.42× throughput
Nemotron-Labs-TwoTower-30B-A3B-Base-BF16 is an open-weight diffusion language model that separates token generation into two towers: a frozen autoregressive context tower and a trained denoiser tower. It achieves 2.42× higher wall-clock generation throughput than autoregressive baselines while retaining 98.7% of quality on benchmark tasks. The model is governed by the NVIDIA Nemotron Open Model License and runs on 2×H100 GPUs in full two-tower mode.
What's new
- Dual-tower architecture separates context encoding from token denoising
- Achieves 2.42× throughput over autoregressive baseline at 98.7% quality retention
- Supports three inference modes: full two-tower diffusion, mock-AR, and pure AR decoding
- Denoiser trained on ~2.1T tokens; frozen context tower uses full 25T backbone pretraining
- Deployed on Nemotron-3-Nano-30B-A3B hybrid backbone with Mamba-2, attention, and MoE
Best for
High-throughput text generation requiring parallel token refinementApplications needing quality-throughput tradeoffs via confidence thresholding (γ parameter)Multi-modal and agentic AI systems using lightweight inferenceTasks where modest quality loss is acceptable for >2× generation speed
Source: https://huggingface.co/nvidia/Nemotron-Labs-TwoTower-30B-A3B-Base-BF16