Skip to content

Nemotron Labs TwoTower 30B A3B Base BF16

nvidia/Nemotron-Labs-TwoTower-30B-A3B-Base-BF16
Open Source · chat · open-weights
Open Alert me on changes
Context
Max output
Weights
Open
API $/1M
Modalities
text
Released
11 Apr 2026
Download image Share on X Share on LinkedIn
AI summary
● machine-written

NVIDIA Nemotron-Labs TwoTower 30B: Open-weight diffusion model with 2.42× throughput

Nemotron-Labs-TwoTower-30B-A3B-Base-BF16 is an open-weight diffusion language model that separates token generation into two towers: a frozen autoregressive context tower and a trained denoiser tower. It achieves 2.42× higher wall-clock generation throughput than autoregressive baselines while retaining 98.7% of quality on benchmark tasks. The model is governed by the NVIDIA Nemotron Open Model License and runs on 2×H100 GPUs in full two-tower mode.

What's new
  • Dual-tower architecture separates context encoding from token denoising
  • Achieves 2.42× throughput over autoregressive baseline at 98.7% quality retention
  • Supports three inference modes: full two-tower diffusion, mock-AR, and pure AR decoding
  • Denoiser trained on ~2.1T tokens; frozen context tower uses full 25T backbone pretraining
  • Deployed on Nemotron-3-Nano-30B-A3B hybrid backbone with Mamba-2, attention, and MoE
Best for
High-throughput text generation requiring parallel token refinementApplications needing quality-throughput tradeoffs via confidence thresholding (γ parameter)Multi-modal and agentic AI systems using lightweight inferenceTasks where modest quality loss is acceptable for >2× generation speed
Sources

Source: https://huggingface.co/nvidia/Nemotron-Labs-TwoTower-30B-A3B-Base-BF16