NVIDIA Nemotron Labs 3 Puzzle 75B A9B BF16
nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16
Open Source · chat · open-weights
Open
Alert me on changes
Context
—
Max output
—
Weights
Open
API $/1M
—
Modalities
text
Released
24 Jun 2026
License: other · nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16
AI summary
● machine-written
NVIDIA Nemotron Labs 3 Puzzle 75B A9B BF16 - Compressed Hybrid MoE Model
NVIDIA Nemotron Labs 3 Puzzle 75B A9B BF16 is a compressed variant of the Nemotron-3-Super model, reducing parameters from 120.7B total / 12.8B active to 75.3B total / 9.3B active while preserving the 88-block hybrid Mamba-Transformer MoE architecture. The model supports Multi-Token Prediction for faster text generation and achieves up to 2.03x server throughput improvement on decode-heavy workloads. It is released as open weights under a BF16 precision format.
What's new
- Reduces from 120.7B/12.8B active parameters to 75.3B/9.3B active via iterative neural architecture search
- Preserves 88-block hybrid layout: 40 Mamba, 40 MoE, 8 attention blocks
- Achieves 2.03x throughput boost on 8K/64K token scenarios at ≥100 tok/s user threshold
- Enables 8 concurrent 1M-token requests on single H100 (vs. 1 for parent model)
- Supports Multi-Token Prediction for faster text generation
Best for
Serving multiple concurrent long-context (1M token) inference requestsDecode-heavy workloads with high user throughput requirementsMemory-constrained deployment scenarios requiring parameter reduction
Source: https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16