Skip to main content
Miles supports NVIDIA’s Nemotron-3 line: a Mamba + Attention hybrid that, in the Super tier, adds MoE and ships natively in FP8. All three variants load via the Megatron AutoBridge path, so there is no offline HF → torch_dist conversion step.

Variants

Fastest path to train

Nemotron-3-Nano (dense, 4 B) is the smallest and runs on a single 8-GPU node:
See the Nemotron-3-Nano page for the dense walkthrough, Nemotron-3-Nano MoE for the 30 B MoE variant, and Nemotron-3-Super for the FP8-native 120 B-A12B recipe.

Which variant do I pick?

Pairs well with