Skip to main content
Miles ships a native Megatron recipe for Inkling, Thinking Machines Lab’s 975 B / 41 B-active multimodal mixture-of-experts model: local and global relative attention, the residual ShortConv, the shared-sink router and experts, and the image and audio encoders. The same backend drives both full-parameter and LoRA RL.

Variants

Fastest path to train

Inkling needs 16 nodes of 4× GB300 and the radixark/miles:inkling image:
See the Inkling page for the architecture summary, HF → Megatron conversion, validated parallelism layouts, training attention backends, LoRA RL, and multimodal RL.