1. Model Introduction
MiMo-V2.6-Flash-RL is Xiaomi’s 309B-parameter MoE model (15B active). miles trains its text decoder. Key highlights:- Hybrid attention: 39 sliding-window layers (window 128, learnable attention sink) and 9 global layers, with different KV head counts (8 vs 4) and RoPE bases (1e4 vs 1e7).
- Asymmetric heads: 192-dim Q/K and 128-dim V, V scaled by 0.707, partial RoPE (64 of 192 dims).
- MoE: 256 routed experts, top-8 sigmoid routing with a fixed score-correction bias; layer 0 is dense.
- Compressed checkpoint: FP8 block-quantized dense weights with a fused, kv-head-interleaved
qkv_proj, and MXFP4 experts. The Megatron bridge reads a BF16 split-q/k/v conversion instead.
2. Supported Variants
The vision, audio and MTP modules stay in the checkpoint but are not trained (text-only RL).
3. Environment Setup
3.1 Download + BF16 conversion
prepare (run by the launcher before training) downloads the official checkpoint to <model-dir>/MiMo-V2.6-Flash-RL and converts it unless <model-dir>/<model-name> already holds a finished conversion:
qkv_proj with its own FP8 block scales, splits it into q/k/v in head order, and decodes the MXFP4 experts. The BF16 model is 622 GB.
3.2 SGLang
The BF16 engine needs the following MiMo-V2 fixes. Thesglang-miles branch has them since 8035002, and the radixark/miles:dev image, built from that branch, includes them:
- upstream sgl-project/sglang#40448 (MXFP4 MoE and BF16 router);
- accepting the
splitattention layout of the BF16 conversion; - allocating the sliding-window KV pools inside the memory-saver region, otherwise a colocated engine cannot release its KV cache;
- passing
layer_idto the MoE top-k, otherwise--use-rollout-routing-replaycrashes the CUDA-graph capture.
4. Launch
4.1 Two nodes
Join the second node to a ray head on the first, then launch from the head:--mode sft --prompt-data <chat jsonl> trains SFT instead. It reads a messages column; put a reasoning trace in reasoning_content, since the MiMo chat template renders <think>{reasoning_content}</think>.
5. Recipe Configuration
5.1 Parallelism
TP is capped at 4 by the four global-attention KV heads. Context parallelism is not supported.
--sequence-parallel is on whenever TP > 1. THD packing with --use-dynamic-batch-size is the default; --qkv-format bshd runs one sample per micro-batch.
5.2 Algorithm
GRPO with--eps-clip 0.2 --eps-clip-high 0.28 --entropy-coef 0.00 and --rm-type math on dapo-math-17k. No KL loss: bridge mode has no Megatron-format reference checkpoint for --ref-load.
5.3 Optimizer
Adam at--lr 1e-6. The model’s Adam state (3.7 TB) fits neither the GPUs nor two hosts’ memory, so it streams through node-local NVMe:
5.4 Notable quirks
- Fused attention only:
--attention-backend fused. The learnable sink on the sliding-window layers is implemented by TE’s FusedAttention (THD needs cuDNN ≥ 9.18) and the unfused path, not by FlashAttention. - Sliding window: a SWA layer attends the query plus
sliding_window_size(128) previous keys, 129 in total, as SGLang does; the bridge setswindow_size = (sliding_window_size, 0). The HF modeling code and Megatron-Bridge’smimo_v2_flashattend 128 keys in total (sliding_window_size - 1previous), so a log-prob comparison against them differs at positions past the window. - Dropout: the bridge sets
hidden_dropout = 0. Bridge mode does not copy--hidden-dropoutonto the provider, so a provider left at Megatron’s default (0.1) trains with dropout silently. - Frozen towers and
--check-weight-update-equal: the trainer never sends the vision/audio weights, so exclude them from the post-update check with--check-weight-update-skip-list visual. audio_tokenizer. input_local_transformer. speech_embeddings. projection.mlp..

