radixark/miles#1825. The scripts and image
below come from that branch. Results and background are in the
LMSYS day-0 write-up.
1. Model Introduction
Kimi-K3 pairs two attention mechanisms in one stack, KDA and MLA chosen per layer, with an 896-expert latent MoE at top-16. The checkpoint ships in MXFP4. miles trains it with native LoRA adapters rather than full fine-tuning, which is what makes the recipe fit at all: the base weights stay frozen and only the adapters carry gradients. The adapters are implemented under TP, EP, PP and CP, with shared-A and per-expert-B factors across the 896 experts, and they are exported to the rollout engines as HF-named chunks over CUDA IPC. Key highlights:- Two attention types per layer: KDA and MLA, with an attention-residual snapshot bank.
- 896-expert latent MoE, top-16,
moe_latent_size=3584, plus a shared expert. - LoRA RL, not full fine-tuning. Rank 16 by default, 32 in the validated run.
- Colocated rollout: trainer and SGLang share the GPUs, with adapters synced over CUDA IPC.
- MXFP4 checkpoint upcast to BF16 once, offline.
2. Supported Variants
The name sets the checkpoint paths under
--model-dir and the megatron_model_type;
--train-mode lora|full picks the recipe.
Architecture, from scripts/models/kimi-k3.py: hidden 7168, FFN 33792, 96 attention heads,
kv_channels=256, MLA with q_lora_rank=1536 / kv_lora_rank=512 /
qk_head_dim=128 / qk_pos_emb_head_dim=64 / v_head_dim=128, 896 experts at
moe_ffn_hidden_size=3072, shared expert 6144, vocab 163840, no position embedding.
3. Environment Setup
Use theradixark/miles:dev image with the Megatron and SGLang changes from
radixark/Megatron-LM#94 and
sgl-project/sglang#37704 until they merge.
The only external asset is the native MXFP4 checkpoint; the BF16 dequantization and the
torch_dist conversion derive from it. scripts/run_kimi_k3.py names the three by model:
{model_dir}/{model_name}, {model_dir}/{model_name}-bf16 and
{model_dir}/{model_name}-bf16_torch_dist, each overridable with --hf-checkpoint,
--bf16-checkpoint and --ref-load.
3.1 Four-layer prune (one node)
Pinaster/Kimi-K3-4layer is the first dense layer plus three MoE layers of the release;
Pinaster/Kimi-K3-4layer-64experts keeps only the first 64 routed experts of each MoE layer
(the router is sliced to match) so full-parameter training fits one node’s host memory. The
run-ci-model-scripts recipes train the 64-expert prune; the 896-expert prune is for layouts
that depend on the release’s expert count.
3.2 Full model
Download the release into{model_dir}/Kimi-K3 and dequantize it shard by shard across the
nodes (--shard-rank/--num-shards, then --finalize-only once):
torch_dist conversion runs on 32 ranks; the output re-shards at load, so the conversion
layout does not have to match the training one:
4. Launch
The launcher usesall-linear, which selects Kimi K3’s HF defaults: MLA query/KV down projections, attention output projections, dense/shared MLPs, and routed experts. attn,mlp selects the same set. Explicit HF targets can omit language_model.model.layers.*.block_sparse_moe.experts.*.w2 to leave routed-expert down projections frozen; other layouts remain unsupported by the native backend.
--train-mode lora (default) or full. Validated on 16 nodes × 4 GPUs: one container per
node, a ray cluster across them, export MILES_SCRIPT_EXTERNAL_RAY=1, then:
--rollout-max-concurrency 8 is passed explicitly: the field default is 64, and the
validated runs pin 8. Off the validated 64 GPUs the full model needs --tp-size-override
(and --ep-size-override).
For a single-node smoke test use the default --model-name Kimi-K3-4layer: the trainer layout
follows the GPU count (TP8/EP8 on 8 GPUs) and the rollout TP/EP default to it; --rollout-tp-size 16 --rollout-ep-size 1 on two nodes reproduces the TP16 Marlin layout of the full recipe.
5. Recipe Configuration
5.1 Parallelism
The resolved config at startup should showexpert_model_parallel_size 8,
max_tokens_per_gpu 8192, colocate_memory_peak_device gpu and
lora_base_cpu_backup True. Checking those four lines is the fastest way to confirm the
run came up in the intended shape.
5.2 LoRA
Rank 32 / alpha 64 in the validated run; the script defaults to 16 / 32. Adapters attach to attention output and both MLA down-projections, the dense MLP, and both expert projections:5.3 Rollout
Rollout is colocated: the trainer and SGLang share GPUs, and adapters sync over CUDA IPC as HF-named chunks.lora_base_cpu_backup keeps a host copy of the frozen base so the
GPU copy can be reclaimed during rollout.
5.4 What a healthy run looks like
From the GB300 validation runs:- 11 to 13 minutes per rollout cycle
- trainer allocated memory returns to about 91 GB after every weight sync
rollout/raw_rewardbetween 0.5 and 0.75 from rollout 0eval/aime0.37 to 0.43 at eval@0 (that spread is temperature-0 nondeterminism), rising by at least 0.06 by eval@9; the measured run went 0.367 to 0.467
5.5 Notable quirks
- The image preloads a small
shm_unlinkshim through/etc/ld.so.preload. It tolerates a benign PyTorch CUDA-IPC unlink race that otherwise aborts colocated weight sync at scale. - Weight conversion for K3 lives in
miles/backends/megatron_utils/megatron_to_hf/kimi_k3.py, and the model itself inmiles_plugins/models/kimi_k3/.

