Skip to main content
The complete Kimi-K3 LoRA RL implementation is open at the Miles pull request: radixark/miles#1825. The scripts and image below come from that branch. Results and background are in the LMSYS day-0 write-up.

1. Model Introduction

Kimi-K3 pairs two attention mechanisms in one stack, KDA and MLA chosen per layer, with an 896-expert latent MoE at top-16. The checkpoint ships in MXFP4. miles trains it with native LoRA adapters rather than full fine-tuning, which is what makes the recipe fit at all: the base weights stay frozen and only the adapters carry gradients. The adapters are implemented under TP, EP, PP and CP, with shared-A and per-expert-B factors across the 896 experts, and they are exported to the rollout engines as HF-named chunks over CUDA IPC. Key highlights:
  • Two attention types per layer: KDA and MLA, with an attention-residual snapshot bank.
  • 896-expert latent MoE, top-16, moe_latent_size=3584, plus a shared expert.
  • LoRA RL, not full fine-tuning. Rank 16 by default, 32 in the validated run.
  • Colocated rollout: trainer and SGLang share the GPUs, with adapters synced over CUDA IPC.
  • MXFP4 checkpoint upcast to BF16 once, offline.

2. Supported Variants

The name sets the checkpoint paths under --model-dir and the megatron_model_type; --train-mode lora|full picks the recipe. Architecture, from scripts/models/kimi-k3.py: hidden 7168, FFN 33792, 96 attention heads, kv_channels=256, MLA with q_lora_rank=1536 / kv_lora_rank=512 / qk_head_dim=128 / qk_pos_emb_head_dim=64 / v_head_dim=128, 896 experts at moe_ffn_hidden_size=3072, shared expert 6144, vocab 163840, no position embedding.

3. Environment Setup

Use the radixark/miles:dev image with the Megatron and SGLang changes from radixark/Megatron-LM#94 and sgl-project/sglang#37704 until they merge. The only external asset is the native MXFP4 checkpoint; the BF16 dequantization and the torch_dist conversion derive from it. scripts/run_kimi_k3.py names the three by model: {model_dir}/{model_name}, {model_dir}/{model_name}-bf16 and {model_dir}/{model_name}-bf16_torch_dist, each overridable with --hf-checkpoint, --bf16-checkpoint and --ref-load.

3.1 Four-layer prune (one node)

Pinaster/Kimi-K3-4layer is the first dense layer plus three MoE layers of the release; Pinaster/Kimi-K3-4layer-64experts keeps only the first 64 routed experts of each MoE layer (the router is sliced to match) so full-parameter training fits one node’s host memory. The run-ci-model-scripts recipes train the 64-expert prune; the 896-expert prune is for layouts that depend on the release’s expert count.

3.2 Full model

Download the release into {model_dir}/Kimi-K3 and dequantize it shard by shard across the nodes (--shard-rank/--num-shards, then --finalize-only once):
The torch_dist conversion runs on 32 ranks; the output re-shards at load, so the conversion layout does not have to match the training one:

4. Launch

The launcher uses all-linear, which selects Kimi K3’s HF defaults: MLA query/KV down projections, attention output projections, dense/shared MLPs, and routed experts. attn,mlp selects the same set. Explicit HF targets can omit language_model.model.layers.*.block_sparse_moe.experts.*.w2 to leave routed-expert down projections frozen; other layouts remain unsupported by the native backend. --train-mode lora (default) or full. Validated on 16 nodes × 4 GPUs: one container per node, a ray cluster across them, export MILES_SCRIPT_EXTERNAL_RAY=1, then:
--rollout-max-concurrency 8 is passed explicitly: the field default is 64, and the validated runs pin 8. Off the validated 64 GPUs the full model needs --tp-size-override (and --ep-size-override). For a single-node smoke test use the default --model-name Kimi-K3-4layer: the trainer layout follows the GPU count (TP8/EP8 on 8 GPUs) and the rollout TP/EP default to it; --rollout-tp-size 16 --rollout-ep-size 1 on two nodes reproduces the TP16 Marlin layout of the full recipe.

5. Recipe Configuration

5.1 Parallelism

The resolved config at startup should show expert_model_parallel_size 8, max_tokens_per_gpu 8192, colocate_memory_peak_device gpu and lora_base_cpu_backup True. Checking those four lines is the fastest way to confirm the run came up in the intended shape.

5.2 LoRA

Rank 32 / alpha 64 in the validated run; the script defaults to 16 / 32. Adapters attach to attention output and both MLA down-projections, the dense MLP, and both expert projections:
The 896 experts share one A factor and carry per-expert B factors, which is what keeps the adapter count tractable at this expert width.

5.3 Rollout

Rollout is colocated: the trainer and SGLang share GPUs, and adapters sync over CUDA IPC as HF-named chunks. lora_base_cpu_backup keeps a host copy of the frozen base so the GPU copy can be reclaimed during rollout.

5.4 What a healthy run looks like

From the GB300 validation runs:
  • 11 to 13 minutes per rollout cycle
  • trainer allocated memory returns to about 91 GB after every weight sync
  • rollout/raw_reward between 0.5 and 0.75 from rollout 0
  • eval/aime 0.37 to 0.43 at eval@0 (that spread is temperature-0 nondeterminism), rising by at least 0.06 by eval@9; the measured run went 0.367 to 0.467
The memory figure is the one to watch. If allocated does not return to its baseline after a weight sync, the adapter export is leaking and the run will die later rather than sooner.

5.5 Notable quirks

  • The image preloads a small shm_unlink shim through /etc/ld.so.preload. It tolerates a benign PyTorch CUDA-IPC unlink race that otherwise aborts colocated weight sync at scale.
  • Weight conversion for K3 lives in miles/backends/megatron_utils/megatron_to_hf/kimi_k3.py, and the model itself in miles_plugins/models/kimi_k3/.

6. Pairs Well With