Enable it
Add these arguments to an existing text-only GRPO training recipe:--score-centering-is tis for weights clipped at --score-centering-tis-clip (default 2). Choose mis to retain ratios in [--score-centering-mis-low, --score-centering-mis-high] (defaults 0.5 and 5), setting other weights to zero. These weights are centered together with the score. Use these options instead of --use-tis or a custom TIS function.
The existing entropy and reference-KL loss options remain available. They are separate regularizers; the score-centering identity applies to the policy-gradient term.
With filtered sampling, sampling-support replay rules prohibit reference KL; entropy is computed over the recorded support.
The loss keeps full-vocabulary actor scores for reference KL, matching the reference forward even if called directly with replayed candidates. This does not remove the startup restriction above.
Reference-KL tokens with non-finite probabilities or absolute log-probability ratios above 40 are excluded before exponentiation to keep gradients finite. With unbiased KL, this also applies to the train/rollout ratio. train/kl_invalid_fraction reports the excluded fraction.
Training logs include train/train_rollout_logprob_abs_diff and train/train_rollout_kl. The latter uses the same masked, sampled-token k3 estimator of KL(rollout || train) as the policy loss. It is a detached diagnostic and is emitted even when reference-KL regularization is disabled.
How it works
When the rollout distribution differs from the current trainer, even a constant reward can produce an unwanted average policy update. Score centering subtracts the expected weighted score under the rollout distribution. Letp be the current trainer distribution, q the distribution that actually sampled the token, H the stored candidates, A the detached advantage, and f the selected importance-weight function. The implementation computes:
H, the approximation models q as rho * p. The sampled token always uses its recorded q[token], including when it is outside H. All weights and correction coefficients are detached. Full-distribution centering cancels the expected constant-reward gradient; the top-k version approximates the true tail and does not guarantee exact cancellation for an arbitrary tail.
--rollout-top-logprobs-num K plays a different role in the two sampling modes:
- Unfiltered sampling (
--rollout-top-p 1.0 --rollout-top-k -1): Miles automatically usesselectedmode. The sampler draws from the full vocabulary, so Miles records only its topKtokens asHand the tail model covers the rest.Kis a real truncation and must be set; a largerKleaves less to the tail model. - Filtered sampling (top-p/top-k): Miles automatically uses
supportmode when candidate recording is enabled. The sampler draws only from the realized support, and Miles records that whole support with its post-filter probabilities.His then the full sampling distribution and nothing is approximated.Konly sizes the arrays: it must be at least--rollout-top-kand hold the realized support, which cutoff ties can make larger; a larger support fails validation. The trainer renormalizes its log probabilities over the same support.
The first row is the sampler under filtering and the second the sampler without filtering. Miles-managed rollout workers set
SGLANG_RETURN_ORIGINAL_LOGPROB=0, so the third row does not occur. Each configuration reads:
For support replay without reference KL, the trainer gathers only sampled and support logits, normalizes over that support, and saves activations proportional to the candidate count rather than the vocabulary size. Other paths retain full-vocabulary normalization, reducing normalization scalars and selected logits across tensor-parallel ranks. It excludes padded vocabulary entries and computes probabilities in float32 for BF16/FP16 models.
--log-probs-chunk-size controls the temporary computation size; --recompute-loss-function can trade computation for saved activations. No full vocabulary is gathered across ranks.
Rollout and data contract
Native SGLang generation, the legacy rollout path, and both session-server versions request candidate probabilities at generation time when--rollout-top-logprobs-num is positive. Miles derives the logprob mode at startup from --rollout-top-p and --rollout-top-k. On these training requests, the configured count and derived mode override any client-supplied top_logprobs, top_logprobs_num, or sampling_logprobs_mode. A count of zero disables candidate recording without requesting support-wide probabilities. Each Sample carries:
rollout_topk_token_ids: int32 array shaped[response_length, k].rollout_topk_log_probs: float32 array of the same shape.rollout_log_probs: the actual sampled-token log probabilities.
POST /sessions with JSON body {"evaluation": true}.
Unused candidate slots and non-trained observation rows contain token ID -1 and log probability -inf. Tool-observation masks, multi-turn merging, retries, trailing-token trimming, and truncation preserve row alignment. Session serialization retains both arrays. For score centering in support mode, training batches share the recorded support IDs when every sample has exactly the same candidate order and prefix padding. A per-row candidate count preserves observation rows and masked generated rows; any mismatch keeps the original arrays for the whole batch. The trainer reconstructs only its context-parallel rows. The source Samples and session payloads still retain both representations.
Custom rollout producers must supply these fields with probabilities from the actual generation call and no repeated non-negative token ID in a row. Conversion always rejects missing candidate IDs, candidate log probabilities, or sampled-token log probabilities, including from custom producers. With --ci-test, Miles also validates the complete sample before training, rejecting duplicate candidates, sampling-support mismatches, and disagreeing sampled/candidate probabilities. This full-sample validation is skipped in normal training because sorting and support matching are expensive on long responses. Rescoring old rollouts with newer weights is not a substitute. With the feature disabled, requests and the session wire format remain unchanged.
Supported configurations and limits
- The shared loss is wired into Megatron and FSDP. Candidate selection supports tensor parallelism, packed (
thd) and padded (bshd) zigzag context parallelism, and packed all-gather context parallelism. - Sampling requires a fixed positive temperature and
min_p=0on every call. Filtered sampling, including top-p filtering, automatically uses support mode and requires a positivetop_k, for example--rollout-top-p 0.9 --rollout-top-k 64 --rollout-top-logprobs-num 128. Global filtered rollout settings automatically enable sampling-support replay; per-request overrides are checked when each request is built. See the sampling-support replay guide for its request and server requirements. - Filtered sampling requires SGLang with support log probabilities (SGLang PR #40932, included in the
sglang-milesbranch by #41047, merge commitae04cb14046896b6d453758c5769d639deedd353or a descendant containing it). External servers must setSGLANG_RETURN_ORIGINAL_LOGPROB=0like Miles-managed workers. OpenAI session responses must expose the SGLang fields above inchoices[0].meta_info; a generic OpenAI-compatible server without that metadata is insufficient. - The SGLang router must include sgl-router-for-miles #21 (
df2c790; seeGit Commitinsmg --version-verbose), which raises the chattop_logprobslimit from 20 to 128 and forwardssampling_logprobs_mode. - Constrained/custom sampling, speculative decoding, true-on-policy mode, OPD, multi-LoRA/Tinker losses, sequence masking, custom policy-loss reducers, custom train-data converters, and logprob recomputation via prefill are rejected. Multimodal token expansion is not supported. The initial advantage estimator is GRPO.
- Retaining
k=128uses about 1 KiB per response position for the two arrays, before transport overhead. Eligible support-mode training batches replace the duplicate ID array with one int32 count per response position, saving4 * (k - 1)bytes per position in training transport and storage. Source Samples and session payloads keep the original storage cost. Largerkimproves the unfiltered tail approximation at additional storage and compute cost.
Metrics and verification
Training logs includesc_correction, sc_train_head_mass, sc_rollout_head_mass, sc_tail_ratio, sc_importance_weight, train_rollout_kl, and train_rollout_logprob_abs_diff, under the usual train/ namespace. Small head mass means more of the distribution is approximated by the tail model. Large tail ratios indicate a substantial mismatch in remaining mass.
For filtered sampling, the head covers the full support, so both head masses and the tail ratio should be approximately one.
The numerical tests compare gradients against an independent dense-distribution oracle for all three weighting modes, including sampled tokens outside the head, constant rewards, and tiny tails. Pipeline tests exercise real native/session producers, serialization, data-parallel splitting, masks, advantage computation, regularization, and checkpointed loss scaling. Run the focused tests in the repository’s test environment:
SGLANG_RETURN_ORIGINAL_LOGPROB=0 on the server before starting it. The probe checks unfiltered native and OpenAI response metadata using the production candidate collector and validator, and checks temperature scaling at 0.7, 1.0 and 1.3. It warms the shared prompt first so that cached and uncached prefills do not confound the temperature comparison. Set MILES_LIVE_SCORE_CENTERING_ARTIFACT_DIR to keep the raw responses.
