Identity: what shares a baseline
Goal: pin down exactly which past values a new number may be compared against — same test, same metric, same rule, same point — so baselines never mix across meanings. The gate compares a number only against earlier numbers of the same kind, from the same test. Two keys decide that:- Run series (the “same test”):
(test_path, backend, suite). Runs differing on any field never share a baseline. A test-file edit does not reset the series (see Notes). - Value within a run:
(metric_key, steps_key, constraint_key, step)— the declaring gate’s literal content plus which point.steps_keyandstepare not redundant: a fanned-out declaration (steps=[0, 1]/steps="all") produces several values in one run — one per selected step — and each must be judged only against its own step’s history, so the literals identify the spec whilestepidentifies the point:- Session TITO metrics put
/v1/or/v2/directly aftertito_session_mismatch_rateat the producer, so wandb and both runs’ history coordinates stay distinct. steps_key/constraint_keyare canonical JSON of the declaration’s rawsteps/constraintliterals: no whitespace, dict keys sorted, list order kept as written, a string keyword stored with its JSON quotes —steps=[0, 1]→[0,1],steps="last"→"last"(quotes included). Built from the raw literal, never the normalized form, so a code-side default change can never silently re-key a series; editing the declaration’s literals changes these keys, so a declaration edit starts a fresh coordinate history by construction.- A field the declaration omitted and the defaults table filled (see Steps & constraint) keys on the table’s literal — the table entry is the declaration source there, so editing a table entry re-keys every declaration that relied on it: one global reset lever for that standard metric’s defaulted baselines, deliberate and heavier than a per-test literal edit.
stepis the point the value came from: stepkfor a per-step value,-1for a whole-series reduction (e.g.steps="last") — a reduced value keys on a constant, never the step it happened to land on, or its history would fragment across runs of different lengths.- Step-0
ppo_klis compared only against past step-0ppo_kl— never against step 1 orgrad_norm.
- Session TITO metrics put
limit for how many recent points to read): recent_trusted_values(test_path, backend, suite, metric_key, steps_key, constraint_key, step, limit).
Steps & constraint: what is compared, and by which rule
Goal: let a test declare, as literal data next to its CI registration, which values of a metric are judged and by what rule — validated at parse time, with missing or non-finite data always surfacing as ERROR rather than a silent skip. A gate declaration composes a step selection and a constraint, both validated at parse time:steps— which value(s) of the metric’s series to compare:"last"(the series’ last point, a whole-series reduction),"all"(every step present), or a list of step indices."all"and a step list fan out to one comparison per step, judged against that step’s own history.- Constraint — whether one value passes against a reference: every constraint is two-sided — the value must land in the corridor
[ref − band_down, ref + band_up], each side’s band written independently asband = max(rel·|ref|, abs_floor); a literal dict of those params (rel_up/abs_floor_up/rel_down/abs_floor_down), at least one param per side written. Bands scale from the reference only, so a deviating value cannot widen its own tolerance. There is no unbounded side — a value far from baseline in the “improving” direction is usually a broken metric, and a trusted run’s values become future baselines, so admitting it would drag the mean; a side meant to be lenient gets a wide band, not no band. - Defaults for standard metrics — a per-
metric_keytable beside the parser (GATE_DEFAULTSinregister.py) suppliesstepsandconstraint, soregister_ci_gate(metric_key="train/ppo_kl")alone is complete. The standard RL metrics usesteps="all", so every captured step is judged and persisted under its own coordinate. Omitted fields pass through the same parser validation; an explicit literal wins, and a metric absent from the table must provide both fields. Table keys stay within the capture whitelist (test-enforced). - Default calibration and resets — band values are shadow-calibration starting points. Changing a table literal re-keys every declaration that relies on it and cold-starts those coordinates (see Identity).
NaN / ±Inf) at a selected coordinate is an ERROR verdict, never a skip — non-finite is judged here, not silently dropped (capture records it faithfully as a strict-JSON string marker the gate-side reader decodes; write_run refuses it at the DB boundary).
A declaration sits at top level of the test file, next to its CI registration — register_*_ci decides where the test runs, register_ci_gate what is judged after it passes. A gate declaration alone does nothing: an unregistered file is never collected.
Roles & data flow
Goal: split the pipeline into three decoupled roles — collector, harness, gate library — connected only by JSONL files and one store, so the training process never blocks on gating and the gate stays a pure, read-only library. Three roles, connected only by JSONL files and one DB — there is no long-lived “metrics manager”; the pipeline is per-test, driven by the harness:- Collector (training process) —
miles.utils.tracking_utils.TrackingManagerfans everylog()out to all enabled backends;WandbBackendandCiHistoryBackendare parallel siblings in that registry, so wandb receives the same data independently and nothing downstream ever reads it back.CiHistoryBackendsnapshots the fixed metric whitelist into per-process JSONL files under the harness-assigned record dir (MILES_CI_GATE_RECORD_DIR, injected by the CI harness; no CLI flag). - Harness / finalizer (CI runner) —
run_suite.pybuilds the store from env (NEON_DATABASE_URL, a CI secret), resolves the baseline-write signal + provenance, and allocates the record dir (CUDA suites only);ci_utils.run_unittest_fileshands each attempt its own record subdir and merges the PASSING attempt’s per-process records into the merged per-run JSONL record (a metric key appearing in several processes gets its series concatenated and sorted by step);ci_utils.run_gate_hookthen assigns identity, runs the gate, and acts on the verdict. - Gate library (pure functions, read-only against storage) —
register.pyparsesregister_ci_gatedeclarations out of the test file’s AST at evaluation time (the call itself is a runtime no-op; nothing registers at runtime),selection.pypicks comparison coordinates,constraints.pyjudges pass/fail,gate.py:evaluate_gatecomposes them over the store’s baseline read.
NaN / ±Inf) is real evidence of the run and is recorded faithfully, encoded in the JSONL as the string marker "NaN" / "Infinity" / "-Infinity" so every line stays strict JSON (the gate-side reader decodes markers back to floats). Judging non-finite values is the gate’s job, not the recorder’s. A wrong type (non-int/float) is an authoring bug, not run evidence, and still fails loud at capture.
The gate: drift against trusted history
Goal: judge every declared coordinate against its own trusted past, so slow drift gets caught while a fresh baseline can seed itself. After a test passes, each comparison coordinate’s value is judged by its spec’s constraint against the coordinate’s own history:- Historical gate — activates with ≥1 trusted point at the coordinate.
ref= mean of the coordinate’s trusted values. Catches drift. - Cold start (0 trusted): the gate is inactive — not an error. Zero active checks means the run is vacuously trusted: that is how a fresh baseline gets seeded (recover a poisoned seed via
mark_untrusted).
steps="all" or a step list) contributes one verdict per step; the run is trusted iff every coordinate’s active checks pass.
The gate’s data input is the run’s merged per-run JSONL record:
- Merged, per-run: three processes (the
train.pydriver, the training actor’s main rank, the rollout manager) callinit_tracking, each snapshotting to its own record file. In practice each whitelisted key has one logging owner, so the merge is a plain union; a key that does appear in several files just gets its series concatenated and step-sorted. - JSONL (JSON Lines): one self-contained JSON line per metric —
{"metric": <key>, "series": [[step, value], ...]}— each line stands alone, so a process killed mid-run still leaves a parseable record. Capture writes the per-process files, the merge produces this one, and the gate only reads it (parse_merged_record, decoding the non-finite string markers back to floats).
(test_path, backend, suite); it and the value coordinate (metric_key, steps_key, constraint_key, step) are defined in the Identity section above.
Storage: two backends, two tables
Goal: persist runs and metric values behind oneMetricHistoryStore contract so gate code stays backend-agnostic (SQLite offline, Neon in CI), with the write boundary guaranteeing only finite values ever enter a baseline.
Backends — one MetricHistoryStore contract, two implementations; callers see only the contract, and for the same inputs both backends must persist the same run and metric fields, return the same trusted baseline rows in newest-first order, and revoke trust for the same runs:
SQLiteMetricHistoryStore— the local/offline backend, for unit tests and in-process development.NeonMetricHistoryStore— the hosted Postgres backend, for CI/prod.
write_run(...)— persists one CI run: its identity/provenance, its run-leveltrustedflag, and all metric values from that run. It rejects (raises on) non-finite metric values before persisting anything: the DB is the write boundary where validity is enforced, soNaN/±Infnever enter a baseline — upstream they are gate-side ERROR evidence, not storable measurements.recent_trusted_values(...)— the historical-gate baseline read: the newest trusted values for one exact run series and one exact value coordinate.mark_untrusted(...)— flips matching runs totrusted = falsebyrun_id,github_run_id, orcommit_sha, so the next baseline read excludes those runs without deleting rows.
runs— one row per CI run of one series: the identity above + provenance (commit_sha,pr_number,github_run_id,github_run_attempt,event_name,ref) +created_at+trusted(run-level).metric_values— one row per value:run_idFK +(metric_key, steps_key, constraint_key, step)+value.- The baseline read is served by the composite index
runs(test_path, backend, suite, trusted, created_at DESC).
NeonMetricHistoryStore never issues DDL). Old-row cleanup policy is a later operational concern, not part of the M0/M1 substrate.
Inspect hosted history
Repository writers can use$neon-access to query or repair hosted metric history. For example:
Trust, cleanup, who writes
Goal: keep the baseline self-protecting — only baseline-writing runs (with recorded provenance) write at all, a run whose metrics fail the gate is still persisted but flaggedtrusted = false so it never enters the baseline, and a point later found bad is revoked by one flag flip instead of deletion.
- A run is
trustediff it passed all active gates. A drifting run is still recorded, withtrusted = false, so it can’t drag the baseline. A test that fails then passes on retry is gated on its passing attempt’s metrics and trusted normally — needing a retry is not itself a trust penalty. - Clean a bad point:
mark_untrusted=UPDATE runs SET trusted = falseon the run. The next gate read excludes it immediately — no rebaseline, no row deletion. - Nightly and weekly runs write baselines — either an explicitly mapped
schedulecron (onmain, post-merge) or a PR carrying thenightlylabel (the PR’s own pre-merge code). Provenance (event_name,pr_number) records whether a writer was scheduled or PR-triggered, so a label-PR baseline can bemark_untrusted’d if it turns out bad. Ordinary PR runs and explicitly called release runs are read-only and only shadow; frozen release dependency SHAs must not enter the rolling baseline. - What one baseline-writing run writes — one
runsrow plus onemetric_valuesrow per value coordinate: two specs sharing a coordinate (identicalsteps+constraintliterals, differing only in policy metadata) collapse to a single row, so a duplicated declaration cannot double-weight the baseline mean; and a file that declares no gate writes nothing at all —run_gate_hookskips the write instead of leaving an emptyrunsrow.
Rollout
Goal: land the gate observe-only first, so it accumulates history and proves its verdicts on real runs before any PR can be blocked; enforcement is a later, reversible switch. Shadow-first: collect, store, and evaluate, but never block a PR initially — a historical-gate failure lands as an untrusted row and is surfaced, not enforced. Enforcement arrives later behind a per-test allowlist + a global kill-switch.Notes
Goal: record accepted caveats and open questions beside the behavior they qualify; planned-but-unimplemented work lives in TODO below.- A test-file edit does not reset the series:
test_file_hashwas dropped from the run-series identity because a tiny edit to a test kept wiping its whole history. A test change that genuinely shifts a metric’s expected level surfaces as gate failures instead; the reset levers are manual —mark_untrustedthe stale runs, or edit the declaration literals (newsteps_key/constraint_key⇒ fresh coordinate). - Baseline writing is selected by the resolved CI policy rather than inferred from
GITHUB_EVENT_NAME; nightly and weekly write, while regular and release runs stay shadow-only. - Open: should a brand-new test’s first baselines need human confirmation before counting as trusted? (v1: no.)
TODO
Goal: collect everything planned but not implemented in one place — a doc-first pass must not conform code to this section; current behavior is everything above it.- Hard gate returns as a pure absolute bound. The removed hard layer mixed two motivations that want different judging logic: a sanity check (“from experience this metric must stay below X” — a plain one-sided limit, no tolerance) and a backstop against implicit drift that the historical gate absorbs (e.g. a logp diff growing a little per PR: each run sits within band of a baseline that itself follows the drift). The old implementation fed
hard_refthrough the same band constraint as a reference value (evaluate_constraint(constraint, value, ref=hard_ref)), giving a sanity limit tolerance semantics it should not have. Target — two unmixed declaration flavors:register_ci_gate(...)withouthard_ref— the relative check against trusted history, exactly as documented above.register_ci_gate(..., hard_ref=X)— a plain absolute bound: the selected value must stay on the right side ofX(an upper or lower limit), no band, no history involved.hard_refis a limit, never a pinned pseudo-history reference value — synthesizing a baseline from it was considered and rejected as hard to implement. It will not read any historical data.hard_refis never filled from the defaults table — an absolute limit is always written explicitly.- Baselines survive both the removal and the return:
hard_refwas policy, never part of the value coordinate, so adding or dropping the absolute flavor never resets a series.
- Finish the M4 sweep outside
tests/e2e/megatron— eligible Megatron CUDA RL-training tests declare the standard one-liners; extend them to the remaining CUDA e2e training tests only for metrics each test’s run actually emits (a spec on a metric missing from the record is an ERROR verdict that untrusts every nightly run — a mis-swept test would never accumulate a baseline); each test owner tunes or vetoes their line in review. - Capture set becomes
TARGET_METRIC_KEYS∪ declared keys — the harness parses specs pre-launch and injects the extras via env. - Self-calibrating constraint — band = k·std of the coordinate’s own history, for heteroskedastic tests; and a
mean(step-average) reduction. - Enforcement — the per-test allowlist + global kill-switch from Rollout; shadow mode is current behavior.

