- No environment — single-turn RLVR: a prompt dataset scored by the built-in rule-based rewards (math, ifbench, …) or a custom reward function.
- Your own environment — plug your code into one of the three rollout layers described in Integration shapes; most environments sit in the agent function, with the session server recording tokens (see Agentic Rollout (TITO)).
- An external ecosystem — adopt a prebuilt connector from the table below, spanning coding, computer-use and tool-calling agents; connectors occupy the same three layers.
Sandbox providers are a different axis: they provision the task containers
inside a connector rather than occupying a rollout layer. Which providers work
with which connector, and how to set one up, is
Sandbox Providers.
Everything above is experimental, and listed alphabetically.
Integration shapes
The rollout stack is three nested plug-in layers (see Customization): each column in the table below wraps the one to its left, so replacing an outer layer also takes over everything an inner one would. A connector replaces exactly one layer. ✓ = the external framework takes it over; ○ = stays in Miles.
¹ Typically by speaking SGLang’s native
/generate (token IDs in and out)
rather than the session-server chat endpoint Miles’ own recording uses.
² The environment may grade an episode itself (Harbor and τ-bench do); the
score still enters training through Miles’ Sample.reward / RM hooks, and
group-level reward handling stays in Miles.
³ HUD’s harness records per-turn token ids and sampling logprobs itself when
the inference server returns them, so the connector’s job is stitching those
into one training sequence rather than recording. Computer-use observations
are screenshots, which Miles’ session-server recording does not carry yet —
once it does, this connector can also sit in the agent function.
