reset / step (and optionally evaluate), so any environment speaking the
protocol can serve any trainer.
Miles integrates OpenEnv as an
agent-function integration: a Miles-side agent
function drives the agentic loop — reset(task_id), repeated steps, then
scoring the episode with the task’s own tests — against an unmodified OpenEnv
server, and the score becomes the sample’s reward through a custom reward
hook.
Try it
The maintained end-to-end recipe is Terminal-Bench-2 GRPO inexamples/experimental/openenv.
It gives every episode its own cloud sandbox, built from that task’s official
image so no resident infrastructure is left behind, on any of the
sandbox providers. One shared Docker env
server is supported as well,
for running without any sandbox platform.
Follow the
recipe README
for prompt-data preparation, environment options, launcher flags, and
operational notes.
