Skip to main content
Two policies trained against each other in one run: a solver answering a gsm8k question and a verifier ruling on its work. Each rollout yields one sample per policy; each policy has its own trainer and its own inference engines inside the same job.

Quick Start

Eight GPUs: 2 trainer GPUs and 2 single-GPU engines per policy. Needs a cluster backend and a namespace, from MILES_SCRIPT_CLUSTER_BACKEND / MILES_SCRIPT_NAMESPACE or --cluster-backend / --namespace: