Skip to main content
Training VLM on geo3k dataset with multi-turn reasoning with interactive environment feedback, using GRPO. For the dataset, we used the processed version. Use the CUDA 13 Miles image from the installation guide and keep its bundled cuDNN version. The multi-turn rollout is implemented through a custom generate function, overriding the original generate function. In terms of the environment interaction, this example initializes a custom interactive environment with the APIs below.
  • build_env(sample: Sample | None = None, args: Any | None = None, **_) -> Geo3kEnv: constructs the env.
  • reset() -> tuple[dict, dict]: clears internal state.
  • step(response_text: str) -> tuple[dict, bool, dict]: parses the actor’s response text and update the state. Return new observation, a flag that marks whether the task is done, and step_info.
  • format_observation(observation: dict) -> dict: converts an env observation into a chat message.

The reward model is the default math RM. VLM multi-turn geo3k reward Rollout megatron

Reproduce

What each file does

  • examples/geo3k_vlm/multi_turn/run_geo3k_vlm_multi_turn.py: downloads model, sets training/rollout args, and launches the run.
  • examples/geo3k_vlm/multi_turn/geo3k_vlm_multi_turn_config.yaml: specifies max_turns and rollout_interaction_env_path for the multi-turn rollout.
  • examples/geo3k_vlm/multi_turn/rollout.py: custom multi-turn rollout that calls SGLang for token generation, builds loss masks/log_probs, enforces max_turns, and early-stops on max_new_tokens.
  • examples/geo3k_vlm/multi_turn/env_geo3k.py: geo3k tool-calling env that parses <tool_call>{…}</tool_call>, scores math answers, and returns tool feedback per turn.