本页研究 verl 如何用 Hydra config group、OmegaConf interpolation 和 shell overrides 构造一条训练任务。目标是能从任意 example script 反推出最终 runtime configuration,而不是只抄参数说明。
本页问题
ppo_trainer.yaml如何组合 actor、critic、reference、rollout、reward 和 engine configs?- example script 中的 override 覆盖了哪一层默认值?
_target_、${oc.select:...}与 dataclass conversion 分别在何时生效?- GRPO/PPO、FSDP/Megatron、vLLM/SGLang 的选择由哪些最小参数决定?
- config validation 如何阻止不一致的 batch size、reference policy 和 critic 配置?
Source Anchors
- ppo_trainer.yaml
- actor config
- rollout config
- algorithm dataclasses
- worker configs
- config validation
- print_cfg.py
Configuration Ledger
完整笔记需要建立以下映射:
shell override
-> resolved OmegaConf path
-> dataclass field
-> runtime consumer
-> observable training behavior重点参数包括 adv_estimator、rollout.n、train_batch_size、ppo_mini_batch_size、use_kl_loss、use_kl_in_reward、actor strategy、rollout backend、trainer mode 和 off-policy threshold。
Completion Criteria
- 保存一份最小 GRPO resolved config;
- 标出 algorithm、resource placement、batching 和 performance knobs;
- 解释 generated config 与 source config 的关系;
- 记录至少三类常见配置错误及其 validation path。