🍊 The Latent Field
Search
搜索
暗色模式
亮色模式
学习
探索
标签: ppo
阅读模式
Hide sidebars
Show sidebars
此标签下有6条笔记。
2026年9月02日
09 PPO and GRPO Codepath
projects
source-reading
verl
ppo
grpo
2026年9月02日
Lab 03 - Compare GRPO and PPO
projects
verl
experiment
ppo
grpo
2026年8月27日
Proximal Policy Optimization Algorithms
source
paper
reinforcement-learning
ppo
policy-optimization
2026年8月27日
Training language models to follow instructions with human feedback
source
paper
rlhf
instruction-tuning
reward-model
ppo
alignment
2026年7月15日
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
source
paper
rlvr
entropy
grpo
ppo
reinforcement-learning
reasoning
coding
2026年3月07日
PPO
post-training
rlhf
ppo