🍊 The Latent Field

标签: ppo

此标签下有6条笔记。

  • 2026年9月02日

    09 PPO and GRPO Codepath

    • projects
    • source-reading
    • verl
    • ppo
    • grpo
  • 2026年9月02日

    Lab 03 - Compare GRPO and PPO

    • projects
    • verl
    • experiment
    • ppo
    • grpo
  • 2026年8月27日

    Proximal Policy Optimization Algorithms

    • source
    • paper
    • reinforcement-learning
    • ppo
    • policy-optimization
  • 2026年8月27日

    Training language models to follow instructions with human feedback

    • source
    • paper
    • rlhf
    • instruction-tuning
    • reward-model
    • ppo
    • alignment
  • 2026年7月15日

    Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective

    • source
    • paper
    • rlvr
    • entropy
    • grpo
    • ppo
    • reinforcement-learning
    • reasoning
    • coding
  • 2026年3月07日

    PPO

    • post-training
    • rlhf
    • ppo

The Latent Field · An AI knowledge atlas built with Quartz © 2026