🍊 The Latent Field

标签: grpo

此标签下有10条笔记。

  • 2026年9月02日

    01 Sync GRPO Run Overview

    • projects
    • source-reading
    • verl
    • grpo
  • 2026年9月02日

    09 PPO and GRPO Codepath

    • projects
    • source-reading
    • verl
    • ppo
    • grpo
  • 2026年9月02日

    Lab 03 - Compare GRPO and PPO

    • projects
    • verl
    • experiment
    • ppo
    • grpo
  • 2026年8月31日

    DeepSeekMath

    • source
    • paper
    • reasoning
    • grpo
    • rlvr
    • math
    • data
  • 2026年8月26日

    MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

    • source
    • paper
    • distillation
    • on-policy-kd
    • grpo
    • rl
    • capability-integration
  • 2026年8月12日

    Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

    • source
    • paper
    • distillation
    • on-policy-kd
    • self-distillation
    • reasoning
    • grpo
  • 2026年8月11日

    Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

    • source
    • paper
    • distillation
    • on-policy-kd
    • opd
    • grpo
    • reasoning
    • code
    • rl
  • 2026年7月15日

    Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective

    • source
    • paper
    • rlvr
    • entropy
    • grpo
    • ppo
    • reinforcement-learning
    • reasoning
    • coding
  • 2026年7月14日

    Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

    • source
    • paper
    • agent
    • reinforcement-learning
    • post-training
    • grpo
    • asynchronous-rl
  • 2026年3月07日

    GRPO

    • post-training
    • grpo
    • reasoning

The Latent Field · An AI knowledge atlas built with Quartz © 2026