🍊 The Latent Field

标签: opd

此标签下有4条笔记。

  • 2026年8月11日

    Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

    • source
    • paper
    • distillation
    • on-policy-kd
    • opd
    • grpo
    • reasoning
    • code
    • rl
  • 2026年8月11日

    Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

    • source
    • paper
    • distillation
    • on-policy-kd
    • opd
    • reasoning
    • recipe
    • post-training
  • 2026年8月10日

    On-Policy Distillation of Language Models

    • source
    • paper
    • distillation
    • on-policy-kd
    • opd
    • gkd
    • post-training
  • 2026年3月08日

    On-policy KD

    • post-training
    • distillation
    • on-policy-kd
    • opd

The Latent Field · An AI knowledge atlas built with Quartz © 2026