🍊 The Latent Field

标签: on-policy-kd

此标签下有8条笔记。

  • 2026年8月26日

    MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

    • source
    • paper
    • distillation
    • on-policy-kd
    • grpo
    • rl
    • capability-integration
  • 2026年8月12日

    Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

    • source
    • paper
    • distillation
    • on-policy-kd
    • self-distillation
    • reasoning
    • grpo
  • 2026年8月11日

    MiniLLM: Knowledge Distillation of Large Language Models

    • source
    • paper
    • distillation
    • on-policy-kd
    • post-training
    • generative-model
  • 2026年8月11日

    Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

    • source
    • paper
    • distillation
    • on-policy-kd
    • opd
    • grpo
    • reasoning
    • code
    • rl
  • 2026年8月11日

    Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

    • source
    • paper
    • distillation
    • on-policy-kd
    • opd
    • reasoning
    • recipe
    • post-training
  • 2026年8月10日

    On-Policy Distillation of Language Models

    • source
    • paper
    • distillation
    • on-policy-kd
    • opd
    • gkd
    • post-training
  • 2026年7月22日

    On-Policy Delta Distillation

    • source
    • paper
    • distillation
    • on-policy-kd
    • post-training
    • reasoning
    • code
  • 2026年3月08日

    On-policy KD

    • post-training
    • distillation
    • on-policy-kd
    • opd

The Latent Field · An AI knowledge atlas built with Quartz © 2026