🍊 The Latent Field

标签: rlhf

此标签下有4条笔记。

  • 2026年8月27日

    Deep Reinforcement Learning from Human Preferences

    • source
    • paper
    • rlhf
    • reward-model
    • preference-learning
    • reinforcement-learning
  • 2026年8月27日

    Training language models to follow instructions with human feedback

    • source
    • paper
    • rlhf
    • instruction-tuning
    • reward-model
    • ppo
    • alignment
  • 2026年5月29日

    Learning to summarize from human feedback

    • source
    • paper
    • rlhf
    • summarization
    • reward-model
  • 2026年3月07日

    PPO

    • post-training
    • rlhf
    • ppo

The Latent Field · An AI knowledge atlas built with Quartz © 2026