🍊 The Latent Field
Search
搜索
暗色模式
亮色模式
学习
探索
标签: rlhf
阅读模式
Hide sidebars
Show sidebars
此标签下有4条笔记。
2026年8月27日
Deep Reinforcement Learning from Human Preferences
source
paper
rlhf
reward-model
preference-learning
reinforcement-learning
2026年8月27日
Training language models to follow instructions with human feedback
source
paper
rlhf
instruction-tuning
reward-model
ppo
alignment
2026年5月29日
Learning to summarize from human feedback
source
paper
rlhf
summarization
reward-model
2026年3月07日
PPO
post-training
rlhf
ppo