🍊 The Latent Field
Search
搜索
暗色模式
亮色模式
学习
探索
标签: instruction-tuning
阅读模式
Hide sidebars
Show sidebars
此标签下有5条笔记。
2026年8月27日
Training language models to follow instructions with human feedback
source
paper
rlhf
instruction-tuning
reward-model
ppo
alignment
2026年8月27日
Self-Instruct
source
paper
instruction-tuning
synthetic-data
2026年5月29日
Finetuned Language Models Are Zero-Shot Learners
source
paper
instruction-tuning
sft
2026年5月29日
Multitask Prompted Training Enables Zero-Shot Task Generalization
source
paper
instruction-tuning
multitask-learning
2026年3月01日
Instruction Tuning
post-training
instruction-tuning