Training language models to follow instructions with human feedback
Introduces InstructGPT: GPT-3 fine-tuned on demonstrations and human feedback rankings via RL to align language models with user intent.
Larger language models are not inherently better at following user intent and can produce untruthful, toxic, or unhelpful outputs. The authors align models by fine-tuning GPT-3 on labeler-written demonstrations, then further fine-tuning with reinforcement learning from human feedback using rankings of model outputs, producing InstructGPT. In human evaluations, outputs from the 1.3B-parameter InstructGPT are preferred over the 175B GPT-3 despite 100x fewer parameters, with improved truthfulness, less toxic generation, and minimal regressions on public NLP datasets.
Based on: Training language models to follow instructions with human feedback · Neural Information Processing Systems