RLHF
RLHF, reinforcement learning from human feedback, tunes a model with reinforcement learning using a reward model trained on human ratings of its outputs.
RLHF, or reinforcement learning from human feedback, trains a reward model on data in which people choose or rank the better of several model outputs, then uses reinforcement learning to adjust the model so that it maximizes that reward.
It is a leading post-training method for shaping large language models to follow human intent and was used in developing conversational AI such as ChatGPT. Unlike supervised fine-tuning on human-written example answers, it uses people's relative preferences as the training signal.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
RLHF, reinforcement learning from human feedback, tunes models using people's ratings of their answers.