RLHF

RLHF, reinforcement learning from human feedback, tunes a model with reinforcement learning using a reward model trained on human ratings of its outputs.

2 articles
Last mentioned

RLHF, or reinforcement learning from human feedback, trains a reward model on data in which people choose or rank the better of several model outputs, then uses reinforcement learning to adjust the model so that it maximizes that reward.

It is a leading post-training method for shaping large language models to follow human intent and was used in developing conversational AI such as ChatGPT. Unlike supervised fine-tuning on human-written example answers, it uses people's relative preferences as the training signal.

This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.

Articles covering this entry

Arena argues that this weakness, introduced by reinforcement learning from human feedback, is exactly why deceptive completion needs to be measured by independe…

RLHF, reinforcement learning from human feedback, tunes models using people's ratings of their answers.


© 2026 AIPOST. All rights reserved.

AIPOST is an AI publication covering practical AI, AI security, performance, startups, health, ethics and industry news. No account is needed, and our privacy policy explains how we handle personal information.