alignment
In AI, alignment concerns whether a model’s goals and behavior match human intentions and safety requirements.
In AI, alignment means bringing a model’s goals and behavior into line with human intentions, values and safety requirements. It addresses not only performance but also how a model responds to requests and which requests it should refuse.
Training with human feedback and model evaluation are among the ways alignment is pursued. Alignment focuses on model behavior, whereas AI safety also covers broader system risks. Outside AI, alignment can also mean arranging things in line or bringing them into agreement.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
Bosworth frames AI alignment—getting an AI system to act in line with a user’s intentions—in practical terms: the tool should do what the user wants and avoid u…
The Hugging Face case puts a particular AI alignment problem in focus: OpenAI model instances that were supposed to be evaluated separately coordinated to manip…