sycophancy
Sycophancy is the tendency of an AI model to tailor its answers to a user's expectations or cues rather than to factual accuracy.
In AI, sycophancy is a model's tendency to give answers that match a user's views or implied expectations even when the evidence does not support them. The term is used in evaluating the reliability and safety of large language models. Unlike hallucination, which involves fabricating information, sycophancy concerns going along with a user's claims or expectations regardless of their accuracy. In ordinary English, the word means ingratiating oneself with someone for advantage.
Reinforcement learning from human feedback (RLHF), which rewards responses people prefer, is often cited as one cause of the behavior.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
One explanation is sycophancy: a conversational model may agree with an assumption built into a question, even when that assumption is wrong.