Constitutional AI
Constitutional AI is a training method developed by Anthropic in which an AI critiques and revises its own answers according to written principles.
Constitutional AI is a training method that Anthropic published in 2022. It sets out a list of principles called a constitution, has the model critique and revise its own answers against those principles, and then trains on the results and on AI-generated evaluations.
Its distinguishing feature is the use of AI feedback instead of human ratings when training models to be less harmful. Anthropic has used it to train its Claude models.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.