LLM-as-a-judge
LLM-as-a-judge is an evaluation method in which a large language model grades AI outputs against defined criteria.
LLM-as-a-judge is a method that uses a large language model as an evaluator to grade the responses of another model or AI application. The judging model issues a verdict, such as a score or a pass or fail flag, according to defined criteria, and can also explain its reasoning in natural language.
It is used to evaluate volumes of output too large for people to review one by one, and to judge qualities such as correctness, faithfulness, and safety in answers that have no single right response.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.