LLM-as-a-judge

LLM-as-a-judge is an evaluation method in which a large language model grades AI outputs against defined criteria.

1 article
Last mentioned

LLM-as-a-judge is a method that uses a large language model as an evaluator to grade the responses of another model or AI application. The judging model issues a verdict, such as a score or a pass or fail flag, according to defined criteria, and can also explain its reasoning in natural language.

It is used to evaluate volumes of output too large for people to review one by one, and to judge qualities such as correctness, faithfulness, and safety in answers that have no single right response.

This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.

Articles covering this entry

The standard answer in 2025 was LLM-as-a-judge evaluation, or evals: a language model grades each trace against defined criteria and returns both a verdict, suc…


© 2026 AIPOST. All rights reserved.

AIPOST is an AI publication covering practical AI, AI security, performance, startups, health, ethics and industry news. No account is needed, and our privacy policy explains how we handle personal information.