golden set
A golden set is a curated collection of inputs and human-verified expected outputs used as a reference for evaluating AI models or agents.
A golden set, also called a golden dataset, is a curated collection of representative inputs paired with expected outputs that people have reviewed and confirmed as correct. It serves as the reference against which the outputs of an AI model or agent are judged.
In machine learning and large language model evaluation, teams rerun the same golden set whenever they change a model, prompt or configuration so they can compare results and catch regressions. Unlike a public benchmark, a golden set is often built by an organization for its own tasks and requirements.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.