Humanity's Last Exam
Humanity's Last Exam is a benchmark of difficult expert-level questions used to assess AI models’ knowledge and reasoning.
Humanity's Last Exam (HLE) is a benchmark of difficult questions across specialist fields that tests AI models’ knowledge and reasoning. It was developed by the Center for AI Safety and Scale AI with subject-matter experts and released in 2025.
It was designed with hard questions because leading models had begun to score so highly on earlier benchmarks that differences between them were hard to measure. Results are sometimes reported separately by condition, such as whether external tools are allowed.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
Humanity's Last Exam with tools 67.7% 65.6% 57.2%