Humanity's Last Exam

Humanity's Last Exam is a benchmark of difficult expert-level questions used to assess AI models’ knowledge and reasoning.

2 articles
Last mentioned

Humanity's Last Exam (HLE) is a benchmark of difficult questions across specialist fields that tests AI models’ knowledge and reasoning. It was developed by the Center for AI Safety and Scale AI with subject-matter experts and released in 2025.

It was designed with hard questions because leading models had begun to score so highly on earlier benchmarks that differences between them were hard to measure. Results are sometimes reported separately by condition, such as whether external tools are allowed.

This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.

Articles covering this entry

Haiku 5.5 also posted 45.9% on Humanity's Last Exam without tools and 57.4% with tools, 39.2% on the agentic coding test Terminal-Bench 4.0, and 46.4% on both F…

Humanity's Last Exam with tools 67.7% 65.6% 57.2%


© 2026 AIPOST. All rights reserved.

AIPOST is an AI publication covering practical AI, AI security, performance, startups, health, ethics and industry news. No account is needed, and our privacy policy explains how we handle personal information.