GDPVal-AA
GDPVal-AA is an Artificial Analysis benchmark, built on OpenAI's GDPval, that ranks AI models on real-world knowledge work using Elo ratings.
GDPVal-AA is a benchmark run by the AI evaluation firm Artificial Analysis. Built on tasks from GDPval, which OpenAI released in 2025, it has models carry out real-world knowledge-work tasks, compares their outputs against one another and ranks them with Elo ratings.
GDPval draws its tasks from actual work in occupations that account for a large share of the US economy, aiming to measure how well AI models perform economically valuable work. GDPVal-AA reports relative head-to-head results as Elo ratings rather than accuracy scores.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
Professional-work rating, GDPval-AA 1,542 1,672
On GDPval-AA, a knowledge-work evaluation spanning 44 occupations, Claude Sonnet 5.5 scores 1844, close to Claude Opus 5.5's 1846 and above Claude Sonnet 5's 14…
An Elo score is a relative rating used to compare performance; a higher GDPVal-AA rating indicates a stronger result in that evaluation.