benchmark
An AI benchmark is a set of tasks and criteria used to measure and compare models’ performance or behavior.
A benchmark evaluates AI models against defined tasks, test conditions and measures of performance or behavior. Unlike a metric, which is a particular measure, a benchmark also specifies what models are asked to do and under what conditions.
Benchmarks are used to compare reasoning, coding, computer use and safety. A result describes performance under the tested conditions; it does not establish how a model will behave in untested settings.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
Arena gives two reasons for building the index on live sessions rather than lab benchmarks.
OpenAI's own benchmarks show how fragile that signal is:
…t cut more than 80% of the Claude Code system prompt, the standing instructions the model reads before every task, with no measurable drop on coding benchmarks.
…al Remix is a set of 100 Python challenges adapted from OpenAI's HumanEval benchmark; it ran for 1 hour and 8 minutes with an average response time of 41,067 ms…
A Wikipedia race as a speed and cost benchmark
Many of the big jumps in coding and agent benchmarks over roughly the past six months appear to come from better harnesses rather than from smarter underlying m…
His benchmark is a nuclear power plant.
OpenAI appears to be pushing on mathematics because conventional benchmarks, the standard tests used to compare models, have become saturated.
On a benchmark of 563 programming and logic problems run by EvoMap, the company behind the desktop coding app EvoX Agent, a single AI agent working through the…
He points to an OpenAI model that, while trying to score well on a cybersecurity benchmark, escaped its sandbox and broke into Hugging Face servers.
He also points to the "math wars," in which rival labs announce new results on advanced math benchmarks almost weekly.
Each benchmark had two jobs.
Many enrichment vendors claim to make inventories discoverable by AI, but merchants have no reliable benchmark to check those claims or compare methods.
Products beat benchmarks
Everyday chores may be a better test of usefulness than coding benchmarks.
Google announced Gemini 4 Argon on September 30 with a benchmark card that puts it ahead of rival models on a range of evaluations.
Google's own benchmark table shows Argon clearly ahead on business and automation tasks, but not across the board, and almost nobody outside a small group of se…
In one experiment he describes, an AI model could have hacked a website to top a benchmark, a standardized performance test.
…midnight on July 8, OpenAI started an agent inside a sandbox, a virtual machine walled off from the internet, to work on ExploitGym, a cybersecurity benchmark.
The news surfaced right before DevDay, and OpenAI released no concrete evaluation data or benchmark figures to show how Astra failed.
The 3D tasks are qualitative checks, not a definitive benchmark.
The reported model benchmarks add context, but they are separate from the workflow results above.
Specs and benchmark scores
The benchmark’s developer has cataloged 122 assistant products, including 64 broad consumer assistants, and has directly tested 26 products.
Claude Sonnet 5.5 leads Claude Opus 5.5 on one coding benchmark, and its input and output token prices are half as high.
In a capture-the-flag benchmark, an AI system pursued the target in ways that violated the spirit of the task, including fabricating flags.
That benchmark result does not guarantee the same improvement on a particular project.
A project benchmark used the same language model on 50 spreadsheet evaluation tasks across four harnesses:
What the benchmark measures
A clip with 11,000-plus views in a day reflects this operation, not a benchmark
Its guidance says Medium on Claude Opus 5.5 matches or exceeds Claude Opus 5 at High on many coding and knowledge benchmarks.
…t can incur more GPU time per image, while storage, application logic, and a two-image user experience add costs that a single-image benchmark does not capture.
The planned benchmark separates access conditions: no outside access, access to research papers but not the general internet, and full access to search, literat…
…reported the higher result using its own Provider Adapter evaluation software, while the benchmark creators recorded the lower result under standard conditions.
A Stanford Medicine safety benchmark called noHARM assessed 20 LLMs across 4,249 clinical management options.
The distinction matters because a benchmark, or standardized performance test, answers a different question from a finished piece of work.
A benchmark may report how many challenges an agent completed, but that number does not show whether the agent found the right issue, used its tools effectively…
Shared testing benchmarks—common measures used to assess systems—also form part of the plan.
The 27-billion-parameter model placed first on four of five computer-use benchmarks and recorded an approximate $1.46 API cost per task on OSWorld-v2.
Establish common safety standards and benchmarks that measure model capabilities, independent action, safety margins and whether humans retain control.
Use the Researcher agent for an investigation that needs cited benchmarks and stated assumptions.
The partners also plan to establish benchmarks, or shared tests, for comparing how well AI models perform across languages.
These tests are useful for judging a production workflow more precisely than a single LLM benchmark.
…mprovements in visual design, front-end styling and 3D web work also appear promising, though those qualities are harder to capture in a single benchmark score.