Terminal-Bench

Terminal-Bench is a benchmark of how well AI agents complete tasks by working in a terminal.

5 articles
Last mentioned

Terminal-Bench is a benchmark released in 2025 by Stanford University and the Laude Institute that evaluates AI agents on tasks they carry out by issuing commands in a terminal. Unlike SWE-bench, which focuses on resolving issues in software repositories, it assesses task completion through terminal-based work more broadly. A related benchmark, Terminal-Bench Science, focuses on scientific research tasks.

Each task comes with an isolated environment and tests that check the outcome, and later revised versions followed. AI companies frequently cite it when reporting the autonomous-task performance of their models and coding agents.

This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.

Articles covering this entry

…m without tools and 57.4% with tools, 39.2% on the agentic coding test Terminal-Bench 4.0, and 46.4% on both FrontierCode 1.1 and the visual reasoning test Char…

Claude Sonnet 5.5 costs half as much as Claude Opus 5.5, yet it scored 70.6% on Terminal-Bench 4.0, an agentic coding benchmark, ahead of Opus 5.5 at 64.4%.

Terminal-Bench 4.0 57.4% 58.2% 57.9% 66.4%

Claude Sonnet 5.5 reaches 70.6% at maximum effort on Terminal-Bench 4.0, compared with a reported 54.4% for Claude Opus 5.5.

Terminal-Bench 4.0 66.4% 55.8% 57.9%


© 2026 AIPOST. All rights reserved.

AIPOST is an AI publication covering practical AI, AI security, performance, startups, health, ethics and industry news. No account is needed, and our privacy policy explains how we handle personal information.