Terminal-Bench
Terminal-Bench is a benchmark of how well AI agents complete tasks by working in a terminal.
Terminal-Bench is a benchmark released in 2025 by Stanford University and the Laude Institute that evaluates AI agents on tasks they carry out by issuing commands in a terminal. Unlike SWE-bench, which focuses on resolving issues in software repositories, it assesses task completion through terminal-based work more broadly. A related benchmark, Terminal-Bench Science, focuses on scientific research tasks.
Each task comes with an isolated environment and tests that check the outcome, and later revised versions followed. AI companies frequently cite it when reporting the autonomous-task performance of their models and coding agents.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
Claude Sonnet 5.5 costs half as much as Claude Opus 5.5, yet it scored 70.6% on Terminal-Bench 4.0, an agentic coding benchmark, ahead of Opus 5.5 at 64.4%.
Terminal-Bench 4.0 57.4% 58.2% 57.9% 66.4%
Claude Sonnet 5.5 reaches 70.6% at maximum effort on Terminal-Bench 4.0, compared with a reported 54.4% for Claude Opus 5.5.
Terminal-Bench 4.0 66.4% 55.8% 57.9%