FrontierCode

FrontierCode is a benchmark from Cognition that evaluates whether code changes made by AI coding agents are good enough to merge.

3 articles
Last mentioned

FrontierCode is a benchmark that evaluates whether patches produced by AI coding agents for real issues in open-source repositories are good enough for the maintainers to merge. It was introduced in 2026 by Cognition, the developer of the AI coding tool Devin. Beyond whether tests pass, it grades correctness, test quality, scope, style and adherence to project conventions against rubrics written by the repositories' maintainers. Version 1.1 reports results on two task sets: Main, with the 100 hardest tasks, and Extended, with all 150.

This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.

Articles covering this entry

…tools and 57.4% with tools, 39.2% on the agentic coding test Terminal-Bench 4.0, and 46.4% on both FrontierCode 1.1 and the visual reasoning test Chartography.

FrontierCode v1.1 49.4% at high effort; 46.2% at max 54.4%

FrontierCode v1.1 54.4% 50.3% 53.3%


© 2026 AIPOST. All rights reserved.

AIPOST is an AI publication covering practical AI, AI security, performance, startups, health, ethics and industry news. No account is needed, and our privacy policy explains how we handle personal information.