OSWorld

OSWorld is a benchmark that measures how well AI agents complete tasks by operating applications in real computer environments.

3 articles
Last mentioned

OSWorld is a benchmark released in 2024 by researchers led by the University of Hong Kong. It evaluates AI agents that observe a computer screen and operate applications to complete tasks in real operating systems such as Ubuntu, Windows and macOS. Unlike static question-answering or web-only benchmarks, it checks the outcome of desktop tasks that can span multiple applications. Later editions include OSWorld-Verified, which reviewed and corrected the tasks and evaluation environment, and OSWorld-v2.

This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.

Articles covering this entry

OSWorld 2.1, computer use 72.4% 83.9% 15.7%

In AutomationBench and OSWorld testing of desktop tasks, it reached 72.6% accuracy and took an average of 40 minutes per task, compared with 75 minutes for prev…

The 27-billion-parameter model placed first on four of five computer-use benchmarks and recorded an approximate $1.46 API cost per task on OSWorld-v2.


© 2026 AIPOST. All rights reserved.

AIPOST is an AI publication covering practical AI, AI security, performance, startups, health, ethics and industry news. No account is needed, and our privacy policy explains how we handle personal information.