GPT-6.1 Sol, which OpenAI introduced at DevDay, produced a working browser desktop with connected apps and two playable 3D games in hands-on tests, but its results varied across visual design and physical control tasks. Its first-pass 3D visuals looked less polished than earlier Claude Sonnet 5.5 work on the same tasks, and a visual reference narrowed that gap.

Specs and benchmark scores

OpenAI positions GPT-6.1 Sol as a lower-cost model with intelligence close to GPT-6 Astra. Its listed API prices are $2 per million input tokens and $10 per million output tokens, matching the prices given for Claude Sonnet 5.5. Input and output tokens are the units used to count material sent to and generated by a model. Price parity does not, by itself, establish which model builds a better app.

The DeepSWE v1.1 software-engineering benchmark provides another reference point. Its reported GPT-6.1 Sol results differ by reasoning effort, a setting that changes how much work the model devotes to a task:

Reasoning effort DeepSWE v1.1 score Reported cost per task
High 75.2% $0.65
Maximum 71.9% $0.97

On that benchmark, GPT-6.1 Sol matches GPT-6 Astra’s baseline and scores 6.4 percentage points above GPT-6 Sol. The higher setting scoring lower reflects a known effect: a model can overcomplicate a task beyond what the evaluation expects. The API model listing also gives GPT-6.1 Sol a 1,050,000-token context window, a 128,000-token output limit and an April 30, 2026 knowledge cutoff; it accepts text and images and returns text.

Those scores measure benchmark performance, not the finish of a game or the success of a robotic grip. The app runs offer a closer look at those outcomes.

A browser desktop with working connections between apps

The clearest software result was Halo OS, generated in 33 minutes and 42 seconds at maximum reasoning effort. It is a self-contained browser desktop in one HTML file, the format browsers read to display pages. It opened in Google Chrome with windows, a dock, a launchpad, notes, mail, a synthesizer and tools for changing the wallpaper and saving sessions. Some dock and sidebar icons failed to render, but the desktop and its main apps worked.

Two games ran inside its windows. Signal City offered low-poly streets, drivable cars, package deliveries and pedestrians. Its vehicle handling and camera were responsive, though police cars sometimes appeared immediately beside or on top of the player’s car. Orbital Run let a player steer a spacecraft through rings and obstacles. Finishing its course awarded 4,216 points and added a certificate to the desktop’s mail app.

Close view of a browser desktop with a low-poly city driving game, ring-flight game, and shared app panels

▲ Two games inside a browser desktop

The connections extended beyond the games. Mail reflected activity from both titles; Notes could export text files; and Continuum saved and restored open windows, game states and notes. That shared state matters more to the assessment than the number of apps alone: it shows the generated desktop could coordinate separate features. The visual judgment was less favorable. Claude Sonnet 5.5’s work from the previous day appeared more polished, while GPT-6.1 Sol generated its browser system noticeably faster and ran it responsively.

Playable games, with uneven visual polish

The pool-game task asked GPT-6.1 Sol to use Blender, a 3D modeling tool, and Godot, a game engine, to build four playable divers, water effects, sound and scoring. An initial run used Fast Mode and took 18 minutes and 12 seconds. Because that setting may have affected output quality, the project was rebuilt after clearing earlier files and disabling Fast Mode.

The clean run took about 51 minutes. Its backyard scene, character selection, charged jumps, airborne tricks and scored splashes worked. Yet the comparable Claude Sonnet 5.5 pool game had richer particles, more expressive animation and stronger visual humor. This is evidence of a visual-quality gap on that particular task.

A skateboarding project showed how much the instructions could change the result. GPT-6.1 Sol first delivered a playable C++ game with tricks and the requested slow-motion replay, but the street environment looked sparse. After receiving a screenshot as a visual target and permission to use more than one source file, it produced a fuller waterfront setting with reflections, scenery and sound. Pedestrian collisions still were not mapped. The revision suggests that a functional first build need not be treated as the model’s final visual result.

Other runs reinforced that distinction between feature coverage and completeness. A browser subway game linked enemy waves to train rides between stations and included working weapons and later enemy variants, though a train’s departure looked unrealistic. An Old School RuneScape combat recreation assembled detailed interface panels and playable fights in 19 minutes and 30 seconds, but special attacks did not respond as expected. These are substantial working builds with specific gaps, not uniformly finished games.

The robot arm stopped without completing its task

A physical test gave GPT-6.1 Sol control of a robotic arm above a grid mat. Its assignment was to move a toy vehicle. During roughly 18 minutes of operation, the arm adjusted its trajectory when the vehicle was moved or turned, showing that its camera feedback and motion planning responded to changes in the workspace.

Articulated robot arm parked upright above a gridded mat, with a toy vehicle outside the gripper’s reach

▲ Robotic arm parked clear of the toy vehicle

The arm never obtained a stable grip and did not transfer the vehicle. It then halted upright, clear of the toy. The outcome is a task failure with a useful safety distinction: responsive tracking did not translate into successful manipulation, but the system stopped rather than continuing uncertain movements. For a physical task, both completion and behavior after failure belong in the evaluation.

What to take from the results

Across the reported tests, GPT-6.1 Sol looks capable of building interactive software with connected features, especially when a task rewards logic and multi-step behavior. Its first-pass 3D visuals were less consistently impressive, and the robot could not complete its transfer. A ChatGPT Pro account’s weekly allowance fell from 98% remaining to 95% during the collection of tests; that is one account’s usage figure, not a forecast for other workloads.

If choosing a model for an app project, compare the same prompt and constraints on the models of interest. Test whether features actually work, inspect visual quality separately, and check how each build handles errors or failed actions. Use visual references when appearance is important. These results make GPT-6.1 Sol a solid, cost-effective contender worth testing for such work, including against top-tier models such as Claude Opus.