GPT-6 Luna produced usable browser software and two playable C++ 3D games in a set of demanding development tasks. The strongest result was an initial skateboarding game that compiled cleanly and included tricks, rail grinding, and scoring. The weaker results matter just as much: a browser desktop needed a code fix before it would open, and a rally game looked more convincing than its steering felt. Luna appears most useful as a low-cost starting point for software that someone can run, inspect, and improve.

What the low price covers

OpenAI lists GPT-6 Luna as the lightweight model in its GPT-6 family. Its API prices are quoted per million tokens, the units used to measure text processed or generated by a model:

Model Input price per million tokens Output price per million tokens
GPT-6 Luna $0.10 $0.50
GPT-6 Sol $2.00 $10.00
GPT-6 Astra $10.00 $50.00

Those rates describe API use, not the price of a ChatGPT Pro subscription. Luna supports text and image inputs, has a 1,050,000-token context window, and can produce up to 128,000 output tokens. Its listed knowledge cutoff is May 19, 2026. Published DeepSWE results put Luna at 66.6% with maximum reasoning effort, compared with 62.2% for GPT-5.6 Luna. That benchmark result does not guarantee the same improvement on a particular project.

A browser desktop that needed one repair

With maximum reasoning effort selected, GPT-6 Luna was asked to make a single HTML file containing a browser-based desktop, several everyday apps, animated wallpaper, and two functional 3D mini-games. Its reported generation time was 45 minutes and 13 seconds, but roughly 40 minutes of that interval involved waiting for approval at an interactive permission prompt.

The first file did not render: a JavaScript syntax error stopped the page. After receiving the browser console error, Luna supplied a patch in 3 minutes and 45 seconds. The repaired desktop opened with a searchable app launcher, a local clock, wallpaper controls, mail, notes, and a calculator. Notes saved text in the browser, while the calculator performed arithmetic despite its awkward, oversized layout.

Blank browser mail, notes, and calculator panels beside a low-poly driving game and abstract wallpaper

▲ Browser desktop with built-in apps and games

Both games ran. One let the player drive through a low-poly city and collect packages; the other was a third-person obstacle-dodging game. A focus timer also changed the wallpaper during timed work sessions. The result met the central goal of an interactive, multi-app file, though some window controls were clumsy and the driving game lacked an in-window exit control. The practical lesson is to treat a generated app that fails on launch as a debugging task, not as a finished deliverable: inspect the console error, pass it back, and test the repaired file.

Two C++ games, two different strengths

The clearest native-code success was a single-file skateboarding game set on a city block. Luna was asked to avoid Raylib, a game-development library, and instead supplied C++ code with build instructions using OpenGL, GLU, and GLUT for graphics and windowing. The code compiled without errors. In 9 minutes and 50 seconds of generation time, the task yielded a playable scene with ramps, rails, an animated skater, ollies, flips, spins, automatic grinding, and combo scoring. Luna gave the user a build command rather than running the compilation itself, so the successful build still required a separate check.

A second native project, Dustline Rally, took 23 minutes and 8 seconds. It used SDL2 and OpenGL rather than Raylib and opened with a menu, a track diagram, and a playable course. Players could switch between a chase camera and a cockpit view while driving past trees and elevation changes. Yet the car snapped to the road path instead of responding with realistic steering and tire behavior. Some gauges did not move, and parts of the track appeared to float. It is a functioning 3D driving prototype, not a finished driving simulation.

A later visual overhaul of the skate game added rain, audio, neon lighting, and moving taxis. That iteration used GPT-5 Luna Max, so its additions should not be credited to GPT-6 Luna’s initial build. The first build already establishes Luna’s more relevant strength: generating a substantial game foundation that compiles and plays.

A 3D website reveals the cost of visual review

A separate website task in the broader set of creative builds produced a single-page watch showcase using Three.js, a JavaScript library for 3D graphics. The page offered selectable procedural watch models and an interactive slider that separated a watch into its component layers. Its full build took 38 minutes and 57 seconds.

Separated layers of a generic three-dimensional watch above a dark pedestal, including face, hands, and lens

▲ Exploded view of a 3D watch model

The interaction worked, but the default watch face had misplaced numeral markers. Its hands initially appeared to be missing; moving the exploded-view slider showed that they existed beneath other geometry. That distinction matters during review: a missing-looking part may be present but hidden by the order or position of 3D surfaces. The page provided a functional starting point, not a polished product page. Unlike the initial browser desktop and native-game tasks, the final website build is not as clearly attributed to GPT-6 Luna alone, so it should not serve as a precise Luna performance measure.

Where Luna fits

After the broader set of coding and 3D tasks, a $200-per-month ChatGPT Pro account still displayed 100% of its weekly allowance. Less than 1% of that allowance had been consumed. This is one account’s subscription indicator, not an API bill or a promise that another workload will use the same share.

For a developer choosing a tool, the evidence favors GPT-6 Luna when the goal is to get interactive code running cheaply and there is time to verify it. Run generated browser files, compile native projects, and check controls, physics, and 3D layer visibility before building further. The low API price makes repeated prototyping appealing; the syntax error, rally handling, and uneven visual details show why testing and refinement remain part of the job.