Google Antigravity coordinated 93 AI coding agents over 12 hours to build an operating system kernel that ran DOOM. The run used Gemini 3.5 Flash, processed 2.6 billion tokens and cost under $1,000 in API credits. It is a concrete example of a large, parallel coding task reaching a visible outcome. It is not, on its own, evidence that the same task will succeed reliably on another run.

What the kernel run established

A kernel is the core software that lets programs run on a computer. In this case, the goal was to create one from scratch and use it to run DOOM, without human code authoring. The resulting kernel was demonstrated running the game at Google I/O.

Measure Reported result
Coding agents 93 subagents
Elapsed time 12 hours
Model requests More than 15,000
Text processed 2.6 billion tokens
API credits Under $1,000

Tokens are units of text a model reads or generates. Together, the figures describe the scale of this particular run, not a price or completion-time estimate for future projects. The spending figure covers API credits; no full accounting of other project costs accompanies it.

Many isolated coding workspaces send abstract streams toward one operating system core beside a subtle clock

▲ Parallel work on a shared kernel

The outcome matters because the task required more than producing a single code file. Multiple agents had to work toward a shared result over many hours. The demonstrated ability to run DOOM is stronger evidence of progress than a count of generated files would be, but it tests only part of what someone might want from a kernel.

How Antigravity coordinates the work

Google Antigravity offers an Agent Teams mode through the /team command. An agent is an AI system that can take actions toward a task, rather than only return a text response. A lead agent assesses the request, can ask for missing constraints and assigns work to specialized subagents. Those agents can operate concurrently in separate worktrees—independent code workspaces—and isolated execution environments. The lead agent can also assemble a custom interface for tracking their progress.

This is the architecture Antigravity offers for large tasks; the reported kernel figures do not specify every role, assignment or coordination decision made during that run. They do establish that the kernel project used 93 subagents and Gemini 3.5 Flash. Antigravity 2.0 separates its IDE, the code-editing environment, from a standalone Agent Manager designed to oversee concurrent work.

Parallelism helps explain how a project could issue more than 15,000 model requests within 12 hours. It does not mean every request advanced the kernel, or that adding more agents would always improve a result. Task division, collaboration and error handling remain areas where multi-agent systems have room to improve.

A successful demonstration is not a reproduction test

The clearest evidence of completion is the kernel running DOOM in a live demonstration at Google I/O. That supports a narrow conclusion: one resulting system ran the game in the demonstrated setting. The resource totals describe what that run consumed.

A computer runs an abstract pixel-style corridor beside a checklist of empty geometric shapes

▲ Demonstrated result and open verification questions

Judging repeatability would take a build recipe, environment specifications, a test matrix and repeat-run results, which a single demonstration does not provide. Without those details, readers cannot assess from this case how often the same process would succeed or whether another team could recreate the outcome under matching conditions. Running one game also does not establish broader kernel quality, such as behavior across other programs or sustained use. These are limits on what the demonstration shows, not findings that the kernel failed those tests.

Google DeepMind also uses Antigravity for a different kind of parallel work: comparing model evaluation results. In that workflow, a research agent proposes roughly 100 explanations for differences between results, and subagents inspect them concurrently. That example illustrates why the system emphasizes delegation, but it is separate from the kernel result and does not validate its reproducibility.

What to take forward

The DOOM run shows that a coordinated group of coding agents produced a working, demonstrated outcome within the reported time and API-credit budget. Antigravity’s team treats it as proof that long-running agent work is economically viable today, while acknowledging that spending hundreds or thousands of dollars to generate a kernel every day is not yet practical. The run also leaves important questions about repeatability and overall software quality unanswered. When evaluating a similar agent-built project, check the demonstrated behavior separately from its resource figures, then look for build details, tests and repeated runs before treating one success as a dependable method.