Google’s Gemini 4 Argon, announced on September 30, is pitched as frontier-level intelligence for enterprise knowledge work, software engineering and cybersecurity defense. The announcement came from Koray Kavukcuoglu, senior vice president at Google DeepMind and chief AI architect at Google. Google’s own benchmark table shows Argon clearly ahead on business and automation tasks, but not across the board, and almost nobody outside a small group of security defenders can use it yet. That means most of the numbers below are self-reported. Here is where Argon leads, where it trails, what it costs and what Google says it has already done with it internally.

The scorecard: strong on agent work, mixed on coding

Google compared Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5.

Benchmark Gemini 4 Argon GPT-6 Astra Claude Fable 5.1 Claude Opus 5.5
Vals Index (knowledge work) 68.9% 63.1% 65.8% 67.0%
AutomationBench 51.3% 41.4% 31.4% 42.5%
Vals Finance Agent v2 65.4% 53.5% 58.9% 58.6%
Harvey Legal Agent Benchmark 19.6% 5.4% 6.7% 3.8%
DeepSWE v1.1 77.9% 74.1% 67.4% 74.2%
Vibe Code Bench 91.9% 89.6% 90.3% 90.3%
FrontierSWE v2 55.0% 65.5% 56.3% 62.3%
Terminal-Bench 4.0 57.4% 58.2% 57.9% 66.4%

The widest gaps show up where an agent, an AI system that carries out multistep tasks on its own, does real business work: knowledge work, workflow automation, financial analysis and legal tasks. On the legal agent benchmark, Argon scored 19.6% while its rivals sat between 3.8% and 6.7%. Coding is a split decision. Argon leads on DeepSWE v1.1 and Vibe Code Bench, but it trails GPT-6 Astra on FrontierSWE v2 and falls well behind Claude Opus 5.5 on Terminal-Bench 4.0, which tests work in the command line. Parity at the top, rather than a clean sweep, is the fairer reading.

Independent results point the same way. In evaluations by Artificial Analysis, Argon running in high reasoning mode took first place on AutomationBench-AA, a test of agentic SaaS workflows, with 78%, and scored 52% on Terminal-Bench 4.0, close to the top. The gains look like broad improvement rather than a model tuned to game particular benchmarks.

A 1-million-token output limit and introductory pricing

Argon can generate up to 1 million tokens in a single response, sixteen times the previous ceiling of 64,000 tokens. Tokens are the small units of text a model reads and writes. The output limit is separate from how much input a model can read at once and from any memory that carries across sessions; it matters for long chains of reasoning and for producing large, multi-file code changes in one pass.

Introductory API pricing is $2 per million input tokens and $10 per million output tokens, with a 95% discount on cached input. Because this is launch pricing, it may not hold for long.

A stylized AI core producing one short paper strip and one extremely long strip side by side

▲ A much larger output limit

What Google says Argon has already done

Thousands of Google engineers already use Argon across engineering and research. Google’s examples include:

  • Quantum computing: optimizing the qubit and gate resources of a quantum algorithm, Argon beat a published baseline by 40% in 40 minutes.
  • Data center memory: Argon agents analyzed fleet-wide profiling data and freed more than 300 TiB of memory, with eventual savings projected at 500 TiB to 1 PiB. At Google’s scale, even tiny optimizations add up to large savings.
  • Rust migrations: agents are moving C and C++ code to memory-safe Rust, including more than 800,000 lines for the Zircon kernel of Google’s Fuchsia operating system.
  • Video decoding: in libgav1, Google’s open-source video decoder, agents reviewed an existing Rust port and replaced 32,000 lines of SIMD code. The result produces identical video output and runs 2.7 times faster than the previous Rust port.

The libgav1 work shows how these agents operate. They ran repeated rounds of profile-guided experiments, examined compiler output and reshaped the Rust code so the compiler could vectorize it automatically, meaning it could turn loops into instructions that process many values at once. Forming a hypothesis, testing it against a measurable target and keeping what works, around the clock, resembles the automated research loop Andrej Karpathy has described. These are Google’s own showcase results, though, and they may be the best outcomes picked from many attempts.

Stylized robots refining code blocks in a loop beside a rising speed gauge and server racks

▲ Iterative, experiment-driven code optimization

Why you probably cannot try it yet

Argon is rolling out in phases. For now, access is limited to trusted cyber defenders in Google’s Fairwind Program, in coordination with US government protocols for pre-release model access. Paid API subscribers on the Ultra plan are next in line, while individual developers and general users have to wait.

That also means independent developers cannot yet check Google’s numbers for themselves. Early results look promising, but a firm verdict on how much better Gemini 4 really is will have to wait for broad, independent hands-on testing.

What to do while you wait

Argon puts Google back in the top tier, with clear strengths in automation and knowledge work and a closer race in coding. A few steps make sense before access opens up:

  1. Shortlist tasks where Argon’s lead is largest, such as workflow automation, financial analysis and legal document work, for your first trials.
  2. If you plan to use it for command-line coding agents, prepare a side-by-side test against Claude Opus 5.5 and GPT-6 Astra on the same tasks.
  3. Treat the launch benchmarks as Google’s own results and re-check them on your real workloads once you get access.
  4. Budget with some headroom, since $2 per million input tokens is introductory pricing.
  5. Identify jobs that need very long outputs, such as large code generation, where the 1-million-token limit could change what is possible in one pass.