Google announced Gemini 4 Argon on September 30 with a benchmark card that puts it ahead of rival models on a range of evaluations. For now, though, only designated cyber defenders can use it. In the same stretch, OpenAI built Dots, a personal agent platform, into ChatGPT, and Anthropic released Claude Sonnet 5.5 at exactly Argon’s price. An agent here means an AI system that carries out multistep tasks on a user’s behalf. Taken together, these launches suggest the contest is shifting away from benchmark scores and toward two questions: how reliably a model’s agents can use tools, and which app people actually work in every day. Google appears to need good answers to both if Argon is going to win back developers.

A model few people can try yet

On Google’s own benchmark card, Gemini 4 Argon beats its competitors across multiple evaluation criteria. Because access is limited to a small group of security defenders, however, outsiders have no practical way to check whether those numbers hold up in real work. When a frontier model opens to a handful of users first, everyone else is better off waiting for broad hands-on experience before trusting the launch figures.

There is history behind the caution. Gemini 3 and Gemini 3.1 also impressed on benchmark charts, then disappointed users in everyday work. Judging Argon by sustained use after it opens up looks like the safer approach.

The pricing is competitive. Argon’s introductory rate is $2 per million input tokens and $10 per million output tokens. Tokens are the small units of text a model reads and writes. That is the same price as Claude Sonnet 5.5.

Model Input, per million tokens Output, per million tokens
Gemini 4 Argon $2 $10
Claude Sonnet 5.5 $2 $10
Claude Opus 5.5 $4 $20

Claude Sonnet 5.5 costs half as much as Claude Opus 5.5, yet it scored 70.6% on Terminal-Bench 4.0, an agentic coding benchmark, ahead of Opus 5.5 at 64.4%. If Argon truly performs at the frontier, its price is reasonable, but it also enters a price tier where a strong rival is already waiting.

The homework: tool calling

The last time Gemini offered a genuinely competitive model was arguably Gemini 3, in January 2026. Gemini 3 was good at generating front-end code, the part of an app that users see, but weak at tool calling, a model’s ability to invoke outside tools to get work done. Without dependable tool calling, it is hard to build agents that act on their own, and that weakness held Gemini back. Google’s AI coding tool Antigravity never caught on widely, apparently because developers and founders did not trust the underlying Gemini model for their daily work.

Gemini does have a clear strength. It accepts video directly as input and can take in and deeply analyze a full video in roughly 10 seconds, something competing frontier models do not match. That strength has not been enough on its own. Through 2026, Google appears not to have shipped a model that founders and advanced knowledge workers rely on every day.

A stylized AI robot connecting cables to tool icons such as a folder, an envelope and a terminal, some cables loose

▲ Agent tool calling

The app where agents live

Competitors are turning their own apps into the place where agents do their work. OpenAI’s Dots sits right inside the ChatGPT desktop app. A user can call Dots by voice from the ChatGPT phone app and hand off a task. Dots can then update a teammate over Slack and email and, when needed, open a session in Codex, OpenAI’s coding tool, to build web page mockups. The fast model that runs Dots does not use plan credits until it hands work off to Codex or GPT models. Alongside Dots, OpenAI introduced ChatGPT Space, where people and agents edit the same documents together.

Anthropic, for its part, lets a main agent in Claude’s Projects open a separate thread for each task and handle several jobs at once, and those threads can run in the cloud or on the user’s Mac. Beyond ChatGPT and Claude, GrokBot and Cursor also offer dedicated desktop apps.

Google’s AI features, by contrast, are spread across Antigravity, the Gemini app, Google AI Studio and NotebookLM, with no clear signal about where a new model should be used. OpenAI took the opposite path, putting the ChatGPT desktop app first so users have a definite place to go whenever a new model arrives. Unless Google brings its features together in a single app, even a strong Argon may struggle to give users a reason to stay.

One tidy desktop app holding documents, chat and code beside scattered small windows, with a user choosing between them

▲ A unified app versus scattered tools

What to do now

Gemini 4 Argon is welcome competition, but it deserves cautious watching rather than excitement based on benchmarks alone. A few practical steps:

  1. Judge new models by steady use on your own work, not by the vendor’s benchmark card.
  2. When Argon opens up, start with agent tasks that depend on tool calling to see whether the old Gemini weakness is fixed.
  3. If you need fast video analysis, Gemini’s strength is already usable today.
  4. Choose agent tools by how well they fit the app you work in every day, not by model scores alone.
  5. Prepare a side-by-side cost and quality comparison with Claude Sonnet 5.5, which sits at the same price.