Google DeepMind has rebuilt the foundation of the Gemini API for agents. The new Interactions API handles quick model calls and long-running agent jobs through a single endpoint, and it can keep conversation and reasoning state on Google’s servers so developers no longer stitch it together by hand. On top of that, Managed Agents give Gemini its own Linux sandbox, an isolated computer for running code, with one API call. For teams building agents on Gemini, the update appears to move three chores to the platform: state management, execution environments and secret handling.
Why the old API fell behind
The way developers use language models has gone through three stages. First came simple completions: send a prompt such as “tell me a joke” and get a block of text back. Then applications needed output they could rely on, which led to function calling, where the model returns a structured JSON object naming a function and its arguments. Ask for the weather in New York City and the model replies with a call to a weather function. A back-end server can act on that payload the same way it handles a sign-up form with a username and password.
Today’s models go further. They take a goal, reason through it, use several tools and may work for a long stretch before answering. Google DeepMind describes an agent as the sum of four parts:
| Part | Role |
|---|---|
| Model reasoning | Interprets the user’s intent and decides the next move |
| Tools | Search, code execution and custom functions |
| Context and memory | Short-term and long-term recall |
| Execution environment | The sandbox where code and functions run |
A developer experience engineer at Google DeepMind argues that strong models need less fragile scaffolding than before. Rather than separate tools for reading and editing files, models such as Claude Opus and Claude Haiku often work through a single bash tool. The harder problem is time. Jobs that once finished in one turn now run in the background, like deep research agents that work for three minutes or more before producing a long report. The older API made developers dig through deeply nested JSON to find results and track conversation history on their own.
One ID carries the conversation and the reasoning
The core of the Interactions API is optional server-side state. Each response comes with an interaction ID. Pass it as previous_interaction_id on the next call, and the server restores the earlier conversation along with the model’s internal reasoning state.
That matters because of thought signatures. Recent Gemini models record their reasoning as opaque encoded values that must be returned on the next turn for the model to perform at its best. According to Google DeepMind, dropping them leads to a noticeable decline in reasoning quality. Rebuilding history by hand is fragile, too: a stray whitespace change can invalidate the server cache. Sending one ID removes that whole class of bugs.

▲ State carried by an interaction ID
Output is cleaner as well. Text, audio and image responses come back as strongly typed objects, so code can branch on output.type instead of hard-coding paths through nested dictionaries. Switching modalities takes only a change to settings such as generation_config and response_modalities. One example from a Google DeepMind team member is a travel simulator: upload a single portrait, and it generates images of that person at places like the Venice canals and Neuschwanstein Castle in parallel, then chains the interaction ID to turn the first frame into an animated clip.
Tool calls and the Steps Data Model
Within a single API call, Gemini can decide what it needs to know, call tools and return an answer. Google Search and a URL context tool bring in web content from after the model’s training cutoff. The same request can include a custom function, such as a file_incident tool that logs a security issue in an internal ticketing system, so the agent can research and act in one pass.
To record these multi-step runs, Google replaced the old outputs array with the Steps Data Model, a strongly typed timeline of everything the model did during an interaction.
| Element | What it holds |
|---|---|
| Step type | Model output, thought, function call, Google Search call and more |
| Status | Done, waiting or in progress |
| Structured fields | Tool names, arguments and thought signatures |
| Content | Raw text, image, audio and video |
Managed Agents put a sandbox behind the API
Managed Agents, which run Antigravity as a remote agent, let Gemini work on its own inside a secure Linux sandbox that Google manages. The default toolset covers bash code execution, file system access, web search, URL context and agent skills. Building a coding agent, an agent that writes and runs code, used to mean tuning the harness yourself, finding and configuring third-party sandbox infrastructure and working out how to keep context. Now a persistent sandbox comes with one API call.
In one example, the agent received a GitHub repository, cloned it, listed its directories, read the source files one by one and wrote a project overview without human help. In another, it evaluated a hackathon entry that built its own programming language and reinforcement learning loop over several turns, using more than 2 million tokens, the units of text a model processes, without losing sandbox state.
Multi-turn work depends on keeping two values:
| Value | What it preserves |
|---|---|
| Interaction ID | Conversation and reasoning context |
| environment_id | Files and runtime state in the existing sandbox |
Sandboxes can be preloaded with Google Cloud Storage buckets, GitHub repositories and inline files such as configuration values or setup scripts.

▲ Sandbox and network proxy
A network proxy keeps secrets away from the model
Developers often ask how to protect agents from leaked credentials and prompt injection, an attack that hides instructions in content the model reads. Google’s answer is a managed man-in-the-middle network proxy. When code in the sandbox sends a request to an outside service such as the GitHub API, the proxy intercepts it and adds the authorization token to the header on the way out.
The model never sees or holds the raw secret. Even if an attacker uses prompt injection to make the model print its memory or environment, the API key cannot leak because it was never on the model’s side. Compared with putting keys in prompts or in environment variables the model can read, this appears to move the line of defense outside the model entirely.
Named agents, a CLI and a skills repository
A tuned setup of instructions, files and dependencies can be frozen into a custom named managed agent in two ways:
- From source: write a base agent spec that lists the skills and configuration files it needs.
- From a snapshot: chat with an agent in a sandbox, install the libraries you need, then save and name that environment.
Each project can have up to 1,000 named agents. Google does not charge for environment storage or idle sandboxes; billing covers model token usage only.
From the terminal, the open-source Gemini API CLI can send prompts to any model and create, test and deploy agents. For moving existing coding agents to the new API, Google published the gemini-skills repository on GitHub. Models tend to fall back to code for older versions they saw often in training, such as Gemini 2.0 or Gemini 2.5 Flash, so a skill that describes the new API steers them toward current patterns. Google says it keeps evaluating these skills to make sure the generated code follows current best practices.
What Gemini developers should change now
The Interactions API and Managed Agents look like an effort to take state, infrastructure and secret handling off developers’ plates. If you build agents on Gemini, these steps are worth checking:
- Stop rebuilding multi-turn history by hand. Pass each response’s interaction ID as
previous_interaction_id, which is the simplest way to keep thought signatures intact. - Handle multimodal responses by checking
output.typerather than hard-coding nested paths. - With Managed Agents, store both the interaction ID and
environment_idso you keep the context and the sandbox. - Inject API keys through the network proxy instead of prompts or environment variables.
- Add the Interactions API skill from the gemini-skills repository to your coding agents, and test with the Gemini API CLI before freezing a setup into a named agent.
For a quick start with no setup, Google AI Studio offers a free tier for trying the models.