The riskiest moment for an AI agent in production may not be when it reasons badly. It may be when it acts on its own reasoning directly. If an agent decides a cluster is unhealthy and then improvises its own restart procedure, nothing guarantees that procedure is correct. The engineer who originally built the Conductor orchestration engine at Netflix argues for a clean split: the model decides what should happen next, and a deterministic execution layer, called the harness, decides how it happens, using only operations that were registered and tested in advance. The model is the brain. The harness is the hands.

Most production agents are not chatbots

The familiar picture of an agent is a request-response loop. A user sends a prompt, a bot calls a tool, and a response comes back. A customer service chatbot is the classic example. That picture is far too small for real enterprise software, because most production agents run with no human watching or prompting them. Teams that think of agents only as chatbots may miss the most valuable background and event-driven automation.

Agent type What it does
Background worker Runs indefinitely and handles processing off the main thread
Scheduled agent Wakes at set intervals, such as an hourly calendar check for new conflicts, then sleeps
Event-driven agent Listens to alerts, logs or webhooks, reacts and then stops
Monitor or reviewer Continuously evaluates system behavior and compliance, and escalates anomalies
Long-running coordinator Coordinates multi-step work over hours or days
Multi-agent system Several specialized agents work toward one business outcome

The less a human is present, the more the question of what controls execution appears to become a safety question.

The agent is a component, the harness is the application

The design mirrors microservices. Companies do not ship one service that tries to handle all application logic. They run many focused services under an orchestration layer that manages state, routing, policy and failure recovery. Agents are following the same path. A single agent loop is a component, not an application. The application is the harness that wraps agents together with humans, databases, APIs and MCP servers. MCP, the Model Context Protocol, is a standard way to connect AI models to outside tools.

Splitting the work also limits hallucination, the tendency of a model to produce plausible but wrong output. Asking one agent to handle every operational step sharply raises that risk.

Two kinds of work

Work inside an agentic system falls into two groups.

  • Non-deterministic work, owned by the agent: planning, reasoning, classification, summarization, choosing the next step and replanning
  • Deterministic work, owned by the harness: payments, cloud infrastructure provisioning, legal approvals, compliance audits, rollbacks and SLA timers, which track service-level commitments

The second group must never be hallucinated or improvised by a probabilistic model. Restarting a Kubernetes cluster is a good example. It requires the same validation and deployment steps in the same order every time. An agent can reasonably conclude that a restart is needed. How the restart happens must be strictly controlled by the harness. In production, the harness should also add a human approval gate, such as a Slack message asking an engineer to approve before any destructive action. Trusting a model to follow safety policy without slipping is unsafe. Execution constraints belong in code.

The agent decides The harness decides
What to do What is allowed
Why to do it What has already happened
How the plan adapts to change Retries, approvals and execution order
Idempotency, meaning a repeated action has the same effect as one run
Compensation, meaning a registered action that undoes another

An AI core inside a sturdy frame that links databases, cloud servers, a human reviewer and outside services

▲ The harness wrapped around an agent

Durability is the cost of admission

Production agent workflows do not fit inside a single HTTP request. They run for hours, days, weeks or months. An order management workflow, for example, may wait days or weeks on shipping carriers and warehouse approvals. During that time the harness has to wait for human reviews, react to asynchronous events, fire on schedules and recover from crashes. On distributed cloud infrastructure, network partitions, container crashes and hardware failures are inevitable.

That makes durable execution a baseline requirement rather than a premium feature. The runtime must resume exactly where it stopped without duplicating work or losing state.

A second rule follows from this: replanning changes the future, never the past. Completed work, side effects, approvals and visible failures stay locked in durable storage. When something unexpected happens, the harness feeds the current state back to the agent, and the agent plans new steps forward from there. The useful control loop sits not inside the agent’s prompt chain but at the outer boundary, where real-world results become the agent’s feedback.

Free planning, fixed vocabulary

If an agent writes new plans at runtime, how can the harness guarantee how an unseen plan behaves? The answer is a bounded alphabet. Planning stays open, but the set of executable operations is registered at deployment time. This pattern is called a late-bound saga. A saga has traditionally been a hand-coded, deterministic workflow. Here, the agent chains tasks at runtime from a pre-registered catalog.

Each operation that changes state is registered with a compensation task that reverses it.

Operation Registered compensation
Trip a circuit breaker Reset the breaker
Provision capacity Deprovision that capacity

This differs from a branching workflow. In a closed-world branching flow, every path must be drawn at deploy time, and the planner only picks an existing branch. A late-bound saga is open-world: the agent can compose new combinations of the available tools, and engineers do not have to hardcode every future path. Because predetermining every runtime permutation is close to impossible, this looks like a practical trade-off between flexibility and guarantees.

An incident response agent, step by step

A demonstration of an SRE remediation agent, an agent that responds to site reliability incidents, shows how the pieces fit. The scenario was prepared in advance for consistency, but the flow is instructive.

  1. Connect an existing agent to the harness. Mapping the agent’s actions to registered tasks takes about 12 lines of code.
  2. In the first cycle, the agent inspects log evidence and queries metrics.
  3. The model receives that context and returns a structured JSON plan listing the tasks to run.
  4. A compile tool turns the plan into an inspectable DAG, a directed acyclic graph that lays out steps and their order, and Conductor executes it.
  5. The plan rolls back the failed service deployment and keeps polling metrics to confirm recovery. Each step moves from in progress to completed, with full execution records.
  6. In the next cycle, the agent sees the recovered state, recognizes the rollback worked and stops without repeating it.

A monitor showing connected workflow nodes turning green beside a server rack returning to a healthy state

▲ Rollback and recovery check flow

Conductor is an open-source orchestration engine first built at Netflix and released under the Apache 2.0 license. It runs agentic workflows, long-running loops and microservice orchestration with persistent state, and it works with agents built on LangChain, the OpenAI Agents SDK or custom SDKs.

A safety checklist for agents in production

The core idea is simple: give judgment to the model and execution to deterministic code. If you run or plan scheduled, event-driven or long-running agents, check the following.

  • Business logic is split into model reasoning steps and deterministic tasks, not mixed in one prompt loop
  • The model cannot directly execute payments, cluster restarts or compliance enforcement
  • The agent chooses only from pre-registered operations and cannot generate arbitrary commands
  • Every state-changing operation has a defined compensation task
  • Risky production actions pass through a human approval step, such as a Slack approval
  • Completed work and side effects are kept in an immutable record, so replanning only changes later steps
  • Model plans are compiled into an inspectable DAG before execution, so they can be traced and safely replayed
  • The execution layer survives network drops, machine restarts and long waits