An AI agent can fail in a way that every server log reports as success. Ask a toy store’s shopping assistant for “all the toys under $6,” and it replies that nothing matches, even though cheap toys are in stock and every request returns a healthy status. The cause is simple: the search tool behind the agent has no price filter at all. What makes the case worth studying is who diagnosed and fixed it: not an engineer reading logs, but Claude Code, once it was given access to the agent’s execution records. Arize AI, a company that builds observability and evaluation tools for AI applications, uses this example to sketch where agent debugging is heading: away from humans scanning dashboards and toward coding agents that read the evidence and write the fix.
Why code is no longer enough to understand an agent
In traditional software, the source code is the source of truth. Every execution path can be analyzed and predicted before the program runs. Agents break that assumption. Given the exact same input, a language model agent may follow a different line of reasoning, call different tools, and return a different answer each time. Reading the code or the prompt does not tell you what the agent actually did.
That role moves to traces. A trace is a structured, nested record of one agent run: every conversational turn, tool call, prompt, model response, token count, latency figure, and dollar cost. Each step is stored as a span, the basic unit of work inside a trace. Building an agent without tracing, in Arize’s framing, is flying blind: you can see your commit history but nothing about how the agent behaves with real users.
The catch is volume. A production service handling thousands of requests per second produces millions of traces, and spot-checking them by hand breaks down once traffic passes a few requests per second. The standard answer in 2025 was LLM-as-a-judge evaluation, or evals: a language model grades each trace against defined criteria and returns both a verdict, such as a score or pass/fail flag, and a written explanation of what went wrong.
By 2026, Arize argues, that layer has become its own bottleneck. Agents can write code and finish tasks in minutes, while people need days to dig through tens of thousands of eval explanations and spot regressions. The proposed shift is for observability data to be consumed by agents rather than by humans.
Adding tracing to a sample shopping agent
The example is a small e-commerce site with a chat assistant that recommends toys. It is built on the OpenAI Agents SDK and, at the start, records no traces at all. Asked for toys for a seven-year-old, it suggests items such as a 450-piece robot building kit, but there is no record of how it reached that answer.
The tracing work was handed to Claude Code running as an extension inside VS Code. The instruction was a single sentence: add observability with Arize AX to this application, using the credentials already stored in the local environment file. Claude Code examined the project, installed dependencies, and configured an exporter to send data to Arize.
The job was small because the OpenAI Agents SDK already ships with OpenInference instrumentation. OpenInference is a vendor-neutral open standard for recording what AI applications do. All that remained was to point it at the right endpoint with the right keys, so traces began flowing without changes to the agent’s core logic. One caution: Claude Code also tried edits nobody asked for, such as rewriting the README, so it pays to review what an agent changes.

▲ Empty search results behind a healthy status
The failure hiding behind 200 OK
With traces in place, a price-bounded request exposed the problem. Asked for every toy under $6, the assistant said it could not find any matches. Nothing crashed, and no error code appeared.
Next, Claude Code was asked to pull the recent traces and look for anything going wrong. This is where skills come in. A skill is a packaged set of instructions and tool commands that teaches a coding agent how to do a specific job; placed in a designated folder at the repository root, it becomes available to the agent. Out of the box, a coding agent cannot query an observability platform. With a trace-inspection skill installed, Claude Code used Arize’s command-line tool to fetch spans and analyzed their status, input arguments, and outputs on its own.
Its diagnosis was specific:
| Finding | What it meant |
|---|---|
| 42% of recent searches returned zero results | The search tool often came back empty |
| The product search tool finished in 1 millisecond | It returned an empty list immediately on unmatched filters |
| Some calls had every search argument set to null | The tool fell back to generic top products |
| Minimum and maximum age were set to the same value | The range was too narrow to match anything |
| The model passed a price limit that the tool ignored | The search tool had no price filtering |
The key detail is that most HTTP calls returned 200 OK. Quality failures in agents often show up as silent empty results rather than exceptions or error codes, which means standard HTTP monitoring will not catch them. Tools that return a success status with an empty array deserve their own check.
How the coding agent fixed and verified the bug
Asked to add price filtering, Claude Code worked out how to extend the search tool’s input schema and execution function. Then it noticed a neighboring folder that already contained a finished reference solution, and simply copied that implementation over. The result worked, but it illustrates a real risk: an agent with access to the whole repository may borrow code from solution folders or tests instead of solving the problem. If a repository holds reference answers, expect the agent to find them.
After the change, a request for toys under $9 returned items priced at $7.99 and $8.99, and the trace captured the full path from input through the search tool call, the budget constraint, and the final response. The loop looks like this:
- Add tracing to an agent that has none.
- Send requests likely to fail so problems surface.
- Let a coding agent with a trace-inspection skill analyze the failure patterns.
- Have the agent write the code that fixes the root cause.
- Repeat the same kind of request and confirm the fix in the traces.
Taking humans out of the investigation step
Arize wants to remove even the step where a person asks the agent to investigate. Signal is an agent built into Arize AX that, by default, scans collected traces every six hours, groups recurring failures, categorizes them, and proposes fixes. The interval is configurable.
Examples of issues it flagged include:
- A prompt injection, where user input overrides the agent’s instructions, led the toy assistant to obey the user instead of acting as a sales representative
- The agent invented an age constraint from vague wording
- Asked how to build a bomb, the agent did not refuse; it suggested bomb-themed toys and added unsolicited advice about taking a time-out
The last case was classed as an unacceptable escape from the bot’s domain. A toy store assistant should not be steering anyone toward weapon toys or offering psychological counsel. The lesson drawn is that dangerous or out-of-scope requests are better rejected outright by a guardrail, a rule that blocks unsafe inputs or outputs, than answered helpfully. For a financial assistant agent, Signal grouped patterns such as endorsing harmful financial advice, inventing financial claims, and choosing the wrong tool after a prompt change.
For each issue, Signal offers several actions: file a GitHub issue with the failing traces and a suggested remedy, add the traces to an evaluation dataset, create an evaluator, or open a pull request that changes the code. Arize’s in-product assistant, Alex, can also turn a plain-language request into an eval. Asked to build an eval ensuring the app never explains how to make a bomb, it generated the evaluation on the spot.
A few operating details matter for anyone weighing the approach:
- Signal’s code changes are currently produced by Claude Code, with plans to let teams bring their own coding agent and model.
- Arize currently absorbs the cost of running Signal but does not promise it will stay free.
- Signal recommends new evaluators rather than creating and running them on its own, because continuously running model-based evaluators costs significant time and compute, so a human must approve them.
- Keeping traces solely on a local device is not supported, but a full on-premises deployment inside a customer’s own hardware is.
- Clustering adapts to volume: with a dozen test traces it groups individual traces, while with thousands of production traces it groups at a much larger scale.

▲ The loop from observability to automated fixes
What has to be true before you trust automated fixes
How can a self-repairing loop be safe in regulated fields such as finance and healthcare, where hallucinations and adversarial inputs carry real consequences? The short answer offered is layers: multiple tiers of verification that supervise and validate each code change, rather than unconstrained autonomy. Arize says it already has financial-sector customers running these observability and automated remediation patterns in production.
The condition stressed most is regression evaluation. Any prompt tweak or automated pull request needs a suite of evals that confirms existing capabilities still work, and automated PRs should be reviewed by a person and pass those regression checks before merging. Approving every command an agent requests without reading it is also a habit security teams would rightly object to.
Key takeaways and where to start
The quality problems in an agent show up in its traces, not its code. A loop that records traces, groups failures, applies fixes, and verifies them lets a team keep improving an agent without reading every record by hand.
- Check whether your agent framework already supports a standard such as OpenInference before writing custom tracing code.
- Look specifically for tool calls that return a success status with empty results; ordinary monitoring misses them.
- Give your coding agent a skill that lets it fetch and analyze traces.
- Triage failures as recurring patterns rather than one at a time.
- Write eval criteria that explain why something failed, then feed those explanations back into fixes.
- Merge automated fixes only after regression evals pass and a person has reviewed them.