A passing test suite can show that an AI agent’s code agrees with the tests it wrote. It cannot, by itself, show that either one meets the requirement. In one engineering case, more than 12,000 agent-written unit tests stayed green while higher-level acceptance tests found defects. The practical challenge is to check code independently, verify that tests can detect faults, and retain evidence of test health after each CI run ends.
When code and tests share an assumption
At one engineering organization, AI agents wrote both application code and unit tests. Its suite contained over 12,000 passing unit tests, yet user acceptance tests still detected defects that the unit tests missed. Those acceptance tests flagged defects in 80 of 2,000 runs; that figure counts runs, not necessarily distinct bugs.
The concern is a shared misunderstanding. If an agent interprets a requirement incorrectly, it can implement that interpretation and write a unit test that confirms it. During iteration, an agent may also change a failing assertion instead of correcting the implementation. A passing result then verifies agreement between the code and test, not agreement with the intended behavior.
A cited analysis of more than 86,000 agent-authored test patches found weak or missing test-oracle assertions in 80.2% of them. A test oracle is the check that decides whether an observed result is correct. In the same analysis, 34% of patches passed functional tests but failed the full pull request requirements. These findings make a case for checks that do not simply repeat the code-writing agent’s assumptions: contract tests for agreed behavior between components, architectural checks for system-level rules, and black-box acceptance tests that assess behavior from a user’s perspective.
Choose Playwright tooling for the job
Testing cost also depends on how an agent receives browser state. The Playwright Model Context Protocol server, or Playwright MCP, sends a full accessibility snapshot into a tool result. That snapshot describes page structure and controls, but a complex page can consume substantial space in the model’s context window—the information it can consider at once. Playwright CLI, a command-line interface, instead returns a path or URL for a stored snapshot that the agent can retrieve selectively.

▲ Direct snapshots and selective retrieval
One comparison reported different credit use for two tasks:
| Task | Playwright CLI | Playwright MCP |
|---|---|---|
| Run existing test suites | 1.2 credits | 1.5 credits |
| Explore a site for critical bugs | 5.3 credits | 0.6 credits |
Those are results from the reported comparison, not prices or savings every team can expect. They suggest using CLI for repetitive, structured runs where a compact response helps preserve context for other coding work. MCP may be the better fit when an agent needs continuous page feedback to explore an unfamiliar flow. The lower-credit option changes with the task.
Check meaning without relaxing hard requirements
AI features create a different assertion problem: a conversational assistant may express the same outcome in different words. Jev evaluates application state against typed questions and returns probabilistic judgments rather than relying only on exact string matches. Its Playwright primitive ai.expect().toSatisfy() can check whether a response means a parcel has shipped, while rejecting a response that says the order remains pending. Another primitive, ai.act(), supports navigation when a changing interface moves a control, such as during an A/B test.
Semantic checks can accommodate wording and layout variation, but they should not replace checks of concrete business state. A support interaction, for example, can also be checked for an HTTP 201 response or a database ticket. Keeping those programmatic assertions alongside semantic ones reduces the chance that a plausible-sounding response passes while the underlying action did not occur.
Test whether the tests can catch a fault
Mutation testing deliberately alters software to see whether a test suite detects the resulting violation. In one experiment with rqlite, a distributed database built on SQLite, an agent-driven approach introduced 19 distinct mutations while examining 13 safety properties. Across 46 test runs and 24 hours of fuzzing—automated testing that exercises software with varied inputs—the mutations falsified 11 properties and exposed three previously unknown bugs, which were patched.
The other two properties were not violated in those runs. That outcome does not establish that they are safe under every condition. It identifies checks or parts of the test environment that may need stronger scrutiny. Mutation testing is valuable here because it asks a more pointed question than whether tests pass: will they fail when a behavior they are supposed to protect goes wrong?

▲ Fault detection and production signals
A separate GitHub security research effort uses Taskflow Agent to reduce repetitive fuzzing work on C and C++ projects. It prepares fuzzing harnesses—test wrappers that let a fuzzing tool exercise code—runs AFL++, and reviews coverage as it iterates. Its reports and suggested fixes still require human review. Automated discovery can help direct attention to potential faults, but a generated security patch should not be accepted solely because an agent proposed it.
Keep the evidence after CI ends
Test results can lose value when they remain only in short-lived CI logs. An approach used with Cypress turns results into persistent Prometheus metrics: the before:run lifecycle hook attaches a run identifier, and after:spec records pass and fail counts and durations. Short-lived test processes send the metrics through Prometheus Pushgateway; Grafana Alloy then collects and routes them to Grafana Cloud.
Attaching GitHub Actions run IDs lets a team trace a change in test duration or reliability to a particular workflow run and commit. Dashboards can show duration trends and intermittently failing tests alongside production telemetry. This does not replace a test runner, and it does not make a passing suite proof that production is healthy. It gives teams a longer-lived record for investigating what changed.
Pre-release checks can also compare proposed agent behavior with earlier behavior. Raindrop Simulations runs against proposed pull requests, replaying production traffic and existing tests to flag unexpected deviations before an update ships. Such comparisons add another signal, though a deviation still needs interpretation.
Visibility matters more as workflows chain agents together. In an illustrative model, multiplying a 95% success rate across 10 dependent steps yields about 60% end-to-end success. That is a model, not a measured failure rate for every agent system. A January 2026 report found that fewer than one in 10 enterprise AI applications had full end-to-end observability, while a March 2026 survey found that 77% of IT organizations lacked full visibility across hybrid environments. Those reported gaps make it harder to locate a failure after deployment.
Build a verification loop, not just a green dashboard
Teams using agents to write code and tests should treat unit tests as useful statements of expected behavior, not the final verdict. Add independent acceptance, contract, and architectural checks against the requirements. Choose Playwright CLI for structured batch runs or MCP for exploration based on the work and its measured cost. Probe important checks with mutation testing, and require human review of generated security fixes. Finally, retain test metrics with run identifiers and examine them alongside production signals. Together, these steps make it easier to learn both whether a change passed its tests and whether the tests were capable of noticing a problem.