An AI agent can edit a product listing, but it cannot assume the product will fit inside a vending machine—or that a completed software task changed anything in the physical world. Those gaps shaped a four-month Prosus experiment in autonomous vending. The project targeted a fleet of six machines; two were operating at an Amsterdam site during the pilot. It tested what happens when an agent must manage stock, sales, prices, and work that still requires human hands.

From digital controls to physical stock

Prosus began with a question relevant to the small merchants on its digital marketplaces: could an AI agent take on routine business operations beyond software? A restaurant was considered too complex for a first test. Vending machines offered a narrower operation, but even that choice complicated the original expectation that setup might take a day.

Hardware was the first constraint. Some advanced machines had quoted delivery times of six months, so the team chose a local Dutch supplier whose machines offered remote monitoring, a management dashboard, and a mobile payment interface accessed by QR code. The dashboard was built for people, not agents. Engineers mapped its functions into tools the agent could call. That work took less than a working day, but the supplier’s customer-facing point-of-sale system, or POS, allowed products only from a fixed catalog. Prosus built a separate POS so the agent could change product descriptions, listings, and prices; that added two days.

This created two distinct control paths. One reached machine functions, such as dispensing an item. The other changed what customers saw and paid. Keeping those paths distinct mattered: a catalog edit is reversible in software, while a dispense command moves real inventory.

A vending machine linked to a cloud workspace and payment interface by separate abstract control paths

▲ Separate machine and catalog controls

An agent that wakes up, acts, and checks its work

The system used the Claude Agent SDK and two main command-line skills—sets of instructions and tools the agent could use to control the machines and the POS. Other connected tools supported communication, browsing, and reporting. Rather than run constantly, the agent followed a scheduled cycle:

  1. Wake: Start a scheduled business task.
  2. Act: Use the available tools to perform it.
  3. Remember: Save context for the next run.
  4. Review: Compare the outcome with the goal.

Cloudflare Workers supported the shared front end, while Cloudflare VMs provided isolated environments for agent execution. Cloudflare R2 held documents and skills that the agent could load, and Cloudflare D1 stored execution records. A credential system supplied access to connected services without placing secret keys in the model’s working context or logs.

The team also changed how the agent kept its memory. Instead of merely shortening a long conversation, it saved decisions, findings, next steps, and references to working documents. That helped preserve business goals between scheduled runs. Shared access became another operational requirement: when one administrator was away and the agent needed help, the system stopped making progress for a week. A common interface let multiple team members inspect its state and respond.

None of this guaranteed that an action had worked. Early dashboards marked tasks complete even when the physical machine had not carried them out. The pilot showed why an agent’s successful run and a verified change in machine state need separate checks.

Where the physical world broke the plan

One test reached live dispensing hardware when it should not have. The agent repeatedly activated the dispenser and sent roughly 30 drinks onto the floor. The practical safeguard is to separate simulated checks from commands that actuate a real machine, then verify physical results independently.

Shopping decisions exposed a different gap. After a suggestion to sell noodles, the agent ordered cup noodles too large for the dispensing coils; the team gave them away in the office. An inventory search returned 1,100 irrelevant discounted listings for soap bottles. Other unsuitable choices included cleaners and oversized snack bags. The agent could follow a request to find products or bargains without accounting for what the machine could hold and sell.

The response was to put slot dimensions and product requirements into documentation the agent could use. The team also found that business communication needed clearer roles. Asked for status updates, the agent produced software-style changelogs rather than sales and margin information. Asked to promote the machine, it initially posted plain text to Slack despite access to an image-generation tool. Separating management from marketing work and giving the agent a broader goal encouraged it to propose and review further tasks instead of treating one post as the whole job.

Oversized food and unsuitable goods beside a vending machine, with packaged snacks on a restocking cart

▲ Physical fit and inventory selection

Humans remained part of restocking. Vague requests to add an item left operators unsure which machine or slot to use. A dedicated operator interface turned those requests into guided steps and let staff ask the agent for clarification. In this setting, autonomy depended partly on making the remaining human work easier to carry out.

Sales exposed a pricing problem

Sales records ran from May 20 through September 15, 2026. The pilot began with free items to observe demand. Once autonomous pricing started, prices first rose high enough to suppress sales. After a summer demand drop, the agent cut them sharply; some protein shakes became cheaper than those at a nearby supermarket. More orders did not solve the merchandise shortfall, and free office snacks and drinks also competed for customers.

Four-month pilot measure Reported amount
Vending sales €800
Wholesale food purchases €1,100
Merchandise cash shortfall €300
Model token cost Approximately €300 per month

The €300 shortfall compares sales with food purchases; it does not include the additional monthly model cost. These are results of this pilot, not a forecast for other vending businesses. They show why an automated pricing system needs limits that protect margins when demand changes.

What the pilot leaves for builders

Prosus has also released Prosus Vending Bench, an open-source simulation with six machines, 30 operational days, and a €1,500 starting budget. It offers a way to test agent designs without moving real stock.

The central lesson is operational rather than conversational: an agent needs accurate physical constraints, independently checked outcomes, pricing safeguards, and a workable handoff to people. Teams considering similar systems should define those boundaries before granting an agent control over inventory, payments, or hardware.