Product catalogs built for human keyword search are a poor fit for AI shopping agents, and fixing them takes structured, focused data rather than more text. That is the core lesson from controlled experiments run by PayPal’s agentic commerce team, the group working on how AI agents find products and complete purchases for shoppers. Enriching product data lifted keyword retrieval for every agent tested, and the merchants with the sparsest catalogs gained the most. But piling on unstructured copy blurred semantic retrieval and, in some tests, made agents hallucinate more often.

From typing keywords to handing off the task

Commerce has moved through distinct waves: physical stores, e-commerce websites, email and SMS marketing, social commerce and now agentic commerce, where AI agents shop on a person’s behalf. PayPal’s director of product for agentic commerce frames each shift the same way: it first feels like an existential threat to incumbents, then becomes a large opportunity, because commerce always adapts when consumer habits change.

Shopping behavior is moving through three stages:

Stage What the shopper says What merchants must do
Search Types “school supplies” and clicks through listings Invest in SEO and their own mobile apps
Intent “Find back-to-school clothes and supplies for the twins starting 5th grade in September” Become discoverable inside AI engines
Delegation “Get the back-to-school shopping done. Keep it under $300” Convert conversational demand safely

In the search era, the shopper carried the whole burden of sifting through irrelevant results, and businesses spent a decade tuning their sites and product taxonomies for keyword search engine optimization. The market is now entering the intent era, in which people describe their needs to large language models in natural language. Brands have to meet shoppers inside those conversations instead of expecting them to visit a merchant’s own website.

Delegation comes next, but it depends on trust. Shoppers will hand over purchasing authority only after agents consistently earn it with accurate, context-aware recommendations.

How big agentic commerce could get

PayPal’s agentic commerce lead stresses that this is not a distant hypothesis but a change already reshaping e-commerce. The forecasts and early data points behind that view:

  • Bain & Company projects that AI agents will execute 15% to 25% of all commerce transactions by 2030.
  • McKinsey, in an October 2025 study, estimates that agentic commerce will reach $1 trillion in US B2C retail and $3 trillion to $5 trillion globally by 2030.
  • More than 2 billion active users across ChatGPT, Google Gemini, Microsoft Copilot, Meta and Amazon are projected to start their shopping journeys with AI agents.
  • Industry forecasts say software agents will outnumber human internet users within the next decade.

Merchants are already seeing effects. During the 2025 holiday season, AI referral traffic to retail sites rose 693% compared with 2024. Some 39% of US shoppers already use generative AI tools to help them shop. Retailers also report checkout conversion rates four times higher when shoppers engage with an AI agent during their visit than with ordinary site browsing, because the agent supplies detailed product information tailored to the shopper’s needs. One mockup shows an AI assistant confirming a $69 order from The RealReal, with home delivery and order tracking.

Small robotic assistants carrying shopping bags between glowing storefronts in a digital marketplace

▲ Agents doing the shopping

Catalogs were built for people, not agents

The central question for merchants is whether their products surface wherever shoppers talk to AI agents on third-party platforms. Agents such as ChatGPT Shopping, Perplexity Pro and Google Gemini route a query straight to individual products rather than to category landing pages.

Most retail catalogs were not designed for that. They rely on keyword-stuffed titles, internal SKU codes, the stock-keeping identifiers merchants use to track inventory, and schemas built for Google Shopping feeds. Several other problems compound this:

  • There is no industry standard for what makes a catalog “AI-discoverable.”
  • Many enrichment vendors claim to make inventories discoverable by AI, but merchants have no reliable benchmark to check those claims or compare methods.
  • Agents switch among different large language models to rank and recommend products, so their behavior is hard to predict.
  • Retrieval from vector databases, which store text as numerical representations of meaning, is neither fully deterministic nor easy to explain across platforms.

In short, agent ranking is a black box that traditional, rule-driven SEO cannot optimize. The practical takeaway is to avoid tuning a catalog to any single retrieval system. To find strategies that hold up across platforms, the PayPal team ran controlled enrichment experiments on merchant product catalogs across multiple agents.

Keyword search and semantic search behave differently

The experiment results make more sense once the two main retrieval methods are clear.

Aspect Keyword search (BM25) Semantic search (embeddings)
How it matches Exact words shared by the query and the product text Meaning, after converting query and text into vectors
Strengths Precise, fast and predictable Finds items that fit the intent even with different wording
Blind spot Misses synonyms and intent Loses focus when one vector holds too much unfocused text
Example “blue running shoes” returns only listings with those words “something comfortable for marathon training” surfaces running shoes

BM25 is a keyword search method that ranks products by matching the exact words in a query to the words in the product text. With keyword search, a query like “footwear for jogging” can return nothing if those words never appear in the product data. Semantic search uses embeddings, numerical representations of a text’s meaning, so it can connect a vague request to the right product. Its weakness is that a single embedding forced to represent a long, scattered description stops pointing clearly at anything. A hybrid approach that combines keyword precision with semantic intent provides the strongest foundation.

Five ways to enrich a product catalog

The experiment treated enrichment not as one technique but as five separate data dimensions, each suited to different inventories.

Dimension What it adds Best fit Weak fit
Attribute Fill Missing required specifications Thin-spec catalogs and small merchants Feeds that are already rich
Description Depth Narrative context and use cases Commodity goods and generically titled SKUs Rigid, structured inventories
Buyer Context Persona, occasion and scenario tags Lifestyle, gifting and fashion Pure-spec utility items
Trust Signals Reviews, brand popularity and sentiment Durable goods and furniture Short-lived flash sales
Product Identity Canonical entities and grouped variants Inventories sold across many channels Not specified

A before-and-after example shows the difference. A sparse listing titled “Men’s Shoe BS43, Blue,” with a one-line description, gives a semantic system almost nothing to match. The enriched version becomes “Men’s Lightweight Running Shoe. Breathable Mesh. Blue” and adds structured attributes, including marathon training as the use case, a lace-up closure, cushioning and mesh material, plus a 4.6-star average across 1,240 reviews. With that data, an agent can connect the shoe to broad, conceptual requests.

A running shoe with a blank tag beside the same shoe surrounded by attribute cards and star ratings

▲ Product data before and after enrichment

What the experiments found

The tests produced three main findings:

  1. Enrichment always helped keyword retrieval. Enriched data beat the unenriched baseline in every head-to-head test and raised the top-10 keyword hit rate across all agents. Longer, more descriptive copy adds term variety, matches more keywords and captures attributes drawn from product images.
  2. Thin catalogs gained the most. Merchants whose product data started out weakest saw the largest percentage lift, so they have the most room for quick improvement.
  3. Unstructured text can backfire. Loosely added content diluted the meaning captured in embeddings. In some tests, overly enriched descriptions led agents to hallucinate, meaning they produced details that were not true, more often and to rank worse in semantic retrieval.

Generic store copy such as “Welcome to our store, we are a family-owned business” adds irrelevant tokens, the small units of text a model processes, and confuses retrieval. Agents appear to reward informative, structured product content over brand equity or marketing claims. Quality, not sheer volume of text or tags, drives search and recommendation performance. Enrichment clearly helps keyword search, but text-stuffing alone is not enough for semantic retrieval, which needs a different data structure.

What merchants should do now

Retailers that keep relying only on legacy Google Shopping feeds built for human search engines risk seeing their discoverability drop sharply. There is no single enrichment formula that works for every catalog, so a targeted approach looks like the safer path:

  • Audit first. Decide whether your catalog is thin, spec-heavy or commodity-focused before choosing an enrichment strategy.
  • Match the method to the inventory. Use Attribute Fill for spec-driven catalogs and Buyer Context for lifestyle categories instead of applying one blanket technique.
  • Stay concrete. Keep descriptions focused on real use cases and specifications, not long narratives or slogans.
  • Remove boilerplate. Strip store-level copy, such as family-owned business statements, out of product feeds.
  • Structure trust signals. Put review averages and review counts directly into product feeds.
  • Plan for both retrieval types. Combine keyword and semantic retrieval so products match exact queries and broad intent.
  • Verify, do not assume. Build a way to check whether AI shopping engines can parse and retrieve your listings, and use it to test vendor claims.

As more shopping starts inside AI agents, a merchant’s edge appears likely to come less from SEO keywords and more from accurate, structured product information that agents can understand. Filling the gaps in a thin catalog with precise data is the most direct first step.