A whole tool category arrived this year to solve one thing: getting you mentioned by an AI assistant. It goes by two labels that mean roughly the same thing, generative engine optimization and answer engine optimization, GEO and AEO, and it sells visibility scores, mention tracking, and share of voice in ChatGPT, Gemini, and the other AI assistants. That is a real problem worth solving. It is also the first of four.

Being mentioned is not being listed. Being listed is not being chosen. Being chosen is not being bought. Four gates, four different ways to fail, and most stacks have nothing pointed at the middle two.

That gap has a name, or it does from here on: the product context layer. It sits between your catalog and the AI assistants recommending from it, and it is the layer almost nobody has assigned to anyone.

What this piece covers:

  • Where the agentic commerce stack actually stands in late 2026, and which layer is missing from it
  • Why Mentioned, Listed, Chosen, and Bought fail for four different reasons
  • What the product context layer holds, where that context comes from, and why the work never finishes
  • How it differs from a product information management system (PIM), feed management tools, a product data layer, and search optimization
  • Why the product context layer, and not the connection to the AI assistants, is what decides the sale

The connection stopped being the hard part

For a year the whole industry bet on the pipe. Whoever owned the connection between a brand's catalog and ChatGPT, Perplexity, Gemini, or Google AI Mode, and owned the checkout sitting inside it, would own agentic commerce. That assumption is behind most of the protocol announcements and partnership press releases of the last eighteen months.

The bet lost, and it lost inside six months. Three separate things settled it.

The protocols were published openly. OpenAI and Stripe put out the Agentic Commerce Protocol, Google and its retail partners put out the Universal Commerce Protocol. A connection method published as a specification is a connection method nobody owns.

Platforms switched it on for free. In March 2026, Shopify made its agentic storefronts channel active by default for eligible stores across ChatGPT, Google AI Mode and Gemini, Microsoft Copilot, and Meta, with no fees beyond standard payment processing. Merchants did not buy the connection. It arrived.

Then the payment layer absorbed it. Stripe's Agentic Commerce Suite takes your catalog by CSV, API, or a connection to your commerce platform and syndicates it across supported agents, so you publish once instead of integrating with each shopping agent. Adyen ships Agentic Feed, pitched as one product feed that adapts itself to each AI platform's requirements. PayPal's agentic commerce services include catalog and order management, with Store Sync putting a merchant's catalog inside large language models directly.

Three of the largest payment providers, all doing catalog intake and agent distribution. When that becomes a bundled feature of taking payments, it has stopped being a product anybody sells on its own.

OpenAI's own retreat is a footnote to that rather than the cause of it: it retired the first Instant Checkout in March 2026 and went back to discovery. The reason everyone is racing to offer the connection is that it is cheap to offer and valuable to be the default for.

So being connected buys you nothing. Everybody eligible is in, at no cost, most of them without asking. And look at what all of those specifications actually carry: identifiers, titles, descriptions, pricing, fulfillment, media. Product data, moved efficiently. Not one of them has a field for who a product suits, where it stops being the right answer, or what backs that up.

The pipe got solved. What travels through it did not, and that is the whole opportunity.

Four-tier diagram: AI assistants where shoppers ask, the connectivity layer that is open and on by default, the product context layer where the recommendation is decided, and your commerce platform underneath

Four gates, and only one of them is measurable

We went deep on this in a separate piece, so here is the short version. What used to be a funnel, discover and then compare and then decide and then buy, now happens inside a single exchange. Four gates are what is left of it, and a product has to clear all four.

Mentioned. Named in the answer at all. This one runs almost entirely off the crawl, the editorial slice of the open web. It fails by your never being named, by being described wrong, or by only competitors turning up.

Listed. Shown as a buyable product, with price and stock attached. Reachable either way today, though only the feed path reliably produces a purchasable card. It fails on missing variants and stale availability.

Chosen. Picked over the alternative, and picked for the right reason. Decided by product context, on whichever path delivered the offer. It fails when your copy is too broad to match a constraint, so the AI assistant guesses.

Bought. Survives the click and becomes an order. It fails on a price or stock conflict at the moment of arrival.

Facts get a product listed. Context gets it chosen. That piece carries the full mechanism, including why generative engine optimization only reaches the Mentioned gate, and what the retrieval data shows about the other three.

The four gates are not equally visible to you, which is the practical reason this layer goes unnamed and unowned. Mentioned can be watched from outside, so tools watch it and sell you a score for it. There is no public benchmark for how often a mentioned product becomes the chosen one, because that happens where the answer gets assembled and nobody has a counter on it. Building a strategy around the gate you can see is a natural mistake and an expensive one.

What the product context layer holds

Define it by what it does rather than by adjective. Three movements: what it draws on, what it produces, and why it never finishes.

1. Where the context actually lives

This is the part that surprises people, because "product context" sounds like a tidier product description. The catalog is one input among many, and usually not the most useful one.

Take one product and keep it in view for the rest of this piece: a $140 merino crewneck sweater, invented for the occasion so the numbers can be specific. Here is where its context actually lives.

  • The catalog itself. Titles, descriptions, attributes, prices, inventory, variants
  • Customer reviews. Ratings, the distribution behind them, the questions buyers ask, the complaints that repeat
  • Support conversations. What people get confused about, what breaks, what they return and why
  • Sales conversations. The objections a real buyer raises, and what actually answers them
  • Product use cases. How, when, and where the thing is used, including the uses you did not design for
  • Brand and product context. Mission, materials, ingredients, technology, standards, the references you can point to
  • Third-party evidence. Independent tests, certifications, editorial coverage, retailer listings that agree with your specifications
  • The comparison set. What the products an AI assistant will place next to yours actually claim: their specs, their price bands, the dimensions where they are weaker
  • External signals. What is trending, seasonality, occasions, the language shoppers are currently using for the thing you sell

Most of that already exists inside the business. It is sitting in a helpdesk, a reviews widget, a sales call recording, a spec sheet, and somebody's head. None of it is in a form an AI assistant can use, which is why a complete catalog and a strong brand can still lose the recommendation.

2. From scattered sources to one machine-readable answer

Everything above arrives in the wrong shape. A support ticket is a conversation. A review is prose. A wash test is a PDF. A shopper's question is a sentence with the constraints buried inside it. None of those line up with each other.

Synthesis is the unglamorous part where they get resolved against each other. The ticket that says "itchy at the collar" and the spec that says 18.5 micron are about the same property, so they have to end up in the same field, with the review count attached as support and the boundary written down somewhere it can be read. What comes out is one representation of the sweater that a machine can act on.

And the test of that representation is not whether four boxes are full. It is whether it can answer the question as a shopper actually types it.

One warm layer that packs small for Scotland in October, under $150.

That one sentence carries five constraints: warmth, packed size, working alone rather than in a system, a season and a climate, and a price ceiling. Nothing in a product title answers it. Four things have to be present before anything can.

Structured attributes. For the sweater: 18.5 micron, 280 grams, next-to-skin rated, machine washable cold, and not suitable for a diagnosed wool allergy. Not "premium merino, beautifully soft". Structured means labeled fields a machine can read, not a paragraph a person is expected to read. That last item, the boundary, is the piece most often missing and the most useful, because it is what lets an AI assistant avoid recommending you badly, which is how brands collect bad reviews on good products.

Evidence. The support behind each claim, checkable by something other than a person. For the sweater: the micron count on a spec, 412 reviews with nine of them saying it runs small, an exchange rate lower than the category, a third-party wash test. "Customers love it" is not evidence. "412 reviews, 4.6 average, the recurring complaint is sizing" is.

Comparability. The dimensions on which this sweater can be placed against the two others the AI assistant is holding: 280 grams against their 340 and 310, $140 against $95 and $190, next-to-skin rated where one of them is not. Not marketing differentiation. An AI assistant asked to choose has to compare on something, and a product that cannot be compared is easier to leave out.

Availability and price certainty. In stock, in that size, at $140, shippable to that address, today. This is the one that converts, and it is the one that decays fastest. An AI assistant that cannot confirm availability has a good reason to recommend one it can confirm, and we have watched exactly that: four AI assistants describing an in-stock product as backordered.

Read those four back against the one sentence the shopper typed. Attributes answer warmth, packed size, and price. Evidence answers whether to believe any of it. Comparability decides the sweater against the other two in the shortlist. Availability decides whether the answer still holds in October, in that size, to that address.

That collection has a name, and it is worth fixing, because a thing without a name does not get a budget line or an owner:

The product context layer sits between a merchant's catalog and the AI assistants that recommend and sell from it. It takes the base catalog, price, title, description, and adds the context a catalog cannot carry: structured attributes, evidence, comparability, and availability. The output is a representation an AI assistant can match against a shopper's stated intent and stand behind as a recommendation.

Run it as a completeness test: for each of the four, either the information exists in a form a machine can consume, or the decision gets made without you.

3. Why the work never finishes

The tempting way to read all of this is as a project. Gather the context, publish it, move on.

That fails, for a reason worth understanding. You do not know in advance which expression of the sweater wins which question. Whether "warm without bulk" or "280 grams" is the phrase that earns the recommendation is an empirical question, and the answer differs by surface, by category, and by how shoppers happen to be phrasing things this season. Meanwhile the surfaces themselves keep changing what they accept and what they favor, as this year has made obvious.

So the layer is a running process rather than an artifact. It keeps producing variants of how a product is expressed, keeps them in front of the surfaces, and keeps whichever ones earn recommendations. A one-time export is a snapshot of your best guess on the day you made it.

Product context is where this layer starts rather than where it ends. What an AI assistant needs in order to sell a merchant's products, and what it would need in order to operate on their behalf, are the same kind of thing: context.

What the product context layer is not

Comparison table: PIM, feed management, a product data layer, and GEO each stop somewhere different, and the product context layer is where the choice gets decided

Four comparisons, because the layer gets mistaken for all four and each mistake leads somewhere different.

It is not a PIM. A PIM is a warehouse, built for internal control, holding what you already decided to record. This layer is built for external inference, holding what someone else needs in order to argue on your behalf. A tidy warehouse full of the wrong fields is still the wrong fields.

It is not feed management. This is the category most merchants already pay for, so be precise about the ceiling: a destination's schema decides what you are allowed to say. Google Merchant Center has a field for material and none for "works over a shirt without bulk." A feed carries only what its receiver agreed in advance to accept, and an AI assistant is not filling in a form.

It is not a product data layer. The closest one, and the words are nearly the same. A product data layer normalizes the facts your catalog already holds, cleaned and mapped so machines read them consistently. Necessary, not sufficient. Data is what your catalog knows. Context is what a recommendation needs and a catalog was never built to hold.

It is not GEO. Generative engine optimization, GEO, makes content retrievable and quotable by AI answer engines. It works, on content, and it is aimed at Mentioned. This layer is aimed at Chosen and Bought. A brand can be excellent at GEO and still lose the shortlist on a missing size chart.

Nor is it schema markup, though schema is one of the ways some of it gets expressed. Schema is a vocabulary for labeling facts on a web page. The question of which facts are worth labeling, and whether the underlying facts are complete and true, sits entirely outside it.

Why the product context layer decides the sale

AI shopping puts the decision earlier than most merchants model it.

Someone opens ChatGPT and asks for one warm layer that will fit in a carry-on for Scotland in October. In a search world they would have landed on your sweater's page and decided there, with your photography, your copy, and your reviews all working on them. In an AI-assisted purchase, the shortlist of two or three sweaters is assembled before any of that, out of whatever the AI assistant could find and trust: weight, packability, whether it holds up damp, the boundary that says it is not a rain shell.

If your 280 grams and your next-to-skin rating are not readable, you are not in that shortlist, and no amount of work on the product page reaches back to change it. A shopper who does arrive on the page arrives already recommended, so the page is confirming a decision rather than making one.

That moves most of your conversion work upstream, into material you may never have written for a human at all.

Two-column comparison: in search shopping the choice is made after the click, on your product page; in AI shopping the choice is made before the click, while the shortlist is assembled, so the page confirms a decision already made

It is also why the pipe commoditizing was good news rather than bad. If the connection were the moat, brands would be renting access on someone else's terms. Because it is not, the durable advantage is in something that belongs to the brand: the quality and completeness of its own product context, portable across every AI surface that appears next.

A call to build the layer

Right now this layer gets absorbed into "AI visibility", where it disappears, because visibility tools measure the mention and stop there. A named category gets an owner, a budget line, and a standard you can hold a vendor to.

Nile builds the product context layer for brands: gathering that context, synthesizing it into intent-matched product cards, distributing them to AI shopping surfaces, and tracking the path from click to purchase. Across the brands doing this work, the pattern as of August 2026 is a click conversion lift of 150% to 300% against the brand's own site.

Two things about that number. Attribution is last click, the most conservative model available, and the comparison is before and after on the same store rather than against a control group. Post-purchase surveys at a few merchants returned AI-influenced orders at three to five times the last-click count, so treat it as a floor rather than a result.

The economics follow the shape of the thing. This is AI context infrastructure: built once, reused across every AI surface, and priced as a small percentage of what it produces rather than as a fee for placement. Nothing here is for sale in a ranking.

We are not claiming the word. We are claiming the gap is real, that it is where the next several years of ecommerce advantage sits, and that treating it as a visibility problem will keep producing mentions that do not convert.

Where to start

One question tells you whether any of this is your problem, and the test takes ten minutes.

  1. Write the buying question your best customer would actually type, constraints included, not the category name. "One warm layer that packs small for Scotland in October, under $150", not "merino sweater".
  2. Put it to ChatGPT, Gemini, and Perplexity, the same sentence in each.
  3. Read the answers for two things: whether you appear at all, and whether the reason given for the products that did appear is a reason you could have supplied.

If you are absent, you are stuck at Mentioned or Listed, and that is a reachability problem. If you appear but the reason is generic, you are stuck at Chosen, which is the gate with the revenue attached and the one no dashboard will tell you about.

Either way you now know which of the four is broken, which is more than an AI visibility score will give you. Fixing Chosen is the work described above, and it is work on your own data rather than on a channel somebody else controls.

If the test lands you at Chosen, there is also an order to the fixes. From our own optimization work so far, and this is practitioner experience on a limited sample rather than a benchmark, the returns run roughly in this order:

  1. Structured facts. Attributes in labeled fields a machine can read, boundaries included.
  2. Availability and price certainty. The pair that converts, and the pair that decays fastest.
  3. Use cases and audience fit. Who the product suits, when it is the right answer, and when it is not.
  4. Visual attributes written out as text. What the photos show that the copy never says.
  5. Reviews and policies. The evidence behind the claims and the terms behind the purchase.

Generic marketing copy contributes the least of anything a catalog carries. Expect that order to move as the surfaces do, which is one more form of the work not finishing. Treat it as a starting sequence, not a law.

For what Chosen looks like when it fails, four AI assistants once described an in-stock product as backordered, and fixing the facts did not make the product recommended.

If you would rather see the answer than run the test, connecting your catalog to Nile takes a few minutes and costs nothing until it earns you a sale.