Your dashboard reports orders from AI channels. Now you have to decide what that number is worth: whether it earns more of your team's time, whether it goes in the board deck, and whether it survives someone asking how you know.
There is no clean answer to that yet. What there is, and what almost nobody lays out, is that the difficulty is not spread evenly. It depends on where the sale started.
Three tiers, and what separates them is not the click. It is whether anything records the sale on your behalf.
A product card is a commerce object, so the channel knows which product it was and the order reaches your admin already tagged. A citation in an answer is an ordinary web link, so all you get is a session that you have to connect to an order yourself, using signals that turn out to be uneven. And when an AI assistant names your brand and nobody clicks, there is nothing to connect.
This is a status report on all three, with the numbers we have and an honest account of the ones nobody has.
Current as of September 2026. This area moves quarterly, so treat every figure here as dated and verify the platform behavior in your own analytics.
Where the sale starts decides whether you can prove it
Sort AI commerce surfaces by how much of the record gets created for you, and the picture resolves immediately.

The two product card paths behave like an ordinary channel. Something arrives, you can see it arrive, and the order is yours to count.
The two mention paths behave like brand advertising, and until recently that comparison flattered brand advertising. A billboard at least reports impressions.
Google now reports the top of the funnel, on its own surfaces only
That has started to change, inside Google and nowhere else. Merchant Center now reports AI performance for conversational queries on AI Mode and AI Overviews. Since 31 August 2026, Search Console has reported generative AI performance separately, giving impressions from those same two features for every site worldwide, split by page, country, device and date.
What Merchant Center adds on top of a raw impression count is your share of voice against competitors, split across three shopping stages it calls Discovery, Evaluation and Ready to buy, plus the query intents behind each stage and the product attributes whose absence costs you appearances. Free, organic only, English only, five countries.

Google's own illustration of the report. The account and every figure in it are theirs and invented, so read the layout rather than the numbers. Its labels also run ahead of the current documentation, which names the third stage Ready to buy and the surfaces AI Mode and AI Overviews. We follow the documentation above.
Those are real counters, and they arrived sooner than most people expected. They are also the only ones. ChatGPT, Claude, Perplexity, and Copilot report nothing comparable, and ChatGPT is the large majority of the AI referral traffic we see.
Then notice what neither counter does. Between them they tell you that you appeared, how often, and roughly how that compares to competitors. Neither connects an appearance to an order. The brand advertising comparison holds after all, one altitude up: on Google's surfaces you now have the billboard's impression count, and you still cannot tie it to a sale.
The visible half is the bottom of the funnel, and most tools sell the top
This is the structural fact worth carrying out of the section, because it explains something about the tooling market. The visible half of AI commerce is the bottom of the funnel, and the invisible half is the top.
A GEO or AEO tool, generative engine optimization or answer engine optimization, sells you the top, scored against a number it produced itself. On Google's surfaces it now competes with a free report built from Google's own impression data.
What decides the bottom is the context attached to your products, and that part shows up in orders.
Paid AI placements come with numbers, organic mentions do not
The second split is orthogonal to the first, and it is wider than most people assume.
A paid placement inside an AI assistant behaves like advertising because it is advertising. It reports impressions, it reports spend, and the platform has a commercial reason to give you both.
An organic mention reports neither. Same surface, same shopper, no counters.
So a brand can be measurable and unmeasurable in the same product at the same time, depending on which of the two paths the shopper took.
Paid measurability tells you nothing about organic standing
You cannot compare the two paths on the same terms. Paid will always look more accountable, because it is the only one being counted.
The inference people draw next is the expensive part. Having bought visibility and measured it, they conclude something about their organic standing.
We went through 3,602 ChatGPT ad placements to test that inference. Paid advertisers appeared in the answer text 8% of the time. Holding the question constant across 91 matched pairs, ad on against ad off, paying moved the naming rate by −0.3 percentage points.
Two brands make the point faster than the averages do. Zoom bought no placements and was named 101 times. Mercari bought 66 and was named zero.
The ad slot and the recommendation slot are separate systems. Paid performance tells you nothing about whether you get recommended when nobody paid.

Limits: that is a Q1 2026 baseline, collected 8 March to 12 April. ChatGPT Ads has expanded and changed formats since, so read it as the baseline the field can be measured against rather than a current snapshot.
What your analytics can see of AI traffic today, and what slips past it
The middle tier is the click that arrives from a citation. Estimation is possible there, within limits.
ChatGPT tags those links with a UTM, widened in June 2025 to cover more surfaces. That is why ChatGPT is the one AI source most properties can already see in a report.
The others do not behave the same way, and published guidance says so only in the vaguest terms.
Two different signals are in play, and they fail independently.
The first is a UTM, which the AI assistant writes into the link. Your URL arrives with ?utm_source=chatgpt.com stuck on the end, and you can see it sitting in the address bar.
The second is the page the visitor came from, which nobody writes and nobody sees. The browser quietly tells your site which page they were on when they clicked, and GA4 files that visit as chatgpt.com / referral. Analytics tools call the signal itself the referrer.
A UTM wins when both arrive, so the page only matters when the UTM is missing. Lose both and GA4 has nothing to go on, filing the visit as Direct, the same bucket as somebody typing your address from memory.
Now the part that matters. A URL copied out of a chat keeps the UTM and says nothing about where it came from. A link clicked on an ordinary web page says where it came from whether or not anybody tagged it. So your analytics can catch one signal and miss the other, in either direction.
In Nile first-party data across 100+ Shopify stores over the 90 days to mid-September 2026, ChatGPT tagged 99% of its sessions with a UTM but named the page it came from on only 23%. Gemini inverted that: it named the page on 95% and tagged 43%. Claude ran Gemini's shape, Copilot ran ChatGPT's, and Perplexity sat between the camps.

The practical consequence is that no single signal sees the whole picture. Key your AI detection on the UTM and you resolve ChatGPT cleanly while losing more than half of Gemini. Key it on the page they came from and the loss runs the other way.
Limits: ChatGPT is the large majority of that sample, so any figure averaged across AI assistants is mostly a ChatGPT measurement wearing a broader label. Our Claude sample is thin and our You.com sample was too small to report. And the denominator is the sessions we could identify as that AI assistant, which we do from the same two signals being measured, so a visit arriving with neither is not in it. Read these as the composition of the AI traffic a merchant can see rather than of everything an AI assistant sends.
Search Console may be showing ChatGPT fan-out queries, unverified
A second route is circulating and worth knowing about at the confidence level it deserves.
Lily Ray has reported seeing what look like ChatGPT fan-out queries inside Google Search Console, identified by a filter that catches queries containing both a site operator and the word official.
To look yourself: Performance, Search results, Add filter, Query, Custom (regex), Matches regex, then (?i)(site:.*official|official.*site:) over 12 months.
The signature is the part worth seeing. On the property she showed, that filter returned thousands of impressions at an average position of 1.1 and not one click. Ranking first and never being clicked is not how people behave, which is what makes a machine the plausible reader.
She is careful to call it a pattern rather than a finding, and no platform documentation stands behind it. We have not verified it on a merchant property, so treat it as something to test on your own data rather than a measurement you can report.
It belongs in a status report because the state of this field includes what practitioners are trying, not only what has been settled.
Counted two ways, the same AI orders move by half again
One more number from the same window.
We counted the same three months of orders twice. Once by last touch, crediting AI only when an AI assistant referred the session the order completed in. Once by any touch, crediting AI when an AI assistant appeared anywhere in that buyer's previous 30 days.
Any-touch counting returns 1.5 times the orders that last-touch counting returns. Last touch, the default in almost every dashboard, finds two thirds of the AI-involved orders that leave a trace at all.
The usual worry runs the wrong way here. The common fear is that AI channels take credit they did not earn, and the comparison points the other direction.
That last clause is load-bearing. Any touch is not the true total either. It only sees buyers who clicked something at some point, which is why the next section exists.
Neither number is incrementality. Whether those orders would have happened anyway is a separate question, and no amount of path data answers it. On the narrower question of involvement, both counts are floors, and the more generous one is about 50% higher than the one you are probably reporting.
Nobody has measured the shoppers who never click an AI answer
Then there is the tier where estimation stops working.
A shopper asks an AI assistant for a recommendation. It names your brand. They do not click the citation. They search your name, or type your address, or come back on another device on Thursday.
The AI created that demand and left no trace anywhere in your analytics. Not as a last touch, not as an earlier one.
Getting named without being clicked is the normal case rather than the edge case.
How large is that population? We went looking for a defensible figure and could not find one.
Three candidate figures for untracked AI demand, and why none of them ship.
| Candidate figure and source | Why it does not qualify |
|---|---|
| 70.6% of AI traffic lands in Direct. Loamly, an attribution vendor. Published November 2025, updated February 2026 | Mostly inferred, not verified. Of the ChatGPT visits behind it, 100 out of 8,874 carried a verified signature. The rest came from methods the vendor itself rates at 65% to 90% |
| A citation click rate near 12%. Ouyang and Narechania at CHI. Published March 2026 | Wrong population. It is ChatGPT's own figure and the highest of the nine systems tested, from 12 undergraduates answering 30 questions on sports, politics, geography, science and literature. No shopping, and the authors call the study preliminary |
| Several times over. Our own post-purchase surveys, unpublished | Too small, and we are not neutral. The sample is below anything we would publish, and we are an interested party here as much as anyone else |
The mechanism is understood. Its magnitude is not measured. Anyone handing you a multiple for it is estimating, ourselves included.
How to report AI sales without inventing a number
The division of labor follows from the tiers, and it is simpler than the measurement literature makes it sound.
Report product card orders normally. They arrive carrying their own records. Count them, attribute them, put them in the deck.
Treat mentions nobody clicked like brand advertising. Set expectations that do not require click-level proof, and do not manufacture an incrementality figure for it. A brand channel without an impression counter is still a brand channel.
Give clicked mentions an interval, not a point. A holdout on matched product groups, control products chosen on price band and trend, or a stated pre-post each get you closer, and each carries known limits: products are not independent, control selection does a lot of silent work, and pre-post cannot separate anything else that changed in the window. Pick one, write the rule down before you look at the outcome, and report the range it produces.
Refuse to report a number without its method. A result stated with its window and its limits gets believed. A result stated alone gets discounted by anyone competent, and correctly.
Keep the finding that argues against you. If a product group went the wrong way, say so. That habit costs nothing and buys more trust than any rule on this list.
Where AI commerce attribution actually stands
AI commerce is roughly a year into being a real channel, and attribution is the open question of this phase. It is being worked on by platforms, by analytics vendors, and by the brands living with it, and nobody has a complete answer. The honest state of the art is a tiered one: prove the bottom, estimate the middle, and decline to invent the top.
We build the product context layer for e-commerce brands, and we track the path from click to purchase, which means we can tell you where an order came from and what your products looked like when the decision was made. We cannot tell you what would have happened without us. Nobody honestly can, and a vendor offering you a clean incrementality figure for their own channel is offering you a marketing asset.
We will keep publishing what we can measure, including the parts that do not flatter the channel.
If you want to start with the part that is provable, connecting your catalog to Nile takes a few minutes and costs nothing until it earns you a sale.
Related: if you are still working out which AI channels you are in and what you control, start with what you actually control in Shopify's agentic storefronts. The nine metrics worth tracking for agentic commerce start with incremental profit, and choosing target queries is how you measure discoverability without fooling yourself.
Frequently asked questions
Which AI commerce sales can I actually prove?
The ones that start in a product card, because a card is a commerce object and the channel records the order for you. It arrives in your admin already tagged, whether checkout happened inside the channel or on your own store. A citation click is different in kind rather than degree: nothing records it as commerce, so you hold a web session and have to connect it to an order yourself. A mention nobody clicked leaves nothing to connect.
Why is a mention in an AI answer harder to measure than a display ad?
Until recently there was no impression count anywhere. Google has since started reporting one for its own AI surfaces, so the gap is narrowing on Google and nowhere else. Even where the count exists it tells you that you appeared rather than what the appearance produced, which is the part a display ad can at least model.
Can I see how often my products appear in Google's AI results?
Yes, for Google's surfaces only, and now through two reports. Merchant Center's AI performance report covers conversational queries on AI Mode and AI Overviews, giving share of voice against competitors across three stages it calls Discovery, Evaluation and Ready to buy, the query intents behind each one, and the attributes whose absence costs you appearances. It is free, organic only, English only, and live in Australia, Canada, India, New Zealand and the United States. Search Console has reported generative AI impressions for those same two features since 31 August 2026, worldwide and for any site rather than a product feed. Both report appearances rather than orders, so they narrow the visibility gap and not the attribution one.
Does paid performance inside an AI assistant tell me how I am doing organically?
No, and assuming it does is a common and expensive mistake. Our own look at ChatGPT ads found the brands appearing in paid placements mostly do not appear in organic recommendations. The two run on separate mechanisms, so buying visibility and measuring it says nothing about whether you get recommended when nobody paid.
Do all AI assistants tag their referrals the same way?
No, and the difference is large. Across our own sessions in the 90 days to mid-September 2026, ChatGPT tagged 99% of its sessions with a UTM but named the page it came from on only 23%, while Gemini was the reverse at 43% and 95%. Claude behaved like Gemini and Copilot like ChatGPT. Whichever single signal your analytics keys on, it undercounts a different set of AI assistants. Those shares are of sessions we could identify as that AI assistant, and the identification uses the same two signals, so a visit carrying neither never enters the denominator.
How much AI-driven demand goes completely untracked?
Nobody knows, and the figures in circulation do not survive checking. The mechanism is clear: a shopper given your brand name who never clicks the citation leaves no touchpoint in any analytics. The magnitude is unmeasured, our own surveys included, so treat any published multiple for it as an estimate rather than a measurement.
What is the minimum I should do?
Two things. Separate your reporting by tier, so product card orders are not averaged together with mention-driven traffic that you cannot verify. And record a written baseline of orders and revenue per product before you start any AI work, somewhere it cannot be quietly revised later. Those two cost an hour and preserve your ability to answer the question at all.
Sources
- Shopify agentic storefronts. Shopify Help Center, retrieved 2026-08-26.
- ChatGPT adds UTM parameters to more links for better analytics tracking. PPC Land, 2025-06-13, retrieved 2026-09-15.
- Lawrence Hitches, utm_source=chatgpt.com explained. 100 ecommerce brands, 21 months of analytics (July 2024 to March 2026). Agency-side analysis, retrieved 2026-09-15.
- The AI Traffic Attribution Crisis. Loamly, published 2025-11-03, updated February 2026, retrieved 2026-09-15. Vendor data: Loamly sells AI traffic attribution. Cited as an example of precision exceeding method accuracy, not as a load-bearing statistic.
- Does ChatGPT Send Traffic to Websites. Loamly, same vendor and dataset as the previous entry, retrieved 2026-09-15. Source of the 100-of-8,874 verified-signature breakdown.
- Jianheng Ouyang and Arpit Narechania, Analyzing the Presentation, Content, and Utilization of References in LLM-powered Conversational AI Systems. CHI EA '26, arXiv:2604.15326, retrieved 2026-09-15. Preliminary user study, 12 participants, 30 non-commerci
- Agentic Commerce, key concepts. OpenAI Developers documentation, retrieved 2026-08-26.
- Nile first-party attribution data, 100+ Shopify stores, 90 days to mid-September 2026. Query text and full results kept at nile-content/research/incrementality-2026-09.
- AI performance insights in Merchant Center. Google Merchant Center Help, retrieved 2026-09-15.
- Generative AI performance report (Search). Google Search Console Help, retrieved 2026-09-16. Rolled out to all sites worldwide as of 2026-08-31. Reports impressions from AI Overviews and AI Mode by page, country, device and date.
- Lily Ray, post reporting suspected ChatGPT fan-out queries in Google Search Console. LinkedIn, retrieved 2026-09-15. Practitioner observation, hedged by its author, with no platform documentation behind it. Cited as a hypothesis rather than a method.