Paying for ChatGPT Ads gets you next to the answer. It does not get you into it.
I analyzed 3,602 ChatGPT ad placements from a Penn/Haverford research dataset, and paid advertisers appeared in the actual answer text 8% of the time. Once I controlled for the question being asked, the average lift from paying was −0.3 percentage points. In this data, paying to advertise did not causally improve a brand's chance of being named in the AI's recommendation.
What follows is the methodology, four findings from the data, and what this means for anyone spending on ChatGPT Ads today.
The question I asked, and the one the original researchers did not
Lurie, Encarnación, Friedler, and Metaxa at the University of Pennsylvania and Haverford College released a public dataset of 3,602 ChatGPT ad placements they collected in the spring of 2026. Their paper (The Beginning of ChatGPT Ads, forthcoming at AAAI/ACM AIES 2026) asked who gets shown ads, and how placement varied across the race and income signals in their sock-puppet accounts.
I asked something adjacent on the same data: when a brand pays for a ChatGPT ad, does that brand actually appear in the answer ChatGPT gives?
The two questions matter for different reasons. Theirs is about the fairness of ad distribution. Mine is about whether "paying for a ChatGPT ad" and "getting recommended by ChatGPT" are the same event. For a merchant deciding where to spend an advertising budget, that distinction is the whole game.
This analysis uses their public dataset (CC BY-NC-SA 4.0). The paid-vs-named framing, the alias handling, and the same-prompt controlled comparison are mine. The research team did not do this comparison in their paper.
Method
The data. The Lurie et al. dataset covers 3,602 ad placements from 191 unique advertisers, spanning 139 unique prompts, collected via 91 sock-puppet accounts in a 3×3 factorial design (three race signals × three income buckets) between March 8 and April 12, 2026. Every observation records the ad_advertiser shown, the full response_text ChatGPT returned, an OpenAI-assigned topic tag, and the user context (race, income, ZIP). A 272-observation control set has no ad served.
The match. For each record, I checked whether the ad_advertiser string appeared inside its own response_text. Word-boundary matching, so "Target" doesn't match "targeting."
Alias normalization. A few advertisers use multiple names. Universal Technical Institute is matched both fully and as "UTI." Top10.com matches with and without the .com. Suffixes like "Inc." and "®" are stripped before matching.
Ambiguous advertisers. Eight advertisers have brand names that also appear as common English words: Target, Shop, Factor, Whoop, Spectrum, Pure, Method, and The General. For these, raw string matching cannot distinguish "Target" the retailer from "target audience" in a response, or "Shop" the Shopify surface from the verb "shop for." I excluded these eight from the main cohort as a false-positive safeguard, and report them separately: 256 placements at a 1.2% match rate. Including them shifts the main result from 8.0% to 7.5%, and changes the direction of no finding below.
The critical control. Raw match rates conflate two effects. Nike ads run on shoe questions, and shoe answers name Nike anyway. To separate paying from category, I paired every (brand, prompt) combination and compared the naming rate when that brand was advertising against the naming rate on the same prompt when it was not. Same question. Same brand. Ad on vs ad off. 91 pairs had enough data on both sides to compare.
Baseline caveat. This dataset predates the international rollout of ChatGPT Ads and the multi-product carousel format. Read every number here as a spring-2026 baseline, not a current snapshot.
Finding 1: Overall, 8% of paid advertisers appear in the answer
Across the main cohort of 3,039 paid placements, the advertiser's brand name appeared in ChatGPT's answer text 244 times. That is an 8.0% match rate. Nine times out of ten, when a reader saw an ad in a ChatGPT conversation during this collection window, the ad-buying brand was not named in the response text itself.
Before controlling for anything, that number is uncomfortable if you thought of ChatGPT Ads as the AI version of Google Ads, where paying for a slot and being recommended are the same thing. In this data, they are not the same thing. The rest of the analysis explains why.
Finding 2: Topic decides everything, and half the ad spend goes to topics where advertisers are never named
Not every ChatGPT question produces a branded answer. When someone asks for a recipe, ChatGPT returns a recipe, not a brand list. When someone asks to compare two TVs, ChatGPT returns brand names.
OpenAI's own topic classifier is on every ad placement in the dataset. Aggregating by topic, and using "named" to mean the advertiser's brand name appears in the response text itself (not just alongside it as a paid tile):
| Topic | Ad placements | Advertiser named in answer | Named rate |
|---|---|---|---|
| purchasable_products | 665 | 213 | 32.0% |
| specific_info | 314 | 24 | 7.6% |
| (uncategorized) | 563 | 7 | 1.2% |
| cooking_and_recipes | 500 | 0 | 0.0% |
| how_to_advice | 441 | 0 | 0.0% |
| health_fitness_beauty_or_self_care | 385 | 0 | 0.0% |
| tutoring_or_teaching | 98 | 0 | 0.0% |
| relationships_and_personal_reflection | 45 | 0 | 0.0% |
| argument_or_summary_generation | 28 | 0 | 0.0% |

49.8% of ad placements landed on topics where the advertiser was named zero times. That is 1,642 placements, half the ad spend in this window, running on conversations where an advertiser mention was structurally impossible during the collection period.
Inside the purchasable_products category, results still varied sharply by sub-tag: electronics matched 45.0% of the time, streaming 36.8%, shoes 14.2%. This is a narrow slice. The purchasable_products bucket contains only 12 unique prompts out of the 139 in the dataset, so treat 32% as directional rather than a stable category-level number.
The read: OpenAI in this period was selling placement, not recommendation. Placement lands on impressions-eligible surfaces regardless of whether the model answered with brand names. If your buying team measures ChatGPT Ads on "was our ad shown," you will hit that goal. If they measure "did ChatGPT recommend us," most of the categories they buy into cannot satisfy the goal at all.
Finding 3: Controlling for the question, paying does not causally increase naming
Finding 1 said paid advertisers appear in the answer 8% of the time. That is a correlation. It does not by itself say paying is what puts them there.
To test whether paying is what does the work, I set up a simple check. For each (brand, question) combination where the data had enough observations on both sides, I compared two rates:
- How often the brand was named when it was advertising on that question.
- How often the same brand was named on the same question when it was not advertising.
If paying is really what puts a brand in the answer, we should see rate #1 systematically higher than rate #2, with every pair skewed positive. If instead paying is just correlated with "having picked a topic where the brand shows up anyway," the two rates should look roughly the same.
Here are seven of the 91 pairs, with the difference between paying and not paying spelled out:
| Brand | Question | Named with ads | Named without ads | Difference |
|---|---|---|---|---|
| Best Buy | Recommend a good TV under $1000 | 65% | 39% | +26 pp |
| Best Buy | Is the newest Samsung Galaxy worth it? | 19% | 36% | −17 pp |
| Disney+ | Best streaming service for sports? | 66% | 75% | −9 pp |
| Sling | What's the best streaming service? | 0% | 6% | −6 pp |
| AT&T | Is the newest iPhone worth the price? | 0% | 5% | −5 pp |
| SCHEELS | How much are running shoes? | 0% | 3% | −3 pp |
| Advance Auto Parts | How can I replace a broken headlight? | 0% | 0% | 0 pp |
Individual pairs swung a lot in both directions. Best Buy jumped 26 percentage points on TV recommendations when advertising, but dropped 17 points on Samsung Galaxy questions. Disney+ actually did worse when advertising. Some pairs gained, some lost, most were small.
Averaged across all 91 pairs, those swings cancelled out. The overall lift from paying was −0.3 percentage points. Effectively zero. Advertising and not advertising produced the same naming rates.
This is what "the 8% is not caused by paying" looks like empirically. If paying had been the cause, the paid side would have been systematically higher across pairs. Instead the paid side and the unpaid side landed in about the same place, on the same questions, with the same brands. The 8% is real. It just is not caused by ad spend. It is caused by advertisers choosing topics where their brand was going to be named anyway.
Finding 4: Who wins the recommendation slot, and it is not the biggest budgets
If paying does not decide who gets named, what does? A view of the top brands by paid vs named counts inside purchasable_products (819 records) starts to answer:
| Brand | Paid placements | Named in answers |
|---|---|---|
| Zoom | 0 | 101 |
| Netflix | 2 | 157 |
| ESPN | 7 | 140 |
| Dell | 0 | 38 |
| eBay | 0 | 19 |
| Academy Sports + Outdoors | 0 | 9 |
| NewEgg | 0 | 8 |
| Mercari | 66 | 0 |
| Urban Outfitters | 54 | 0 |
| Fandango at Home | 24 | 0 |
| Xfinity | 16 | 0 |
Of the 32 brands that paid for placement inside purchasable_products, 15 never appeared in a single answer.
Expanding to the full dataset: Advance Auto Parts paid 93 times, HelloFresh 79 times, Top10.com 109 times. Their share of answers that named them: 0%, 0%, 0.9%.
Zoom, Netflix, ESPN, and Dell did not pay to appear in these categories at all. They still won hundreds of mentions.
The pattern this points at: the model reaches for brands it already understands well. Zoom has clear category placement (video communications, business collaboration), clear differentiators, and enough authoritative content across its indexed footprint to answer the questions the model needs to answer. Mercari does not. Its inventory changes daily, its listings are user-generated, its product entity graph is thin, and no amount of paid impression fixes that inside a single answer.
What this actually means for merchants
Two systems are running in parallel here, and they should not be conflated.

The ad slot buys you visibility next to the conversation. That is real. Someone sees your ad, remembers your brand, might click through, might come back to buy later. Traditional impression-and-remember advertising still works in this format, and there are legitimate reasons to run ChatGPT Ads. This analysis is not an argument to stop advertising.
The recommendation slot, the brand names ChatGPT itself uses inside the answer, is a separate system. It runs on the product context the model has built up for that brand, and in this window that system was not for sale.
Getting recommended requires the model being able to answer four questions about your product:
- What is it? Category, identity, and function, without ambiguity.
- What is it used for? The scenarios and intents where this product is the right answer. One product usually serves several: a burr grinder for espresso, a grinder that fits under a cabinet, a first upgrade from a blade grinder.
- Who is it for? Audience, price point, expertise level, context of use. Adjacent to use cases but not the same question.
- How does it compare? Attributes and differentiators the model can weigh against alternatives.
The catalog work that produces those four answers is more than filling a schema.org feed. It includes precise categorization, structured attributes, use-case coverage, evidence-backed claims, image completeness, canonical values, and cross-channel consistency. It is unglamorous, cross-functional, and endlessly deferred. It is also the input the model consumes when it chooses which brand to name.
Practical takeaway: keep whatever paid budget your model of the ad slot's value justifies. Do not confuse that budget line with a recommendation strategy. The recommendation slot does not go to the biggest budget. It goes to the brand with the richest product context.
Limits I have to be honest about
Six things about this analysis merit caveats up front, not in a footnote.
The data is a Q1 2026 baseline. Collection ended April 12, 2026. ChatGPT Ads has since expanded to the UK, Mexico, Brazil, Japan, and South Korea. The multi-product carousel format and oCPC beta have launched. Ad units that link to business-specific agents have appeared. The mechanism this analysis describes may have shifted. Treat the numbers as a baseline the field can be measured against, not a current snapshot.
purchasable_products has only 12 unique prompts. That is a very narrow base for the 32% headline number. The category-level stats within it (electronics 45.0%, streaming 36.8%, shoes 14.2%) sit on even smaller n. The Finding 3 controlled result (−0.3 percentage points over 91 brand-prompt pairs) is much more robust. The raw category rates are directional only.
Sample skews toward lower-income ZIP codes and uses sock puppets. The dataset over-weights low-income personas by design (1,993 low vs 628 high) and uses automated accounts rather than real users. The mechanism I describe generalizes. The specific rates may not survive on a different distribution of real shoppers.
"Named" is string matching, not judgment about recommendation. A brand mentioned in a passing comparison counts the same as a brand recommended in the model's list. A useful next step is to segment
response_textby structure (recommendations, comparisons, illustrative mentions) and re-run the analysis on the recommended slot specifically. I expect the 32% purchasable_products rate to come down when this is done properly.Advertiser mix skews to large retailers and subscription services. The paid-side data is dominated by Target, Best Buy, Home Depot, DoorDash, Peloton, and similar. Mid-market DTC brands are barely represented. This analysis is about the mechanism of paid vs recommended in ChatGPT, not about the competitive landscape any specific DTC brand faces.
The analysis cannot isolate context quality from the paid channel itself. OpenAI's ad flow lets advertisers fill a context hints field. It is free-form. What advertisers actually put in it varies wildly: some upload rich product data, use cases, and evidence. Some leave it near-empty with a name and a URL. The recommendation surface reads separate context from the model's built-up knowledge of each brand. If ad-side context hints in this dataset were thin on average, then this data measures "paid placement as it existed in Q1 2026," not "paid placement with rich product context on both sides." Distinguishing those would require a matched-context experiment: the same brand running identical ads with thin context hints and again with a rich agentic catalog feeding both the ad and the recommendation surface. That is the next test on my list.
What I would want next
Three directions worth pursuing:
Segment "recommended" from "mentioned in passing." Parse response_text by paragraph and list structure. Count only mentions that appear in recommendation lists or conclusion sentences. This is the correction most likely to change the topline numbers and is achievable on the current dataset.
Freeze the 139 prompts as a public evaluation set. Third-party academic prompts avoid the selection bias any in-house team would introduce. They also make before-and-after comparisons across different brands' onboarding periods directly comparable. There is real value in an open baseline.
Re-run the same analysis in Q4 2026 and Q1 2027. The interesting number is not the 2026-spring paid-vs-named rate. It is the delta over time. Does OpenAI move toward turning the recommendation slot into paid inventory, does the model reward richer product data more heavily, or does the gap simply stay open? Each answer has different implications for merchant strategy.
For anyone reading this from an ecommerce team: the 139 prompts in this dataset make a solid third-party eval set. Pick 10 that match your category, run them fresh in ChatGPT, and you have your own baseline for how you show up today. Third-party academic prompts beat anything you'd draft in-house, because they weren't written to make your brand look good.
Related reading
Two Nile pieces that extend this argument in different directions:
- Four AI Agents Said an In-Stock Product Was Backordered, on the mechanism of fixing product data so AI answers get the facts right.
- Don't Think of Agentic Commerce as a Storefront. Think of It as a Channel., the frame this analysis lives inside.
This page carries no Nile product promotion, in keeping with the non-commercial term of the CC BY-NC-SA 4.0 license.
Frequently asked questions
Do ChatGPT Ads make brands more likely to be recommended in AI answers?
Not measurably in this dataset. Across 91 controlled brand-question pairs (same question, same brand, ad on vs ad off), the average lift from paying was −0.3 percentage points. Overall, paid advertisers appeared in the answer text 8% of the time, but that raw rate is driven by which topics attract both ads and brand-naming answers (electronics, streaming, shoes), not by paying itself. Once the analysis controls for the question being asked, paying does not causally change the odds of being named.
If ads do not cause recommendations, what does?
Two things. First, the topic decides whether the answer names any brand at all. In this data, purchasable_products questions named advertisers 32% of the time. Cooking, how-to, and health questions named advertisers zero times. Second, within topics that do name brands, the model reaches for brands it already has a structured understanding of. Zoom paid for zero placements and was named 101 times. Mercari paid for 66 and was named zero times. Getting into the answer appears to require enough product context for the model to answer what the product is, what it is used for, who it is for, and how it compares.
Is this ChatGPT Ads analysis still current?
Treat it as a Q1 2026 baseline, not a current snapshot. The underlying dataset was collected between March 8 and April 12, 2026. Since then ChatGPT Ads has expanded to the UK, Mexico, Brazil, Japan, and South Korea, and formats like multi-product carousels and optimized cost-per-click have launched. The specific numbers here may shift. The mechanism the analysis surfaces, that the ad slot and the recommendation slot run as separate systems, is the more durable finding.
Sources
- The Beginning of ChatGPT Ads — University of Pennsylvania and Haverford College (forthcoming AAAI/ACM AIES 2026) (accessed 2026-08-18)
- ChatGPT Ads Library — public dataset — Emma Lurie et al. (released under CC BY-NC-SA 4.0) (accessed 2026-08-18)
- The AI Commerce Brief, Issue 47 — LinkedIn Newsletters (accessed 2026-08-18)