The assistants split between Netflix and Hulu, each named in about 42% of these queries: no single winner.
Mapou Research · Issue 002 · Streaming · 28 July 2026
What AI is telling your subscribers to cancel.
Your subscribers are already asking AI whether to keep paying you, and what to cancel first. We put the exact questions they ask, "is it still worth it," "what should I drop," "which bundle is cheaper," to ChatGPT, Gemini, Claude, and Grok, then measured what each one recommended, the reason it gave, and the price it quoted. 443 measurements across the four engines, plus 20 full buyer conversations.
Three of the findings below should change what your retention and pricing teams ship this quarter. Each one comes with the specific move it implies, and who owns it. Start at the top.
Money is what AI weighs above everything else: 41% of every reason it gives.
Across all four engines, the sticker price and bundle savings together drove 41% of every streaming recommendation; price alone was 29%, the single biggest reason by far. Catalog depth, original shows, the live-sports and cultural moments streamers spend billions competing on all sit well behind. Across this study, AI competed largely on cost.
WHAT TO DOWhatever else you invest in, AI funnels the keep-or-cancel decision through price. The next finding shows what happens when the price it quotes is wrong.
Read the full finding ↗It inflates the price of the exact services people are quitting. Hulu gets quoted at $20; it's $18.99.
The market leaders are quoted accurately: Netflix, Disney+, and Paramount+ all sit near zero drift. The errors cluster on the cancel-bubble services. Hulu's ad-free tier is cited at a median $20 against an actual $18.99, a 5% overstatement. ESPN+ runs +30%, HBO Max Premium +17%, Prime Video +27%. When AI overstates your price, it manufactures a cancel reason that was never real.
WHAT TO DOGround-truth the prices on the surfaces AI reads, then re-measure monthly. A subscriber pricing the keep-or-cancel call against a number that's half-again too high is being pushed out by bad data, not by your actual value.
Read the full finding ↗AI's number-one source for streaming advice is YouTube: 8% of every citation.
We pulled every source the four engines cited across the study. The single most-cited domain was YouTube at 8%, ahead of any streamer's own page or review site. The streamers' own help and pricing pages were 14.9% combined. So the place AI forms its keep-or-cancel opinion is mostly creator reviews and comparison videos, and the content behind those citations was about four months old by publish date.
WHAT TO DOThe leverage is upstream of your own site. Find the creator reviews and comparison videos AI actually cites for your category, and check which ones carry your current price and lineup. That is where the recommendation gets shaped.
Read the full finding ↗TWO MORE FOR THE ANALYSTAI is weakest exactly at the cancel moment: it grounds (searches the live web) on 100% of "what's trending" queries but only 34% of cancel-and-keep queries (finding 01). And when a real subscriber talked their stack through with AI turn by turn, 50% of conversations ended in cancel or pause (finding 02). AI's default posture in the retention conversation is to help people leave.
Cadence
Monthly · v2 ships Oct 2026 (UK / FR / AU)
Method
50 prompts · 4 engines · 5 personas
Confidence
Wilson 95% on every proportion
Read on
Full report ↘Grounding rate by engine
The gap is on cancel and keep questions, not on what just dropped.
Grounding, in plain English
When an assistant grounds, it stops and searches the live web before answering, instead of replying from memory, its training data, which can be a year or more out of date. A grounded answer is built from pages that exist right now, pages you can change. An ungrounded answer just repeats whatever the model already believed. That is why grounding is the whole game: the more an engine grounds on your category, the more your current pages, prices, and sources decide what it tells a subscriber.
ChatGPT grounded 84% of queries against live web sources. Gemini, Claude, and Grok grounded 85 to 100%on the same prompts. We expected the gap to be wider on time-sensitive queries ("biggest show right now," "what just dropped") because that is where stale training data hurts most. On those, ChatGPT grounded 100% of the time. The gap showed up almost entirely on cancel and keep queries ("should I cancel Netflix this month," "is HBO Max worth keeping"), where pricing and current lineup are just as load-bearing. These rates also come through our standardized retrieval harness, one shared search layer across all four engines; on its native consumer search the same model can ground more or less often, so the cross-engine gap and the monthly trend are the durable read, not any single rate.
ChatGPT only · split by query type
Time-sensitive queries (n=9)
100%
grounded
Cancel, keep, and bundle queries (n=44)
34%
grounded
All engines · overall grounding
CHATGPT
85%
86 of 101 prompts grounded
95% CI: 77–91%
GEMINI
100%
100 of 100 prompts grounded
95% CI: 96–100%
CLAUDE
90%
91 of 101 prompts grounded
95% CI: 83–95%
GROK
84%
84 of 100 prompts grounded
95% CI: 76–90%
Hypothesis · not finding
Two readings consistent with the data. First: ChatGPT's router may classify time-tagged queries ("right now," "this week") as factual lookups that need search, while reading cancel and keep queries as opinion or recommendation tasks where the model answers from training data. Second: web-search invocation is gated on a confidence threshold, and decision queries hit that threshold less often because they don't contain time tokens. Both are testable. If the gap closes when identical decision queries are rephrased with explicit time anchors ("at today's price"), the second reading is closer to right.
What you can do about it · with mapou
When a subscriber asks AI whether to keep paying, ChatGPT is least likely to check the live web. The one moment your current price and lineup matter most is the moment AI is most likely to answer from stale memory.
We split your cancel-and-keep visibility from your what's-new visibility and show which queries AI answers without looking anything up. Then we name the live, structured surfaces you own that pull the engine back onto the web at that moment: current-offer pages, help-center pricing, FAQ schema. Those are yours to change.
The exact cancel-and-keep queries where AI isn't checking the live web, and the owned pages to fix so it starts.
mapou services
How subscribers decide
When subscribers talk it through, most walk away from something.
Five personas, one conversation against each engine, 20 multi-turn dialogues in total. We let each persona press the AI for specifics and then recorded where the conversation landed: Kept, Canceled, Paused, Skipped, or Inconclusive. 50% ended in Cancel or Pause: the subscriber walked away from at least one service they were paying for. Only 45% ended in a clean Keep. In these conversations, left to its own framing, AI was at least as much a churn accelerant as a retention aid.
One mechanism worth noting: the engines ground far more reliably inside these conversations than on cold queries. Across the 20 dialogues the four engines grounded 96% of turns, against ChatGPT's 84% on a cold single-shot query (finding 01). The presence of a real buyer asking real follow-ups is what pulls AI onto the live web. The cell below shows where each persona-engine conversation landed; cancel, pause, and skipped share iron because all three end in a walk-away.
| PERSONA | CHATGPT | GEMINI | CLAUDE | GROK |
|---|---|---|---|---|
Marcus, 31 · sports-stacker, currently churning | Canceled (no service named) 5 turns | Canceled Apple TV+ 3 turns | Paused Peacock 3 turns | Canceled HBO Max + Apple TV+ 5 turns |
Priya & Sam, late-30s couple · prestige-drama, two-service ceiling | Kept Netflix 5 turns | Kept HBO Max 5 turns | Canceled HBO Max 7 turns | Kept Netflix + HBO Max 5 turns |
Karen, 42 · parents of two, Disney-locked | Canceled Netflix 3 turns | Canceled Netflix 5 turns | Paused Netflix 3 turns | Paused Netflix 3 turns |
Liam, 26 · ad-tier price-shopper, churn-friendly | Kept Paramount+ 3 turns | Kept (no service named) 3 turns | Kept Paramount+ 3 turns | Kept Paramount+ 3 turns |
Dave, 58 · recent cord-cutter, live-TV replacement | Canceled Peacock 7 turns | Inconclusive 0 turns | Kept (no service named) 5 turns | Kept YouTube TV + Disney+ + Hulu 9 turns |
Kept
9
45% of conversations
Canceled
7
35% of conversations
Paused
3
15% of conversations
Skipped
0
0% of conversations
Inconclusive
1
5% of conversations
AI's default keep-or-drop hierarchy
Asked what to keep and what to cut, AI has favorites.
We judged every cancellation-prompt response for which services it steered the user to keep versus drop, across 84 responses. The bar shows drop-steers (left) against keep-steers (right) per service. Netflix is the most polarizing: the most-recommended to keep (13) and the most-recommended to cut (9) at once, the price anchor everyone weighs.
YouTube TV sits at the bottom because AI tends to cut the live-TV bill first, before any on-demand service. Live-TV bundles are a separate category from the on-demand services this study measures; they get their own report.
Where AI sends a cancel-ready subscriber
Across 1011sources cited on these prompts, the destinations are YouTube reviews, tech-press explainers, and third-party "discount when you cancel" aggregators. The streamers' own pages that surface are help-to-cancel and sign-up pages. Not one save or win-back offer page appears. When a subscriber asks AI how to get a deal to stay, a listicle answers, not you: at the cancel moment the engine is least likely to be reading your live pages at all, so it answers from stale memory instead of your current retention offer.
Directional: per-service counts run single digits to ~16 across 84 responses. The signal is the order, not the exact gap.
What you can do about it · with mapou
When a subscriber takes the keep-or-cancel question to AI, the modal outcome is to drop a service. Which personas walk, on which engine, and on what stated reason, is the churn signal that doesn't show up in your own funnel until the cancel already happened.
We re-run these buyer conversations monthly, flag which personas AI walks toward Cancel or Pause and the reason it gives, then trace that reason back to the source AI read it from. When the reason is a wrong price or a stale lineup, that is a fixable input on a page you control, not a lost customer.
A persona-level map of where AI talks your subscribers out of staying, with each losing conversation traced to the input that caused it, so you know what to fix rather than just that you're losing.
mapou services
Why it matters
In these conversations, AI talked subscribers out of a service more often than into one. What it told them was built on facts that were months out of date.
Sources, freshness & stale facts
One in five citations is a YouTube video.
of every source AI cited was YouTube, ahead of any streamer's own page or any review site.
6,510 citations · 802 domains · July 2026
We pulled every citation the four engines attached to a streaming answer and ranked them by who actually gets cited. The most-cited source by a wide margin is a video platform, not a streamer and not a reviewer. On the cancel-and-keep queries in this study, the citations behind those answers leaned hardest on creator reviews and comparison videos.
The streamers' own help and pricing pages account for 14.9% combined. Editorial tech press (Business Insider, Tom's Guide, TechRadar, TV Guide) is another 21%. And 1.2% of citations point at gray-market subscription resellers, sites that sell shared or region-arbitraged logins, which AI surfaces as if they were legitimate ways to pay.
Most-cited domains · share of all 6,510 attributed citations
The actual hostnames AI cited, ranked. This is the answer to "where does AI get its streaming information," and where a retention or distribution team would place to be in the room when the recommendation is formed.
What kind of source · every citation bucketed
How old is the content behind those citations
84d
median age by publish date, when the page was first written
-37d
median age by last-updated, when the publisher last touched it
12%
stale-but-refreshed: written over a year ago, re-touched in the last six months
21%
of dated citations were first published more than 12 months ago
The gap between 84d and -37d is the whole story. The content AI leans on was written about four months ago, but the pages were re-touched roughly a week and a half ago. To a freshness signal it reads as current. The pricing, lineup, and naming facts inside it are as old as the original write. That is the same mechanism behind the wrong prices in finding 05, and behind the dead services and out-of-date names below: fresh metadata wrapped around stale facts.
The facts that go stale · the AI confusion tax
Services that no longer exist still get recommended 38 times.
When the content behind a recommendation is months old, the facts inside it are too. Showtime folded into Paramount+ in April 2024; Discovery+ folded into HBO Max in May 2023; CNN+ shut down in April 2022. All three are still recommendable, today, by an AI assistant answering "what should I watch." Same root cause as the wrong prices: stale source, fresh-looking page.
Showtime
31×
standalone folded into Paramount+, April 2024
Discovery+
7×
folded into HBO Max, May 2023
The same staleness shows up in naming. HBO Max was renamed back from "Max" in July 2025; across 140 responses AI used the current name only 44%of the time. ESPN+'s rebrand to ESPN Select, which had a high-profile launch, caught on far faster. Press coverage of the change, not the size of the service, decides how quickly AI catches up.
Max → HBO Max
Overall current-name share: 44%
Re-rebranded from "Max" back to "HBO Max" in July 2025
ESPN+ → ESPN Select
Overall current-name share: 94%
Renamed from "ESPN+" to "ESPN Select" alongside ESPN Unlimited launch in late 2025
A finding for trust & safety, not just marketing
78 citations (1.2% of the total) point at subscription-reseller and account-sharing sites. When a price-sensitive subscriber asks AI how to keep a service cheaply, one in thirty sources it surfaces is a gray-market login vendor. That is revenue leakage AI is actively routing toward, and a brand-safety exposure most retention teams aren't watching.
CHATGPT · top sources
GEMINI · top sources
CLAUDE · top sources
GROK · top sources
What you can do about it · with mapou
The sources AI cites are where viewers form the keep-or-cancel opinion before they ever open the app. Right now that's YouTube reviews, tech-press comparison pieces, and a tail of gray-market reseller sites, not your own pages.
Every month we map the domains AI actually pulled for your category and split them three ways: the pages you own (help center, pricing, FAQ), the sources you can influence (the creator videos and comparison sites it cites), and the ones to dispute (gray-market resellers). You get a ranked fix list: which of your pages to refresh, which creators and outlets to get current facts into, and what to flag to trust-and-safety.
A ranked, refreshed list of the sources moving AI's recommendations, sorted by what you own, what you can influence, and what to dispute, with publish-vs-update age attached.
mapou services
Why AI recommends what it recommends
Nearly half of every reason AI gives is about money.
of every reason AI gave for a recommendation was about cost: the sticker price or bundle savings.
764 reasons tagged · 4 engines · July 2026
For every recommendation, we recorded the reason the engine gave and sorted it into one of fifteen categories. The two biggest are both about cost: price & value (29%) and bundle savings (12%). They're distinct: price-and-value is whether the standalone sticker is worth it; bundle-savings is paying less by combining services. Both put cost ahead of content.
The next tier is about what there is to watch: catalog depth and live sports. Original shows, hit titles, and the cultural moment, the axes streamers actually compete on at the executive level, sit in the long tail. The full list of reasons, with what each one means, is below.
Price & value
29%Whether the service is worth what it costs. AI calls it cheap, fairly priced, or a good deal for the money, judging the sticker price on its own.
“At $7.99 a month, Paramount+ is the most affordable of the three.”
Bundle savings
12%Paying less by combining services rather than buying one outright. AI points to the Disney+/Hulu/ESPN bundle, a phone-plan perk, or a student deal instead of the standalone price.
“You can get Disney+, Hulu, and ESPN+ together for less than two of them separately.”
Live sports
10%Access to live games and sports rights. The one reason that most often overrides price in the data.
“If you want the NBA this season, ESPN is the one you need.”
How much there is to watch
10%The size or breadth of the library as a whole, not any one title. AI recommends on volume: more shows, more movies, something for everyone.
“Netflix still has the deepest overall catalog.”
Every reason, ranked (764 tagged across all four engines)
Top reasons per engine. All four lead with cost.
CHATGPT
GEMINI
CLAUDE
GROK
Hypothesis · not finding
The convergence on price is consistent with a citation-pool bias. Price-comparison content (reviews, cord-cutting blogs, plan-comparison listicles) is the most-indexed and most-linked corner of the streaming web; originals coverage is more fragmented across entertainment outlets, trades, and service-owned editorial. If engines weight by citation density, the recommendation defaults to the dimension with the densest source set. Testable by comparing the reason mix between queries that include genre-specific language ("best prestige drama," "best anime") and queries that don't. If genre-specific prompts shift the reason mix toward originals and hit-shows, the citation-pool reading is closer to right.
What you can do about it · with mapou
AI recommends on price first. If the price it has for you is wrong, or your value story never reaches the sources it reads, you are being judged on a number you didn't set and a pitch you didn't write.
We show the reasons AI gives for your service versus competitors, by engine, and flag where the gap is a fixable input: a price it has wrong, a bundle it misstates, a value point absent from the pages and reviews it cites. You get the specific surfaces to correct, ranked by how often that reason decides a recommendation.
A reason-by-reason map of where AI's pitch for your service runs on inputs you can fix, and which to fix first.
mapou services
The catch
If AI recommends mostly on cost, the cost it quotes has to be right. On the cancel-bubble services, it often isn't.
Pricing accuracy
The prices AI quotes are right two out of three times.
of the prices AI quoted were off by more than ten percent. That is the math a subscriber uses to decide whether to keep paying.
565 prices checked against official pages · July 2026
We asked the four engines to quote current monthly pricing for the eight major streaming services and compared every dollar amount they cited against a ground-truth table sourced from each service's official pricing page (or the Wikipedia article reflecting the most recent documented change).63% landed within ten percent of the real price. The errors cluster on the cancel-bubble services, and they mostly run high.
CHATGPT
61%
accurate · 82 prices cited
95% CI: 50–71%
GEMINI
60%
accurate · 149 prices cited
95% CI: 52–67%
CLAUDE
68%
accurate · 179 prices cited
95% CI: 60–74%
GROK
63%
accurate · 155 prices cited
95% CI: 55–70%
Accuracy at a glance · every tier, worst first
Each tile is one real tier on sale right now. Iron stripe = under 60% accurate. Cane = 60-75%. Moss = above 75%. The first row is the math a retention team should fix this quarter.
Per tier · every service expanded to its full ladder
Each row is a real tier in the market right now. A retention exec at Netflix is asking three different accuracy questions (one per tier), not one. The table is grouped by service; rows highlighted in iron are tiers where AI's median citation drifts more than $1 or where accuracy falls below 60%.
| SERVICE / TIER | ACTUAL | MEDIAN CITED | DRIFT | ACCURATE |
|---|
Missing from this run: ESPN+, ESPN flagship streaming, 4K-tier-explicit prompts (Netflix Premium and HBO Max Premium are the closest proxies). Adding ESPN+ is ~$0.50 of additional pipeline cost. Adding explicit 4K-tier prompts is the same.
What this means for a retention org
A Peacock subscriber asking AI "is it worth keeping" is making the call based on a price that was true two years ago.
Peacock Premium is currently $10.99/month. The median price the four engines cited across 50 mentions was $7.99, a drift of -27%. That difference is real money over a year. If a subscriber decides to stay because the AI told them the price was lower than it is, the first renewal bill is a trust event. If they decide to leave because the AI exaggerated a competitor's deal, the comparison was never apples-to-apples to begin with.
AI doesn't cause the cancel decision. It frames the math the subscriber uses to make it. The distinction matters for what we measure and what we don't claim.
Ground truth sources (July 2026)
What you can do about it · with mapou
Subscribers are deciding whether to stay or cancel using AI-quoted prices that drift up to a third off the real number. Retention math built on those quotes is wrong before the conversation ends.
We benchmark every cited price for your service and its closest competitors against ground truth, flag drift by engine and tier, and trace each wrong quote back to the page AI read it from, usually one you own or can correct. You fix the page; we re-measure on the same prompts next month and show the drift close.
An accuracy dashboard that tells your retention team which engines quote you right, which quote last year's number, and which pages to fix to move it, verified by next month's re-run.
mapou services
The throughline
None of this needs AI to be hostile. It only has to be out of date at the moment a subscriber asks whether the service is still worth keeping.
You don't own the engine. You own most of what it reads.
You can't change a model's training or its algorithm. You don't need to. On the questions that decide whether a subscriber keeps paying, these assistants don't answer from memory, they search the live web and assemble the answer from pages. Across the buyer conversations in this study they grounded 96% of the time. Change the pages, change the answer.
This is SEO's lesson, one layer up: you never owned Google's algorithm either, and optimizing what it surfaced was still the work. The loop below is that work for AI answers.
01
Find where AI looks
We separate the queries AI answers from the live web from the ones it answers from memory, and pull the exact sources it cited for your category.
AI Visibility Monitor · Source & Citation Map
02
Sort by what you control
Your own pages (pricing, help center, FAQ schema) you fix outright. The reviews, comparison sites, and creator videos AI cites you influence through PR and partnerships. Gray-market sources you dispute.
Source & Citation Map
03
Fix the inputs the findings flagged
Correct the price AI quotes. Update the name it still gets wrong. Get your current lineup and offers onto the surfaces it actually reads. Where you can't earn the placement, we buy it.
AI Content & Schema Studio · AI Media Planner
04
Re-measure, and prove it moved
Same prompts, same engines, next month. You see whether your pricing accuracy, recommendation share, and cancel-conversation outcomes moved. The before-and-after compounds month over month.
Tests & Lift
The honest limit: some of what AI believes is baked into training and only updates on the model's clock. The outdated service names in Finding 03 are an example. We measure that lag rather than promise to erase it. Everything else, the prices, the current offers, the sources, moves when you move the pages.
We don't invent a dollar number. Here is the math you plug your own assumptions into.
Asking "what does a one-point drop in AI recommendation share cost us" is the right CMO question. The wrong answer is to make up a percentage. Here is the framework. Pick your own assumption ranges. The output is a band, not a point estimate.
annualized revenue at risk =
(RSOV gap, percentage points)
× (AI share of category discovery)
× (annual revenue per subscriber at risk)
× (subscribers in the at-risk decision cohort)
RSOV GAP
Your recommendation share against the category leaders, measured across the four engines.
AI SHARE OF DISCOVERY
Conservative 10% (industry analyst floor). Aggressive 30% (Gartner 2028 projection). Use your own panel data if you have it.
REVENUE PER SUB
Your ARPU × retained months. Adjust for ad-tier vs. premium mix.
AT-RISK COHORT
Subscribers entering the keep/cancel decision window in the period (typically renewal anniversaries plus price-change cohorts).
We will not publish a fabricated dollar headline because the underlying assumptions are yours, not ours. We measure the RSOV gap, the variance band, and the pricing-drift exposure to the dollar. The rest is your finance team's number.
AI is now part of the keep, cancel, pause, and stack decision. The numbers are wrong often enough to matter.
A subscriber asking ChatGPT or Gemini "is this worth keeping" gets back a recommendation built on whatever pricing, brand name, and competitive set the engine believes is current. This report shows that 36% of the pricing AI quotes is materially wrong, that AI uses your current brand name only 43% of the time after a recent rebrand, and that the variance hour-to-hour on the same query is 38%. Those numbers are the math the cancel decision runs on.
If your retention model assumes AI is showing subscribers your real price and your current name, the model is mispriced. The size of that mispricing is what this report measures, for your service, every month.
AI Retention Intelligence · for streamers
The layer that influences keep, cancel, pause, and stack decisions. Built for retention and subscriber strategy teams at streamers.
- ·Pricing-drift monitoring against your current published rate, refreshed weekly, with per-engine accuracy bands.
- ·Cancel-and-keep query coverage. Where each engine sends subscribers asking whether to drop your service, and what alternatives the engines surface.
- ·Churn risk by audience. How a sports-stacker, a parent household, and a price-shopper hear about your service differently when they ask the same question.
- ·Rebrand checks. How often each engine cites your current name vs. a past name, and where that risk lands on your service tree.
- ·Monthly check-ins scoped to your retention KPIs. Variance bands and month-over-month deltas, not single snapshots.
Also working with content studios & production companies
A parallel engagement focuses on title-level visibility across the four engines: whether your titles surface where they live, the framing AI gives them against competitors, and the editorial moves that change the frame. Same methodology, different KPI set.
Book a 30-min title audit →Methodology v2 · before / after
What changed when the AI models upgraded.
This month the measured engines moved to a new model generation. To tell a real category shift from a model change, we re-ran June 2026's models on a frozen subset of the same streaming prompts the same day. Held constant, the category barely moved. What did move was the models themselves.
The clearest effect is grounding, how often an engine searches the live web before answering. On the same prompts, the upgrade changed it sharply, and in opposite directions by engine.
The upgrade also made answers more concentrated: 10 of the 16 tracked streaming brands were named less often by the new models than the old ones on the same prompts. Newer models give shorter, more decisive answers, so a few names hold, and the long tail thins. That is a model property, not a change in the category, and it is worth watching: the shelf AI shows a shopper is getting narrower.
Before/after measured on a frozen 20-prompt subset, old and new models run the same day. Grounding change is the new model minus the old model on identical prompts.
How this study was run.
Window
Data captured July 2026. The page is dated and refreshed monthly. Numbers reflect the state of streaming as of the capture date, including the July 2025 HBO Max re-rebrand, the May 2023 Discovery+ fold-in, and the 2025-26 NBA rights move.
Engines
ChatGPT, Gemini, Claude, and Grok. All four routed through a single shared retrieval layer (Perplexity Agent API) with web search enabled, so cross-model differences reflect the models themselves, not differences between their search engines. Source-mix findings describe the supply that shared layer feeds the models, not each vendor's native consumer search. A June 2026 route-comparison test (same prompts run on this shared layer and on each vendor's native search, with same-route re-runs as the noise floor) confirmed the boundary: grounding rates and cited sources are properties of the retrieval path, while brand-level recommendation findings held across both paths within re-run noise. This July 2026 edition is a methodology-v2 reading: the measured engines were upgraded to the current generation of each model (OpenAI GPT-5.6, Google Gemini 3.5 Flash, Anthropic Claude Haiku 4.5, xAI Grok 4.5), and the prior month's models were re-run on a frozen subset of the same prompts so month-over-month change is separated from the model upgrade.
Prompts
101prompts, voiced as a real subscriber would write them (no benchmark phrasing). Mix of brand-direct, category, occasion-and-intent, comparison, and purchase-decision queries. Split across two subscriber types: people stacking multiple services, and people deciding whether to cancel. Three sub-groups track time-sensitive queries ("what just dropped"), rebrand confusion ("is HBO Max the same as Max"), and social-proof queries ("reddit keeps saying drop X").
Personas
5 buyer personas spanning household type, genre lead, price elasticity, engagement pattern, and current subscription state. The axis descriptions render here. The detailed briefs (income, specific service stacks, backstories) are server-side only.
Outputs measured
For each engine response: grounding (was web search actually fired), brand mentions (matched against a canonical service list), citations with publication and last-updated dates, and the reason behind each recommendation tagged by an automated analysis pass against a fixed 14-reason list.
Conversation outcomes
Each persona-by-engine conversation closes on one of five states: Kept, Canceled, Paused, Skipped, or Inconclusive. If the persona doesn't decide on its own by the second-to-last turn, the runner injects a forced close so every conversation lands on a parseable outcome rather than running out the clock.
Scope
This is a US-market study built on five persona archetypes, run monthly. It measures engine behavior (what AI surfaced, with what reasoning, against what citations) on consumer-voiced prompts. Engine behavior shifts month to month, which is why the methodology is set up to be re-run. Re-running the same prompts hours apart already rotates a meaningful share of the recommended set, so any single capture carries a noise floor; the value of the study is in the month-over-month delta, not in any one snapshot.
We report 95% confidence intervals on per-engine proportions wherever the sample size is small. The headline numbers in the 60-second card are point estimates; the CIs sit next to the per-engine bars so you can read both the headline and the uncertainty.
What v2 adds
October 2026: UK, France, and Australia corpora (localized prompts, geo-routed engines). Q1 2027: paid-surface measurement once ChatGPT and Gemini ship sponsored-answer formats at consumer scale. The diagnostic-to-fix bridge (why each engine behaved this way, and what changes the behavior) is what mapou delivers as paid engagements rather than as a public report.
What this report does not do
We measure observable behavior, not model internals. Causal claims about why a given engine surfaced a given service require access to model weights and retrieval pipelines that no third party has. The honest deliverable is hypotheses anchored to observed behavior; those live in the finding sections and are labeled as hypotheses.
Study slug: svod-2026-07 · 443 measurements · refreshed monthly.
What mapou measures. How AI assistants alter commercial decisions inside high-intent consumer categories. The streaming wedge is AI Retention Intelligence: the layer that influences keep, cancel, pause, and stack decisions. The beauty wedge is AI Discovery Intelligence, which influences purchase. Both sit inside the same company-level frame: Conversational Commerce Intelligence.
Each monthly study applies the same methodology (multi-turn buyer conversations, four engines, decisions measured) to a different category. Streaming is the second category in the franchise after beauty.