Nicolas Sitter
September 2026AI Search · Methodology

Change the question, and Google AI Mode starts linking the web again.44 place types, 10 question forms, two switches.

TL;DR: Last week I reported that 99.96% of Google AI Mode’s citations were searchviewer wrappers. That corpus was hotel prompts in list form. Run the same detector across a crossed factorial — 44 Google Places types × 10 question forms × 4 cities, 1,868 captures with a 1,496-capture replicate two hours later — and AI Mode wraps 76.2%. What decides it is recommendation intent, not place-ness: the same noun in the same city runs 0.876 on “best cafés in Berlin” and 0.198 on “why are there so many cafés in Berlin?”. Google’s own Geographical Areas types are exempt, at 0.086.

NS
Nicolas Sitter
Published 6 September 2026

Summarize with AI

ChatGPTPerplexityClaudeGeminiGrok

Executive summary

0.876 → 0.198
same noun, list frame vs. why frame
median estimator: 0.909 → 0.000
76.2%
of citations wrapped, wave 1 (pooled)
median capture: 0.400
0.086
Geographical Areas types, wave 2
0.273 in wave 1, pooled
1,868 + 1,496
AI Mode captures, two waves
fired two hours apart

I have to correct myself before I can tell you anything else. The searchviewer study I published off the run of 2026-08-31 measured 29,556 wrapped citations out of 29,568 and reported the share to two decimal places. Every prompt behind it asked for a ranked list of hotels in a city. Vary the question instead of holding it fixed and the wrapper stops looking like a property of AI Mode and starts looking like a response to what was asked.

So this study varies both axes at once: 44 primary types from the Google Places taxonomy — restaurants and cat cafés at one end, school districts and zip codes at the other — crossed with 10 question forms, across New York, Berlin, Seoul, Helsinki, then fired again two hours later so that every difference has a measured floor to clear.

Two conditions move the outcome and they move it independently. The frame of the question does most of the work: the 10 forms run from 0.876 down to 0.000, while the 7 category strata cover a much shorter distance, from 0.851 to 0.273. And inside the recommendation frames, one family of Places types sits the whole thing out: the 6 types Google files under Geographical Areas.

The correction, stated plainly. 99.96% was a hotel-list-query figure. It was measured correctly and it is still true of that corpus, whose residue is 0.041% of citations — 12 rows out of 29,568. Across this factorial the residue is 26.9%. Both readings are right, the difference is the design of the prompt set, and nothing changed at Google in between. Any sentence that turns 99.96% into a claim about AI Mode as a whole is over-reading it, including sentences of mine that a reader could reasonably have read that way.

1. What the 99.96% was measuring

A figure can be exact and still describe something smaller than its wording suggests. The earlier corpus was hotels, asked for as a ranked list, in a city, week after week. That is one cell of a large grid, and it happens to be the cell where wrapping is most complete. Widen the grid and the citation layer looks like this:

What carries an AI Mode citation across the factorial

Wave 1 only, 15,943 citation rows. The six buckets are mutually exclusive and their shares sum to 100%.

Counted over citation rows, 73.1% of the wave-one layer is the entity wrapper. Counted over captures — which is what a person actually sees — it is 76.2% in wave 1 and 73.3% in wave 2. Either way, the distance to 99.96% is not a rounding matter: the hotel run left 12 unwrapped citations in 29,568, and this factorial leaves 26.9% of its rows pointing somewhere else, a good part of it at ordinary websites.

Two estimators, and they disagree. Two estimators, and they do not agree. The pooled rate divides wrapped citations by all distinct citations across the run; the median rate is the middle capture’s own ratio. Pooled weights an answer with thirty citations thirty times as heavily as one with a single citation, and the median does not. Every share on this page says which one it is. The gap is not cosmetic: the same wave-one data is 76.2% pooled and 0.400 at the median capture, because a handful of long, heavily-cited place answers carry a lot of pooled weight while the typical answer cites less and wraps less of it. The pooled figure is the one that compares with the hotel study, so it leads — and it never travels without its median.

2. Same noun, four frames

The cleanest thing in the study is a ladder. Hold the place noun constant, hold the city constant where there is one, and change only the shape of the question. Each rung is generated mechanically for all 44 types, so this is four columns of the factorial rather than four prompts.

Wrap rate by question form: the intent ladder

Rungs one, two and three all contain the place noun. Rung two removes the city; rung three keeps the city and removes the request for a recommendation. The vocabulary is identical across the three, so whatever selects the entity wrapper is responding to what is being asked for. Rung four — an informational question with no place ask in it — returns 0.000 in both waves, on both estimators.

Pooled, the two ends of the ladder are 0.6774 apart, against a form-level noise floor of 0.0328 measured from the replicate — 20.7 times the floor. The why_question rung reproduces to within 0.0027 between waves and transactional to within 0.0014. This is the one result here I would be surprised to see overturned.

All 10 question forms. Slots are filled mechanically from the same per-type table for every one of the 44 Places types, and the Places API type string is never sent in a prompt. A two-wave difference has to exceed 0.0655 at this level to mean anything. The median column is the middle capture's own ratio, and it is a harsher reading of the same run.
Query formPrompt as sent (café × Berlin)W1 pooledW2 pooledΔ pooledW1 mediann, W1 / W2
list_citybest cafés in Berlin0.8760.831−0.0440.909427 / 176
near_districtcafés near Kreuzberg, Berlin0.8720.868−0.0040.909176 / 176
transactionalcafé in Berlin open right now0.8580.857−0.0010.899176 / 176
compare_citiescafés in Berlin or Seoul - which city does them better?0.7990.776−0.0230.000176 / 176
event_anchoredbest cafés in Berlin for the Berlinale0.7460.722−0.0240.500176 / 176
single_citywhat is the single best café in Berlin, and why that one?0.6870.710+0.0230.292176 / 176
no_localitybest cafés0.6630.563−0.1000.00044 / 44
persona_practicalI have one afternoon in Berlin - which café should I go to, one that has wifi and seats for three hours?0.5940.595+0.0010.250176 / 176
why_questionwhy are there so many cafés in Berlin?0.1980.201+0.0030.000176 / 176
info_genericwhat is the difference between a flat white and a latte?0.0000.0000.0000.00044 / 44
The city matters less than the intent, and that is as far as I will push it. Dropping the city hurts less than dropping the list intent — but no_locality ran at n = 44, moved 0.100 between the two waves, and sits level with why_question under the median estimator. Take the direction; do not quote the size of that gap. Pooled, that is about 21 points against roughly 68 for the intent — a shape rather than a measurement. I would rather hand you a soft number labelled soft than a firm one quietly resting on one estimator.

3. Where the wrapper stops

The 44 types were grouped into 7 strata before any data existed, on the axis I expected to matter: how strongly one instance of the type has an identity of its own. 6 of those strata land on top of each other. The one that does not is the group Google’s own taxonomy calls Geographical Areas.

Wrap rate by category stratum, both waves
S1Commerce with heavy review density — restaurants, cafés, bars, hotels, gyms, hair salons.
S2Civic and institutional — hospitals, universities, libraries, courthouses, post offices, fire stations.
S3Landmarks and strong single entities — museums, stadiums, airports, castles, zoos, churches.
S4Natural features — beaches, peaks, lakes, rivers, national parks, viewpoints.
S5Weak or transient identity — ATMs, car parks, bus stops, public restrooms, storage, EV chargers.
S6Odd but real — cat cafés, karaoke, roller coasters, ferris wheels, opera houses, psychics, cemeteries, ski resorts.
S7Google Places “Geographical Areas” — towns, regions, counties, countries, zip codes, school districts.

6 of the 7 strata move less between waves than the stratum-level standard deviation of 0.0598. The exception is S7, and it moves downward — 0.273 to 0.086, while the commercial stratum sits at 0.851 and 0.813. Whether that fall is the edge of the noise or a rollout still in motion, two waves cannot tell you; a third would.

All 44 Google Places Table-A primary types, pooled across both waves and ascending. Per-type n runs 68–84 captures, which makes this a screen rather than a bound: read the ordering and the block structure rather than the third decimal of a single row. The noun column is what the prompt actually said; the API type string never appeared in one.
Places typeNoun in the promptStratumWrap rateCaptures
countrycountryS7 · Geographic areas0.02968
administrative_area_level_1state or regionS7 · Geographic areas0.04268
administrative_area_level_2countyS7 · Geographic areas0.05268
school_districtschool districtS7 · Geographic areas0.09168
postal_codezip codeS7 · Geographic areas0.13184
international_airportinternational airportS3 · Landmark0.45768
localitytownS7 · Geographic areas0.48384
universityuniversityS2 · Civic0.49868

The 5 lowest-wrapping types in that table are all Geographical Areas: country, state or region, county, school district, zip code. There are 6 area types in the sample, and the sixth breaks the pattern: locality at 0.483, which is higher than international_airport. Asking which town in a region is worth moving to is apparently half a recommendation and half a geography lookup. At the other end sit hotel, stadium, national park, lake, castle, between 0.878 and 0.916.

A prediction of mine that came back wrong. I put post_office in the sample as the weak-identity case in the civic stratum: hundreds of interchangeable branches, no reviews worth the name, nothing an entity resolver ought to find distinctive. I expected it near the floor. It came back at 0.868, above both restaurant and cafe. Branch-level entities resolve perfectly well, and neither fame nor review density is the axis this runs on. Publishing the miss is cheaper than quietly dropping the type from the table.

4. Two switches, and both have to be on

Question form and category are separate conditions, and crossing them shows that neither one is sufficient on its own. Ask for a recommendation about a geographic area and the wrapper stays away. Ask an explanatory question about a restaurant and it stays away too. It wants the frame and the type.

The crossed grid: 10 question forms × 7 category strata

Each cell sits at n ≈ 24, where the measured standard deviation of the two-wave difference is σ = 0.0955 and a difference has to clear 0.191 to mean anything at all. Read it as a picture of where the two conditions overlap: no individual cell value is quotable, and the two waves are two hours apart.

Read the blocks. The top-left region — recommendation forms crossed with establishment strata — is dark almost everywhere. The right-hand column is pale down its whole length, whatever the question. The bottom rows are pale across their whole width, whatever the type. And the corner where both conditions fail is at the pale extreme. That is what “two independent switches” looks like once you draw it.

Do not read a cell. Each square rests on about 24 captures, where the replicate puts the standard deviation of a two-wave difference at 0.0955 and the threshold at 0.191. I quoted one of these squares out loud mid-analysis as though it were a finding, and the replicate is the only reason it is not in this article. The grid is a picture of the interaction; the row and column margins are the numbers, and they are in the two sections above.

5. What Google still links to

The hypothesis worth killing was the strong one: that a citation resolving to a Places entity is always rendered as a wrapper. It fails on its own best ground. Here are 6 citations from the cell where wrapping should be most complete — a recommendation form, an establishment stratum — in which AI Mode linked a page about one specific venue. Five of the 6 are the venue’s own website; the sixth is a Michelin guide entry for a single Berlin hotel.

Pages about a single named venue, cited unwrapped, inside recommendation-shaped answers about establishment types. Drawn from the adjudicated sample described below, rather than picked out of the interesting-looking rows.
HostWhat the page is aboutPlaces typeStratumQuestion form
sidgolds.comSid Gold’s Request Room New YorkkaraokeS6single_city
panoramapunkt.dePanorama Point Berlinscenic_spotS4event_anchored
humboldtforum.orgRoof Terrace, Humboldt Forumscenic_spotS4event_anchored
paarautatieasema.fiLuggage storage, PäärautatieasemastorageS5persona_practical
guide.michelin.comChâteau Royal BerlinhotelS1list_city
tools.usps.comJAMES A FARLEY, USPSpost_officeS2persona_practical

Six examples prove the rule is not absolute. They do not say how often it breaks, and that needed the one step in this study a machine could not do. A citation that is not a wrapper is not automatically a place: a Reddit thread about bars, a “ten best cafés” round-up and an ATM-fee explainer are all residue, and none of them is a venue that went unwrapped. Somebody has to read the URLs. 200 of them were drawn by a seeded hash before anyone saw a title — 150 from the cell where wrapping should be most complete, 50 from the rest of the publisher layer as a check on whether that cell is doing anything — and two raters labelled them blind.

How often an unwrapped citation is a single venue. Across the 150 sampled residue rows, 37.3% are a page about one specific named venue — 56 rows, 95% interval 0.300 to 0.453, spread across 45 different hosts. In the control slice it is 10.0%, so the sampling cell is really selecting for something.

This is the share of the UNWRAPPED residue that is a single venue’s page, not the share of all AI Mode citations. It is also a lower bound: ambiguous rows were resolved away from the falsifying label, and a personal blog about one venue was labelled forum, not a venue page.

Only a page about one specific named venue falsifies the claim; a listicle, a forum thread or a fee explainer is residue but not a place cited unwrapped. Two raters labelled the sample blind (Cohen’s κ = 0.935, 10 disagreements in 200), and ties were resolved away from the falsifying label.

That pooled figure is the least interesting thing in the table, because the sample is not homogeneous and the spread inside it runs to twenty-seven-fold. Where Google is asked about a commodity — a cash machine, a public toilet, somewhere to park — the residue almost never contains a venue’s own page. Where it is asked about a cat café, a university library or a stadium, better than half of what escapes the wrapper is exactly that.

Share of the sampled residue that is one named venue’s page, by stratum. Cells are small — S1 commercial is only ten rows, so this sample supports no hotel-or-restaurant-specific rate. Read the ordering, not the decimals.
StratumWhat it coversSampled rowsA single venue’s page
S6curiosities2665.4% (17)
S2civic3259.4% (19)
S3landmarks1957.9% (11)
S1commercial1030.0% (3)
S4natural2222.7% (5)
S5commodity facilities412.4% (1)

The wider residue has a shape of its own. Across 2,821 publisher citations in wave 1, what survives is community and reference material rather than the travel-media layer people optimise for: the 6 hosts at the top of the list account for 32.7% of it between them.

Top hosts among 2,821 wave-one publisher citations, deduplicated within each capture. Shares are of the publisher slice alone rather than of all citations, and the 12 here are a long way from covering it.
HostCitationsShare of the publisher layer
reddit.com2247.9%
en.wikipedia.org2167.7%
youtube.com1866.6%
facebook.com1184.2%
tripadvisor.com1124.0%
instagram.com672.4%

Reddit is worth naming on its own. It is the largest single survivor of the wrapper — 7.9% of the publisher layer, more than any travel title or review site — and it turns up in 10.7% of answers, roughly one in nine. It is also densest in exactly the stratum where the wrapper is strongest: among commercial places, hotels and restaurants and cafés and gyms, 16.3% of everything that escapes as an ordinary link is a Reddit thread. For local queries, the forum did not go away when the wrapper arrived. It is what is left standing.

One caution before anyone builds a plan on that list. A citation that is not a searchviewer wrapper is not automatically a link you can attribute. The third carrier in the mix, 9.1% of wave-one rows, is google.com/goto?url= — Google’s Search-wide passthrough redirect, which Google confirmed publicly and has been rolling out since July 2026. That is prior art; it is no finding of mine, and it has been written up by Relevant Audience, Symphonic Digital and Digital Applied.

What this study adds is where the two wrappers sit relative to each other. Cut by question form, the goto share is highest exactly where the searchviewer share is lowest: the forms that never produce an entity wrapper are the forms where the redirect does most of its work. They divide the citation layer between them, and they are not equivalent in what they leave behind — an svid decodes offline to a Knowledge Graph ID, while the goto payload is encrypted and decodes to nothing at all. The absence of one wrapper buys you an attributable link some of the time and a differently opaque one the rest of it.

6. Three things that turned out not to matter

Where the request comes from

A sub-factorial fires byte-identical prompts through 4 exit countries, 192 captures in total. If the wrapper were a regional rollout, this is where it would show.

0.855
United States exit node
median 0.905 · n = 48
0.893
Germany exit node
median 0.875 · n = 48
0.883
South Korea exit node
median 0.911 · n = 48
0.868
Finland exit node
median 0.909 · n = 48

3.8 points separate the highest from the lowest, and the ordering does not even survive the change of estimator. That is a null, and it is a null I needed, because the 15-item pre-flight probe had shown a German-versus-American contrast large enough to look like a real geographic effect. The German captures were also the Berlin captures and the American ones were mostly New York, so proxy, city and category were perfectly confounded; that is what a probe is for. Never read a geography contrast off prompts that differ in anything else.

Whether a Knowledge Graph entity exists at all

The tempting explanation is that the wrapper appears whenever an answer assembles named entities. Control forms asking for lists of products and of people return no entity wrapper at any rate the detector can see — and products and people have Knowledge Graph IDs of their own. Inside the main factorial, the informational form sits at 0.000 across all 44 types, in both waves and on both estimators. Whatever this keys on, “there is an entity here” is not it.

Whether the detector is finding itself

The same prompt set went to 2 other assistants as a mirror. Of 478 citations returned by ChatGPT and Perplexity, 0 carried a searchviewer link:

0 / 348
ChatGPT: wrapped citations
120 captures
0 / 130
Perplexity: wrapped citations
120 captures

Which is the boring result you want from a control. The detector reads a feature of one surface, and it does not fire on strings that happen to turn up elsewhere.

7. Which claims clear their own noise floor

The point of firing the whole factorial twice was to buy a number for how much an AI answer moves when nothing has changed. Two hours apart, same prompts, same proxy, same batch mechanics. The spread between the waves is the floor everything else has to clear.

Measured from the two-hour replicate rather than assumed. Read down the last column: it is why this article reports at form and stratum level and refuses to report at cell level.
Level of reportingCaptures per cellSD of the two-wave differenceA difference must exceed
Question form44–4270.03280.0655
Category stratum219–2880.05980.1197
Form × stratum cell≈ 240.09550.191
Clears the floor comfortably
  • The intent effect: 0.6774 against a floor of 0.0328, 20.7×.
  • The area exemption: S7 against every other stratum, in both waves.
  • The overall rate: 76.2% and 73.3%, both a long way from 99.96%.
  • The proxy null and the cross-engine null, at the level they are reported.
Does not, and is labelled that way
  • The size of the city effect — direction only, at n = 44 and a two-wave move of −0.100.
  • Every individual cell of the crossed grid, at about 24 captures each.
  • The depth of the fall in S7 between waves, 0.273 to 0.086: the stratum result is solid, the trajectory is one pair of points.
  • Any single category’s third decimal, at 6884 captures.

One standard, applied in both directions. If a cell is too small to report when it disagrees with me, it is too small to report when it agrees, which is why the heatmap in the previous section carries no cell labels at all.

8. The parts I cannot account for

Three results in the tables above have no mechanism behind them, and I would rather leave them bare than dress them in a plausible story.

  • international_airport at 0.457. It sits in the landmark stratum, whose other members run from 0.722 (zoo) up to 0.900 (stadium). An international airport is about as strongly identified as an entity gets — it has an IATA code — and here it behaves like an administrative region.
  • university at 0.498. The same story one stratum over: the civic type I picked as its stratum’s strongest entity came back among its weakest.
  • One cell that reproduces near zero in both waves while its whole stratum sits near the top. It is the only square in the grid whose oddity survived the replicate, which makes it the only one worth going back to — and it is still a cell, so it is described here without being quoted.

The limits are shorter than the findings, but they bound them. This is one surface, read through one vendor’s scraper, on one day, in two waves two hours apart; a third wave a day later would separate a live rollout from stratum-edge noise on S7, and I do not have it. The residue was adjudicated from a seeded stratified sample rather than a census, because hand-labelling several hundred rows by eye is how a biased subset turns into a headline. And every per-type figure rests on tens of captures, which is enough to rank a taxonomy and short of enough to certify a type.

One boundary is worth naming outright: this measures what the citation layer of an answer looks like. It does not measure what anyone clicks, whether a wrapped citation sends a visitor anywhere, or what a business gains from being the entity inside one.

Methodology

Design. 44 Google Places Table-A primary types, assigned to 7 strata before the run, crossed with 10 question forms across 4 cities (New York, Berlin, Seoul, Helsinki). Prompts are generated mechanically from a per-type slot table of singular, plural, district, scope and practical qualifier; the Places API type string itself is never sent. Fired against Google AI Mode through Bright Data in raw mode, on a US exit node except in the proxy sub-factorial, with the arms seeded-shuffled into one batch so that no control arm sits at the end of the queue.

Captures. 1,868 in wave 1 and 1,496 in the two-hour replicate, plus 120 ChatGPT and 120 Perplexity mirror captures and the 15-item pre-flight probe — 3,619 in total. The form and stratum tables above do not sum to the wave totals, because the proxy sub-factorial, the non-place control arms and the hotel positive control sit outside the crossed block.

The outcome, and the two estimators. One capture is one observation, so that the clustering of citations inside an answer cannot inflate precision. Within a capture, citations are normalised and then deduplicated — wrapper rows on their decoded entity ID, publisher rows on the normalised URL — because a raw-row denominator would inflate every share by that answer’s duplication rate, and the duplication rate itself varies by category. The pre-registration requires both a pooled and a median reading of every headline share, and this article carries both.

Carrier buckets are host-parsed, and that is a fix. The shipped detector tested for the substring google.com/searchviewer, which misses google.de/searchviewer and matches a third-party URL that merely contains the string. Regional Google domains do occur in this corpus, and the failure ran in the direction that manufactures this study’s headline, so the analyzer parses host and path instead. Google URLs that are not the wrapper get their own buckets: reading “not a wrapper” as “a publisher” would have been wrong on a meaningful slice of the residue.

What this deliberately does not do. No named-entity recognition, no entity resolution, no Knowledge Graph or Places API call, no seed list of venues, and nothing written back to a database. Every quantity is a function of Bright Data payload fields plus an offline protobuf decode, which is what makes the run reproducible from the stored captures alone. The analyzer refuses to emit its reportable tables unless a fixed set of fatal guards passes — among them that the place-card layer never leaks into the citation population, that the cited flag is still boolean, and that the carrier buckets sum to the row count.

Pre-registered, then amended in public. The design was frozen before the first trigger and two of its decisions did not survive contact with the payload: the card-to-citation join it treated as a primary endpoint is impossible, because those two layers now key on different identifier systems, and the residue adjudication it had sized as an afternoon needed sampling instead. Both are recorded in the addendum with the date they changed, alongside the probe that got a geography effect wrong.

What would make this stale. If a later wave shows the area exemption closing, the mechanism is a rollout boundary rather than a taxonomy one. If the form-level standard deviation rises much above 0.0328, the surface is drifting and nothing at cell level can be read at all. If Gemini or another Google surface starts emitting the same wrapper, that is a separate study.

FAQ

Recommendation intent, far more than the subject of the question. Holding the place noun and the city constant and changing only the frame, the pooled share of an answer's citations carried in a searchviewer wrapper runs 0.876 on “best cafés in Berlin” and 0.198 on “why are there so many cafés in Berlin?”, with a purely informational question about the same category at 0.000. The median capture puts it more bluntly: 0.909 on the list query and 0.000 on the why-question. Removing the city instead of the recommendation costs about 21 points against roughly 68 for removing the recommendation, but treat the first of those as directional only: that rung ran at n = 44, moved 0.100 between the two waves and sits level with the why-question under the median estimator. The list-to-why gap is 20.7 times the 0.0328 form-level noise floor measured from the replicate. The second condition is the category: establishment types get the wrapper, and the types Google files under Geographical Areas largely escape it.

Where this sits

This is the second half of a finding I published in two goes. The searchviewer study established what the wrapper is and how to decode it; this one establishes when Google applies it, and puts my own headline figure back inside the corpus it came from. The February hotel study is the domain-level reading from before any of this started, useful now mainly as a record of what that layer used to look like. And the live AI Mode dashboard is where a third wave would surface, whenever I get round to firing one. If it closes the gap on areas, you will hear about that too.