September 2026 · 516 questions · replicated
Gemini answers hotel questions from a booking tool, not the web
Which product it reaches for depends on the category. Whether it reaches at all depends on how you ask — and that second half is the one nobody is measuring.
Ask Gemini “best hotels in New York” and it returns five properties, a price for each, and not one source. Ask it “why are there so many hotels in New York?” — same noun, same city, same minute — and it returns 5.94 cited publishers.
Nothing about the hotel changed between those two questions. What changed is that the first one triggered a tool and the second one didn’t.
The ladder
Accommodation questions only — ten phrasings across four cities and six lodging nouns, pooled over two capture waves (n = 408).
The two columns are near-perfect mirrors. Where a tool fires, sources go to zero; where none fires, they come back.
The top rungs are the same request in different clothes: give me somewhere to stay. The bottom rungs keep the word hotel and drop the ask. Losing the city barely matters — “best hotels” with no city named still fires 83.3% of the time. Losing the recommendation is what switches it off.
It isn’t reading the web and hiding its sources. There are no sources, because it never went looking.
One phrasing deserves singling out. “Best hotels in New York for the New York City Marathon” is event-anchored — the shape that, on Google’s other AI surface, pulls answers back towards blogs and forum posts and away from Google’s own panels. Here it does nothing: 97.9%, indistinguishable from a bare list. Whatever escape hatch exists elsewhere is not open on this one.
The category picks the tool
So far this could just be “commercial questions don’t get sources”. To test that, the same phrasings ran against six things that are not places to sleep — restaurants, cafés, bars, gyms, hair salons and bookshops.
L = accommodation, C = control categories. Bars show which tool answered.
The control never once reached the Hotels tool. It reached Maps instead — and lost its sources just the same.
This is the part that reframes the story. Cafés and gyms don’t keep their citations either; they route to a different Google product. The intent decides whether a tool fires. The category decides which one. Hotels get the booking module, everything local gets Maps, and either way the publisher who used to be read stops being read.
The control was captured twice as well, and the result is absolute: across 83 control questions in two waves, not one reached the Hotels tool. The Maps routing held too — questions about somewhere near a district went to Maps on 100% of captures in both waves.
There is one asymmetry, and it favours nobody. Across the accommodation questions, 1,002 of 1,006 attached links pointed at a Google Hotels search page. Across every control question, the number of attached links was zero. The booking module at least links somewhere. Maps answers just end.
It holds on a second run
Every accommodation question was asked again in a separate capture window. An effect this abrupt is exactly the kind that turns out to be one bad afternoon, so the replication was planned before the first wave was read.
| question phrasing | wave 1 | wave 2 | move |
|---|---|---|---|
| best X in {city} | 100.0% | 100.0% | 0.0 |
| X near {district} | 100.0% | 100.0% | 0.0 |
| X available tonight | 100.0% | 100.0% | 0.0 |
| I need an X for two nights | 100.0% | 100.0% | 0.0 |
| best X for {event} | 95.8% | 100.0% | +4.2 |
| single best X, and why? | 87.5% | 87.5% | 0.0 |
| best X (no city named) | 83.3% | 83.3% | 0.0 |
| X in {city} or {city2}? | 29.2% | 8.3% | −20.8 |
| why so many X here? | 4.2% | 0.0% | −4.2 |
| how does X work? | 0.0% | 0.0% | 0.0 |
Eight of the ten move by four points or less, and six do not move at all. The ladder is the same ladder. One rung is genuinely unstable — the two-city comparison fired a tool 29.2% of the time in the first wave, 8.3% in the second and 20.8% in a third. It sits in the middle of the ladder where the engine looks undecided, and it is the one number here I would not quote.
What it is not
It is not the word “hotel”. Six accommodation nouns were tested and they land within eighteen points of each other. There is no keyword to avoid.
| noun asked for | questions | tool fired | citations / question |
|---|---|---|---|
| budget hotels | 34 | 79.4% | 0.71 |
| hotels | 34 | 73.5% | 0.74 |
| boutique hotels | 34 | 61.8% | 1.24 |
| hostels | 34 | 61.8% | 1.24 |
| luxury hotels | 34 | 61.8% | 1.29 |
| resorts | 34 | 61.8% | 0.68 |
Pooled across all ten phrasings, which is why none reaches 100%.
And it is not “Gemini stopped citing”. On the explanatory questions it cites perfectly well — 97.9% of them came back with sources. The engine reads the web when it decides the question calls for reading. It has simply decided that “where should I stay” is not that kind of question.
What this changes for a hotel
The working advice for two years has been: be the page the assistant reads. Publish the guide, earn the citation, collect the link. That advice quietly assumes the assistant is reading pages at the moment someone asks where to stay.
On these questions it isn’t. The property still gets recommended — the answers name real hotels, with reasons — but the recommendation is assembled from what the model already knows plus a booking feed, and the link beside it goes to a Google search page rather than to the hotel. The travel writer whose ranking used to inform that answer is not cited, because they were not consulted.
Which moves the question from are we cited to are we known: whether the model already holds the property, and whether the property is described well enough in Google’s own hotel data to be pulled into the panel. Those have different fixes, and only one of them is about publishing pages.
The citations have not disappeared. They have moved to the questions people ask before they are ready to book — the why, the how, the compare. That is a narrower and earlier place to be visible than it used to be, and it is still open.
Limits
Read this before quoting any of it
One engine, two days. 516 questions across 7–8 September 2026, zero capture errors. Both halves were captured twice and both reproduced. This is Gemini only — nothing here says whether other assistants, or Google’s other AI surfaces, behave the same way.
Four cities, one vantage. New York, Berlin, Seoul and Helsinki, all asked from a United States location. Nothing here licenses a claim about how this behaves for a user in France or Japan.
Tool detection is a marker scan. The tool that answered is read from the answer payload. It is reliable for the tools we can name and silent about any we cannot — so “no tool fired” means “no tool we recognise”, and the citation count is the sturdier signal in both directions.
Method
Six accommodation nouns and six non-accommodation control categories, each asked in up to ten fixed phrasings across four cities. 516 questions in total across three capture waves, zero errors; the accommodation ladder pools two full waves (n = 408) and the control block was captured twice (n = 83).
Every question was generated mechanically from one template set, so phrasing is the only thing that varies between comparable cells. Two phrasings read as nonsense for a place you sleep in — “a hotel open right now”, “which hotel should I go to for one afternoon” — and were rewritten for accommodation to preserve the intent rather than the wording; the control categories deliberately ride only phrasings whose wording is identical across both halves, so no comparison here rests on a rewritten cell. The prompt list was fixed and hashed before any of it was sent, and shuffled so that no arm occupied a contiguous block of the batch.
Related: AI Mode stopped linking publishers — the same loss of attribution on Google’s other AI surface, by an entirely different mechanism.