AI Hotel Volatility 2026:the top 10 rewrites itself every week
Since May, the AI Hotel Landscape pipeline has asked five AI engines the same 616 hotel questions every Monday. That gives us something a snapshot study can’t: the ability to hold the questions still and watch the answers move. Eleven consecutive weeks, one prompt library, five engines — how much of an AI’s hotel ranking is still there seven days later?
TL;DR. Averaged across the five engines, only 29.2% of a weekly top-10 survives into the next week — roughly 7 of 10 top picks are replaced every Monday. Depth of churn is an engine personality: Gemini’s week-one cohort is still 60% intact ten weeks later, ChatGPT’s erodes to 24%, and Perplexity kept zero hotels in its top 100 for all eleven weeks. Meanwhile the pool itself is calm: a top-10 hotel almost never falls off the 500-deep list — it just loses its slot.
Back in February, the rankings-consistency study re-ran identical queries minutes apart and found 50.5% position-1 stability — ask twice, get a different winner half the time. The open question was whether that shakiness averages out once you aggregate hundreds of prompts into a weekly leaderboard. It does not. Even at the aggregate level — 616 prompts, thousands of captures per engine per week — the composite top-10 keeps just 29.2% of its members from one Monday to the next, and the top-100 keeps 49.7%.
The interesting part is how each engine churns. Some jitter around a stable core, one drifts steadily away from where it started, and one behaves close to a weekly lottery. The label on the churn matters more than its size: jitter is noise you can average away; drift means the model’s picture of the market is actually moving under you.
1. The setup
Every Monday the landscape pipeline fires the same 616-prompt library (56 destinations × 11 templates) at ChatGPT, Gemini, Perplexity, Copilot and Google AI Mode, then resolves every recommended hotel to a canonical hotel ID. Ranking hotels by weekly mention count per engine yields a 500-deep leaderboard per engine per week. This study compares those leaderboards across the 11 consecutive weeks from May 11 to Jul 20, 2026 — the full run of the current prompt library, which went live on May 11 (earlier weeks used a smaller library and are excluded, because a prompt-set change reshuffles the hotel universe all by itself).
- Identity across weeks is the canonical hotel ID, so a renamed listing or alternate spelling doesn’t count as churn.
- Same questions, same cadence. With the prompt side held fixed, any movement comes from the engines’ answers themselves.
- Grok is absent: its scraper broke in late May and the platform has been suspended from the panel since; it only has two weeks inside the window and is excluded rather than imputed.
- Depth guardrail: at rank 100 a hotel has 3–5 mentions in a week, so tie-order noise is real deep in the list. Headline numbers therefore use the top 10 and top 100, and section 4 shows how churn scales with depth.
2. The weekly rewrite
Take an engine’s top-10 on any Monday and check the following Monday: on average only 29.2% of it is still there. ChatGPT is the steadiest at the very top (34%), Perplexity the wildest (17% — in a typical week it keeps fewer than two of its ten). Widen to the top 100 and the spread between engines opens up: Google AI Mode retains 57.9% week over week, Perplexity just 28.4%.
| Engine | Top-10 kept | Top-100 kept | Median rank move | Anchors (11/11 wks) | Wk-0 cohort alive @ wk 10 | Never-seen newcomers |
|---|---|---|---|---|---|---|
| ChatGPT | 34% | 51.8% | 80 | 5 | 24% | 2.8% |
| Gemini | 29% | 56.7% | 55 | 18 | 60% | 2.5% |
| Perplexity | 17% | 28.4% | 78 | 0 | 19% | 10.5% |
| Copilot | 33% | 53.5% | 58 | 6 | 37% | 2.1% |
| Google AI Mode | 33% | 57.9% | 60 | 12 | 44% | 2.7% |
3. Three volatility personalities
Same-size churn can hide opposite behaviours, so we tracked each engine’s May 11 top-100 as a cohort: week by week, how many of those exact hotels are still in the top 100? The shape of that curve — flat, sloping, or cratered — separates noise from drift from chaos.
Gemini: noise around a core
The curve is flat: 53% after one week, 60% after ten. Hotels leave and come back — 42 of its 60 surviving cohort members dropped out of the top 100 at some point and returned. 18 never left, the most of any engine.
ChatGPT: slow drift
Decent weekly retention (51.8% at top-100) yet the cohort erodes almost monotonically: 59% → 40% → … → 24%. Hotels that leave tend to stay gone. Its list isn’t jittering around a fixed set; the set itself is migrating.
Perplexity: near-lottery
The cohort collapses to 31% in a single week and ends at 19%. Not one hotel held its top 100 for all eleven weeks, and 10.5% of each week’s top-100 are hotels never seen before in any prior week’s top-500 — the other engines sit at 2.1–2.8%.
Copilot (37% at week 10) and Google AI Mode (44%) sit between the flat and the drifting curves. Worth noting against the engine-personality picture from the landscape study: Copilot may be the entity engine in what it cites, but its hotel picks churn like everyone else’s.
4. Membership is calm, position is not
Here’s the reassuring counterweight to all that churn. Ask a different question — not “did the hotel keep its slot?” but “is it still anywhere on the 500-deep list?” — and the picture inverts. A hotel ranked in the top 10 this week stays somewhere in the top 500 next week 94–99% of the time on four of the five engines (Perplexity, again the outlier, 86%). The deeper the current rank, the leakier it gets, down to 30–45% for ranks 251–500 — the zone where hotels sit on 1–2 weekly mentions and a single conversation can flip membership.
| This week’s rank | ChatGPT | Gemini | Perplexity | Copilot | Google AI Mode |
|---|---|---|---|---|---|
| #1-10 | 99% | 96% | 86% | 97% | 94% |
| #11-50 | 91.8% | 93.2% | 68% | 93% | 94% |
| #51-100 | 85.4% | 81.2% | 53.4% | 86.2% | 88.8% |
| #101-250 | 71.5% | 64.1% | 42.7% | 67.9% | 72% |
| #251-500 | 45.2% | 41.3% | 30.1% | 45% | 43.2% |
Reading: of the hotels an engine ranks #1–10 this week, the share still anywhere in its top 500 next week. Averaged over all 10 week-pairs.
5. The anchor hotels
Across five engines and eleven weeks there were 41 engine-seats that never turned over — held by just 29 distinct hotels. 8 of those hotels held a permanent top-100 seat on two or more engines, and 4 managed it on three: Hotel Monteleone (New Orleans), the Cairo Marriott, Marina Bay Sands, and The Langham Gold Coast. The Cairo Marriott’s median position on ChatGPT over the whole run: #1.
Gemini hosts 18 anchors, AI Mode 12, Copilot 6, ChatGPT 5 — and Perplexity zero, so no hotel can anchor more than four engines even in principle (none reached four). These are disproportionately branded flagship properties: 28 of the 41 seats belong to hotels carrying a chain brand.
| Engine | Hotel | Country | Brand | Best | Worst | Median |
|---|---|---|---|---|---|---|
| Google AI Mode | The Taj Mahal Palace, Mumbai | IN | Taj | #4 | #66 | #5 |
| Google AI Mode | Hotel Monteleone | US | independent | #1 | #25 | #8 |
| Google AI Mode | Cairo Marriott Hotel | EG | Marriott | #1 | #45 | #10 |
| Google AI Mode | Marina Bay Sands Singapore | SG | independent | #1 | #57 | #11 |
| Google AI Mode | Hotel Adlon Kempinski Berlin | DE | Kempinski | #6 | #57 | #13 |
| Google AI Mode | JW Marriott Gold Coast Resort & Spa | AU | Marriott | #3 | #27 | #15 |
| Google AI Mode | Hotel Brunelleschi | IT | independent | #1 | #63 | #20 |
| Google AI Mode | The Langham Gold Coast | AU | Langham | #9 | #53 | #22 |
| Google AI Mode | Four Seasons Hotel Seoul | KR | Four Seasons | #8 | #45 | #24 |
| Google AI Mode | Four Seasons Hotel Hong Kong | HK | Four Seasons | #16 | #97 | #39 |
6. Chains vs independents
The anchor list hints at it and the retention split confirms it: on four of five engines, chain-branded hotels in the top 100 survive to the next week more often than independents. Google AI Mode shows the widest gap (61.6% vs 53.8%). Perplexity is the lone inversion — chains actually churn slightly more there (27.3% vs 29.7%), consistent with an engine that re-rolls most of its list regardless of who you are.
| Engine | Chain-branded | Independent | Gap |
|---|---|---|---|
| ChatGPT | 53.6% | 49.3% | +4.3 pts |
| Gemini | 57.8% | 55.3% | +2.5 pts |
| Perplexity | 27.3% | 29.7% | -2.4 pts |
| Copilot | 56.2% | 50% | +6.2 pts |
| Google AI Mode | 61.6% | 53.8% | +7.8 pts |
Week-over-week top-100 retention for chain-branded vs independent hotels. Positive gap = chains stickier.
7. What it means
- Judge AI visibility on a rolling window, never a single week. With the median surviving hotel moving 55–80 ranks between Mondays, a 4–8 week average is the shortest honest read. One-week “we entered the AI top 10!” claims — including celebratory vendor screenshots — will usually un-happen by themselves.
- Pick the metric that matches the layer. Pool membership (are you in the top 500 at all?) is stable enough to track weekly and to alarm on. Position is only meaningful smoothed.
- Know which engine you’re reading. A Gemini drop is likely jitter that mean-reverts; a sustained ChatGPT slide is more likely real drift worth investigating; a Perplexity swing, in this dataset, is close to weather.
- The anchor tier is small and brand-heavy. 29 hotels worldwide held a permanent top-100 seat anywhere this season, mostly chain flagships. For everyone else, volatility is the normal condition of AI visibility in 2026 — plan measurement (and expectations) around it.
Methodology
Data: the weekly AI Hotel Landscape pipeline — 616 prompts (56 destinations × 11 templates, EN, mixed personas/budgets) fired each Monday at ChatGPT, Gemini, Perplexity, Copilot and Google AI Mode; answers entity-resolved to canonical hotel IDs; hotels ranked per engine per week by mention count. Analysis window: the 11 consecutive weeks 2026-05-11 to 2026-07-20 — every week of the current 616-prompt library. The two prior weeks (library v1, a smaller prompt set) are excluded because churn across a prompt-library change mostly reflects the new seeding. Grok is excluded: absent from the data since 2026-05-25 and suspended after persistent scraper failures (2 usable weeks in the window); its absence is disclosed, not imputed.
Pool & metrics: all comparisons use each engine’s global top-500 per week (the uniform retention cap in the published table; per-country floor rows beyond rank 500 are dropped so every week is compared under the same policy). Week-over-week retention = share of this week’s top-N present in next week’s top-N, averaged over the 10 consecutive week-pairs. Cohort survival tracks the May 11 top-100 membership forward. Rank displacement = median |Δrank| for hotels present in consecutive top-500s. Newcomer share = fraction of a week’s top-100 absent from every earlier week’s top-500 in the window. Anchors = hotels in an engine’s top-100 in all 11 weeks.
Caveats: mention counts around rank 100 are 3–5 per week, so tie-ordering contributes noise deep in the list — the depth-gradient section quantifies exactly this, and headline claims stay in the top-10/top-100 where counts are higher. Results describe English prompts at weekly cadence on this prompt library; different phrasings or cadences could churn differently. Scrape-side variance (which conversations a platform returns) is part of what “volatility” means here — it is what a hotel tracking its own AI visibility would experience too. Analysis code and per-metric CSVs: summary.csv plus the full retention/survival/gradient/anchor tables in the same folder.