# nicolassitter.com — Full Content Index > Data-driven experiments on AI search, the web, and digital platforms by Nicolas Sitter ## Author Nicolas Sitter is a tech enthusiast running data-driven experiments on AI search, the web, and digital platforms. Content is organised into four areas: Research (nicolassitter.com/research) — the core body of work on how AI search discovers, ranks, and cites hotels, plus cross-industry AI-search infrastructure methodology; Playground (nicolassitter.com/playground) — the same measurement methodology pointed at non-hotel local businesses across cities and languages, lighter than the core research but still real data; Guides (nicolassitter.com/guide) — practical AEO/GEO playbooks synthesising the research; and Projects (nicolassitter.com/projects) — live dashboards and interactive tools. The "Published Experiments" section below is itself grouped the same way: Hotel AI Research, Methodology, Flights, then Playground. ## Guides ### AI Search for Hotels — The Complete Guide - URL: https://nicolassitter.com/guide/ai-search-for-hotels - Date: June 2026 - Summary: A practical, data-backed GEO (Generative Engine Optimization) guide for hotels, synthesising 30+ of the studies below into a four-step framework: Diagnose → Owned Media → Earned Media → Measure. - Core thesis: Most AI engines don't read hotel websites — they ground hotel answers on structured place and review data (overwhelmingly Google Maps/Places, ~89% of entity cards, plus TripAdvisor, Yelp and OTAs). So hotel GEO is mostly about controlling those sources, your Google Business Profile, and your reviews — not page copy. Schema markup helps indirectly because it feeds the systems the LLM calls (Knowledge Graph, Places), not the LLM itself. - Key actions: allow AI crawlers (only 3.3% of hotels block any); complete and correct the Google Business Profile; add consistent Hotel/LocalBusiness + FAQPage schema (36.3% of hotels have none); keep one canonical name everywhere; earn citations in the sources AI quotes (TripAdvisor, Yelp, OTAs, listicles, YouTube, Reddit); measure via GA4 AI referrers, branded/direct-traffic growth, and a monthly prompt-panel citation rate. ### AI Visibility for Hotels — The Complete Optimization Guide - URL: https://nicolassitter.com/guide/ai-visibility-for-hotels - Date: June 2026 - Summary: Combines AEO (Answer Engine Optimization) and GEO into one resource: how AI models recommend hotels, the data sources they use, the 5 pillars of AI visibility, an 8-step optimization checklist, and how to measure results — built on 1.2M+ AI citations across 6 models and 25 cities. - Key actions: audit the Google Business Profile (79% of Google AI Mode hotel links go to GBP); enforce entity consistency (89% recognition when names match); manage reviews on TripAdvisor (cited in 86% of AI Mode hotel queries); build multi-source presence so Reciprocal Rank Fusion rewards you; add structured data; monitor mention share and respond to model updates. ### ChatGPT Hotel Optimization - URL: https://nicolassitter.com/guide/chatgpt-hotel-optimization - Date: June 2026 - Summary: A technical guide to ChatGPT's hotel search internals — 12 systems, 7 data providers, and the Sonic classifier — and what each implies for optimization. - Key thesis: ChatGPT grounds hotel answers on Google Places and review providers, fuses sources via RRF, and rewards multi-source consistency over on-page copy. ### Google AI Mode for Hotels - URL: https://nicolassitter.com/guide/google-ai-mode-hotels - Date: June 2026 - Summary: Where hotel clicks go in Google AI Mode — 79% to Google Business Profile, 3.6% to OTAs despite OTAs being ~46.6% of citations — and the optimization playbook that follows. - Key actions: GBP is the lever for AI Mode; OTAs and TripAdvisor influence mention, not destination. ### Schema Markup for Hotels - URL: https://nicolassitter.com/guide/schema-markup-hotels - Date: June 2026 - Summary: The definitive Schema.org guide for hotels — Hotel, LodgingBusiness, HotelRoom, Reviews, FAQ, and Offer types with full JSON-LD examples and AI-visibility impact. - Key thesis: Schema doesn't feed the LLM directly — it strengthens the entity record (Knowledge Graph, Places) the model ultimately grounds on. With ~36% of hotels carrying no schema, it's a cheap competitive edge. ### How to Prompt-Track Your Hotel's AI Visibility - URL: https://nicolassitter.com/guide/prompt-tracking-hotel-ai-visibility - Date: June 2026 - Summary: The measurement layer. AI hotel rankings look random but are structured by city, guest persona and phrasing — so they are measurable if you measure around the variance instead of using a rank tracker. Anchored in the rankings-consistency study: run-to-run, only ~1.1 of the top 3 hotels repeat; position-1 stability ranges 17% (competitive markets) to 96% (constrained queries). - Method (6 steps): (1) build a frozen persona × location prompt panel across 4 intent tiers (branded / persona / neighbourhood / generic); (2) run every prompt ~5× per engine weekly, engines tracked separately, country/IP fixed, logged-out; (3) report mention/citation rates with a margin of error, plus sentiment and attributes, not booleans; (4) when you score zero, climb the is-the-zero-real ladder — L0 prompt health, L1 run 10×, L2 more prompts, L3 branded, L4 home-country proxy — and record the level that surfaced you; (5) measure the booking journey (Problem → Exploration → Comparison → Validation → Selection) as one conversation and track persistence; (6) log every answer's co-cited domains to see which sources each engine trusts and which lever to pull. ### How AI Retrieval Works - URL: https://nicolassitter.com/guide/how-ai-retrieval-works - Date: June 2026 - Summary: The mechanism under GEO. How AI answer engines retrieve and quote content, and what it means for getting a hotel into the answer: the retrieval pipeline, chunks as the real unit, lexical vs semantic vs hybrid retrieval, the audition chunk, and the levers that actually move AI hotel recommendations. ### Does an AI Visibility Score Actually Mean Anything? - URL: https://nicolassitter.com/guide/do-ai-visibility-scores-matter - Date: June 2026 - Summary: Reframes the "a visibility score is not a booking" debate around a distinction the debate usually misses — two families of number. REAL MEASURES are ground truth you trust as outcomes (lagging, under-counted): bookings, the "how did you hear about us?" survey, branded/direct search, AI referral sessions, and server logs (ChatGPT-User/Claude-User/PerplexityBot — the only direct proof a model fetched you). BUILT STATISTICS are constructed proxies you trust as leading indicators (early, movable): prompt-panel mention rate, citation, and source mix. The job is to make the two families check each other. - Key thesis: (1) A built score is synthetic but not noise — 50.5% average top-spot stability up to 96.1% in tight markets (consistency study), and a panel modelled on real fan-out (location queries trigger search 98% vs 8% definitional) samples real intent. (2) Getting better is not "make more content" (unfalsifiable); it is controlled experiments — baseline both families, change one thing on a logged date, watch the per-engine statistic, confirm in logs + branded search, claim timing not cause. Works "a bit": the built stat is the closest observable to your action. (3) Grubby tactics genuinely work because engines read ABOUT you more than they read you — Grok grounds on Reddit 54.5% / Facebook 63.5%, ChatGPT drew 14% of hotel sources from Reddit — but they are fragile: ChatGPT 5.3 (5 Mar 2026) collapsed UGC from 21% of answers to 2% overnight. (4) Every real measure undercounts AI in the same direction (logs miss the recommendation, referral misses non-clickers, analytics buries it in branded/direct), so the visible number is always a floor — and the channel is enormous (~1B weekly ChatGPT users, Gemini approaching ~1B monthly), so the asymmetric mistake is reading the channel as smaller than it is. In a black box, the highest-fidelity probe is asking guests directly ("how did you hear about us?"). (5) Methodology frontier: better prompt panels, server-log traffic measurement, and deliberate AI attribution via the booking-form survey. - Proof point: Hotel Ranque, built from scratch with AI visibility as effectively the only channel — the built score led (week-4 long-tail Perplexity → week-20 generic Paris) and the real measures followed. On the booking form's "how did you hear about us?", 21 of 52 guests said AI search — against 27 for Google and 4 for other (real screenshot embedded). ## Published Experiments #### Hotel AI Research - Core body of work: how AI search discovers, ranks, cites, and sends traffic to hotels. ### 1. The AI Hotel Landscape 2026 - URL: https://nicolassitter.com/research/ai-hotel-landscape-2026 - Date: January 2026 - Summary: The most comprehensive study of AI hotel recommendations. Analysis of 1.2M+ citations across ChatGPT, Gemini, Perplexity, Claude, Copilot, and Grok. Covers 12,500+ prompts across 25 cities. - Key findings: OTAs dominate AI citations, Booking.com leads, independent hotels struggle for visibility. ### 2. Google AI Mode: Where Do Hotel Clicks Go? - URL: https://nicolassitter.com/research/google-ai-mode-hotel-study-2026 - Date: February 2026 - Summary: Analysis of 4,000 hotel queries in Google AI Mode. 79% of hotel clicks go to Google Business Profiles. 84K+ references analyzed across 1,146 hotels. - Key findings: GBP dominates click distribution, hotel websites get minimal direct traffic from AI Mode. ### 3. Do French Hotels Blog? A 15,000-Hotel Study - URL: https://nicolassitter.com/research/french-hotel-blog-study-2026 - Date: January 2026 - Summary: Study of 15,155 French hotel websites measuring blog adoption. 49.3% have blogs, but only 1 in 4 are actively maintained. - Key findings: Blog presence correlates with higher AI visibility, but most hotel blogs are dormant. ### 4. Anatomy of a ChatGPT Hotel Search - URL: https://nicolassitter.com/research/anatomy-chatgpt-hotel-search-2026 - Date: March 2026 - Summary: Technical teardown of how ChatGPT builds hotel recommendations. 12 systems, 7 data providers, 424 A/B tests analyzed. - Key findings: ChatGPT uses a complex multi-provider pipeline including Yelp, TripAdvisor, and Google Places. ### 5. How Consistent Are AI Hotel Rankings? - URL: https://nicolassitter.com/research/ai-hotel-rankings-consistency-study-2026 - Date: February 2026 - Summary: Replicating SparkToro's consistency methodology for hotels. Only 50.5% position stability across reruns of 4,000 queries. - Key findings: AI hotel rankings are volatile — the same query produces different results each time. ### 6. Hotel Schema.org Adoption Study - URL: https://nicolassitter.com/research/hotel-schema-adoption-study-2026 - Date: March 2026 - Summary: Scanned 121,425 hotel websites across 7 countries for structured data. 36.3% have no schema at all, 41% use the wrong type. - Key findings: Most hotels are invisible to AI due to missing or incorrect structured data. ### 7. Yelp in ChatGPT: Hotel Data Study - URL: https://nicolassitter.com/research/yelp-chatgpt-hotels-study-2026 - Date: February 2026 - Summary: How Yelp is integrated into ChatGPT hotel queries. 14 destinations analyzed, 33% Yelp integration rate in US markets. - Key findings: Yelp integration is US-focused, European hotels rarely appear via Yelp in ChatGPT. ### 8. What Hotels Are Actually Called: A Naming Study - URL: https://nicolassitter.com/research/hotel-naming-study-2026 - Date: March 2026 - Summary: Analysis of naming conventions across 121,425 hotels in 7 countries. 8 analysis angles including word frequency, length, and star-rating patterns. - Key findings: "Hotel" is the most common word, luxury properties use longer names, regional naming patterns vary significantly. ### 9. Hotel robots.txt & AI Blocking Study 2026 - URL: https://nicolassitter.com/research/hotel-robots-ai-blocking-study-2026 - Date: March 2026 - Summary: 105,002 hotel robots.txt files parsed across 7 countries. Only 3.3% block any AI crawler. GPTBot most blocked at 2.9%. France leads at 7.5%. - Key findings: Hotels overwhelmingly allow AI crawlers. 2.1% use the "smart strategy" of blocking training bots while allowing search bots. ### 10. Hotel llms.txt Adoption Study 2026 - URL: https://nicolassitter.com/research/hotel-llms-txt-adoption-study-2026 - Date: March 2026 (updated 2026-05-09) - Summary: 105,002 hotel websites scanned for llms.txt files. Only 6.3% have one. US leads at 12.4%, France trails at 3.8%. WordPress SEO plugins drive 33% of adoption. - Key findings: llms.txt adoption is very early-stage. Higher-star hotels adopt at 2.5x the rate of 1-star. 7.3% misuse the file for access control rules. - Update (May 2026): Shopify now ships llms.txt, llms-full.txt, and agents.md by default on every store, exposed via /sitemap_agentic_discovery.xml. Example: https://respire.co/sitemap_agentic_discovery.xml. Not yet reflected in adoption numbers above (March crawl). ### 11. What Hotel Footers Reveal — 98K Study - URL: https://nicolassitter.com/research/hotel-footer-analysis-study-2026 - Date: March 2026 - Summary: 98,423 hotel footers parsed across 7 countries. Instagram is in 40.8% of footers, overtaking Facebook in the US and UK. 23.8% of copyright years are 3+ years stale. - Key findings: 9.9% of hotels link to OTAs from their own site. TikTok adoption at 11.8% in the US. Only 54.9% link to a privacy policy. ### 12. ChatGPT Hotel Data Sources: 100K Entity Study - URL: https://nicolassitter.com/research/tripadvisor-chatgpt-hotels-study-2026 - Date: March 2026 - Summary: 99,538 ChatGPT map entities tracked over 90 days. Google Places fell from 100% to 70.3% as Yelp (10.2%), TripAdvisor (0.1%), and Foursquare (0.2%) entered. - Key findings: TripAdvisor descriptions are 8.8x longer than Google. Yelp is US-city dependent. TripAdvisor links to its own pages, not hotels directly. ### 13. How Dirty Is Google Maps Hotel Data? 179K Study - URL: https://nicolassitter.com/research/google-maps-hotel-data-quality-2026 - Date: April 2026 - Summary: 178,647 Google Maps hotel listings across 11 countries analyzed. 17% fail basic quality checks. 8,167 are OYO vacation rentals listed as hotels. - Key findings: Belgium loses 54% of listings after cleaning. ChatGPT uses Google Maps as its primary data source (88.8% of map entities). Chain hotels average 3x more reviews than independents. ### 14. ChatGPT Hotel Index vs Live Web — What Changes When Search Goes Offline - URL: https://nicolassitter.com/research/chatgpt-hotel-index-vs-live-web-2026 - Date: April 2026 - Summary: We ran 400 hotel queries on GPT-5.4 and GPT-5.3 in live and cached mode. 83% of cited domains differ. The index and the live web recommend different hotels. - Key findings: OpenAI maintains its own search index alongside live web access. GPT-5.4 runs ~2 searches per response with advanced operators. Free-tier users likely never get live web results. ### 15. ChatGPT Hotel Ads Are Live — CPC Pivot, Ads Manager, $50K Entry - URL: https://nicolassitter.com/research/chatgpt-hotel-ads-live-2026 - Date: April 2026 (updated April 29, 2026) - Summary: Sponsored ads now appear in 20-35% of ChatGPT hotel queries for US users. Booking.com dominates at 43.5%. April 29 update: OpenAI quietly launched a self-serve ads manager, pivoted from CPM to CPC, and dropped the entry to $50K. Reportedly ~$100M annualised revenue six weeks into the test. - Key findings: Booking.com owns 43.5% of ad slots, followed by Airbnb (21.2%) and Expedia (17.6%). OTAs collectively control 87.7%. Ads first appeared March 31, 2026. CPC pivot resolves the early "scraper impression" problem: bots don't click, so they don't cost. ### 16. Hotel YouTube Channels — Activity Study 2026 - URL: https://nicolassitter.com/research/youtube-hotel-visibility-2026 - Date: April 2026 - Summary: We analyzed 2,583 YouTube channels linked from 98,423 hotel websites across 7 countries. Only 10% of hotels link to YouTube. 43.7% of those channels are ghost accounts (2+ years silent). Only 11.3% post monthly. Median views: 412. - Key findings: US hotels are 4x more active than Italian hotels on YouTube. 5-star hotels are 2.3x more likely to be active than 3-star. Medium-form video (1-10 min) dominates at 54.6%. 68% of channels have fewer than 100 subscribers. ### 17. ChatGPT 5.3 Halved Its Hotel Sources — March 5, 2026 Cutover - URL: https://nicolassitter.com/research/chatgpt-hotel-source-shift-2026 - Date: April 2026 - Summary: Daily ChatGPT UI runs of 140 world hotel prompts from 4 country locales (US, GB, DE, ES) across 87 days. On 2026-03-05 ChatGPT UI switched to GPT-5.3 (5.4 in API). URLs per answer dropped 49% (24→12), unique domains per answer 46% (18→10), domain pool per prompt 54% (132→61), inline-cite rate collapsed 100%→24%. - Key findings: Drop is uniform across locales (−47 to −53%) — a model/pipeline change, not a geography rollout. Losers: Booking −82%, Expedia −76%, Hotels.com −77%, TripAdvisor −69%, Reddit −93%, Wikipedia −93%, Four Seasons −94%. Gainers: a Stockholm-operated network of 17 Booking-affiliate SEO listicle sites (luxuryhotel.guide, all-boutique-hotels.com, hotels-with-balcony.com, couples-hotels.com, etc.) sharing Gandi registrar, Cloudflare NS pairs ANDY/RITA and BRENDA/GRAHAM, and image host images.luxuryhotel.guru. Ted Valentin named as curator on one site — Stockholm directory-site entrepreneur; other curators (Maja Holm, Elain Olsson, David Bachmann) are likely personas. Luxury queries hit hardest (−54%), affordable least (−41%). St Barts worst city (−64%). ### 18. How Claude Searches Hotels - URL: https://nicolassitter.com/research/how-claude-searches-hotels-2026 - Date: May 2026 (updated 2026-05-04) - Summary: Captured event streams of several Claude conversations about hotel discovery, across budget tiers, languages, brand filters, and date-anchored stays. The pipeline has two branches gated by the Connector Discovery setting (off by default). With Discovery off, almost everything goes through one direct Google Places call (places_search) plus a thinking-step re-rank by rating × review count and a map render. With Discovery on, Claude returns a small curated panel of OTA connectors (Booking.com, Tripadvisor, Trivago and a few others) plus a Browse-all link. Anthropic has stated it will not run sponsored ads, so the curation logic itself is the product. - Key findings (default mode): 8 of 9 captures called places_search; the only outlier used web_search. Number of parallel search calls scaled with prompt fuzziness — sharded by price tier, neighbourhood, synonym, or brand. Two captures used a second round to verify named hotels pulled from training. A French prompt was searched in English but answered in French. Things Places has no field for (pet-friendly, sauna, brand affiliation, star tier) were filled in from review snippets or Claude's own training knowledge. Dates appeared in map titles and closing offers but never in the search call. - Key findings (connector mode): asking Claude to book with Discovery off produces an inline opt-in prompt; turning Discovery on returns a curated picker. The visible-list problem (3-4 slots, growing supply of chain loyalty programmes and direct-booking platforms wanting in) is the open question — without an auction mechanism, the selection logic is closer to App Store editorial than to Google Ads. ### 19. How Mistral Searches Hotels - URL: https://nicolassitter.com/research/how-mistral-searches-hotels-2026 - Date: May 2026 - Summary: Captured Le Chat event streams across French, Italian and English prompts; specialist niches and a generic city-tier query; named-hotel sentiment, brand-vs-brand comparison, and accessibility-sensitive intent. The pipeline is the simplest of any AI we've captured for hotels: one web_search call (Brave-backed, toolType "rag") per entity, snippet paraphrase, inline references. No place data, no map, no booking surface. - Key pipeline findings: One Brave call by default; parallelised only when the prompt names multiple distinct entities to compare (Mama Shelter vs 25hours fired 2 parallel calls 135ms apart, wall-clock latency stayed ~1.3s). Mistral does not follow Brave rank order strictly — there's an LLM-side selection layer over the SERP. Reference grouping is sophisticated: per-entity for list answers, per-argument for two-sided summaries, with the same source cited twice when it carries opposing claims. - Query-rewrite rule (per-term): Globally-indexable English terms get translated ("un hôtel boutique" → "boutique hotel"); named entities stay (Saint Pierre, Gare de Lyon, Toscana); structural connectors keep the user's language ("hors quartier"). Response language always matches the prompt. - Year-injection: the vast majority of our captures had "2026" appended to the Brave query — often unprompted. Relative-time words ("recent", "ce week-end") get resolved to specific dates / current year. The lone exception we saw was a hard-negation prompt where the constraint was already pruning the result set. - Niche vs generic SERP shape: Niche specialist queries (cyclists, family-Tuscany, vegan-Lisbon, train-station-Paris) surface real specialist editorial sites (freewheelingfrance, its4kids, veganfamilyadventures, igares, hotelaparis). Generic queries ("best 3-star Marseille") surface programmatic SEO-spam aggregator networks (3-star-hotels.com, marseillehotel24.com, marseillefrhotels.com). Safety-sensitive queries override both with official institutional sources (visitberlin.de for accessibility). - Authority-laundering (four patterns): One review → "many guests"/"significant number"; self-marketing copy → "known for"; editorial coy-label → presented as brand name; outright fabrication of a non-existent property (Mama Shelter has no Vienna location, Mistral described "Mama Shelter Vienna" anyway using brand-homepage and Prague review snippets). - Intent gate on closing offer: Imminent + dated transactional → hard OTA punt (Booking/TripAdvisor/Hotels.com mention); brand-vs-brand → soft drill-in on prices/amenities; generic best-of, subjective sentiment, niche aspirational → refinement question only; safety-sensitive → refinement plus explicit recommendation to an external authority site. Le Chat has no booking connector or app picker. - Same-prompt head-to-head with Claude on "best 3-star hotels in Marseille": one overlap (Alex Hotel & Spa). Claude returns boutique-leaning Google Places winners ranked by rating × review count; Mistral returns chain and apartment-hotel-leaning paraphrases of SEO-spam aggregator pages. ### 20. The Schema.org Debate (2026): Why It Still Matters for Hotels - URL: https://nicolassitter.com/research/schema-org-grounding-loop-2026 - Date: May 2026 - Summary: AEO oversold schema as an LLM unlock. The pushback is right for the general case — transformers don't read JSON-LD as JSON-LD, they read tokens. But for local / hotel intent, every major AI grounds against Places / KG / OTA aggregator surfaces, all of which sit downstream of schema. So: schema → entity graph → grounding source → LLM. Same effect, different mechanism than the AEO pitch implies. - The four schema fields that move the needle for hotels: @type: Hotel (category gate — not LocalBusiness or Organization); sameAs (cross-platform identity, reconciles your hotel to its Booking / TripAdvisor / Wikidata / Maps profiles); typed starRating with Rating object (the category bucket gate — AI queries are sliced by "best 4-star", "luxury", "budget"); alternateName (cross-time identity for rebranded properties — Hôtel de la Paix → Alfred Hotel Beaune). - GBP framing: a Places API response in an LLM tool call exposes a narrow set of fields (rating, review_count, hours, photos, address) coming from Google Business Profile. That's what is visible to the model at call time and makes schema look redundant. But GBP has no amenities field, no room categories, limited typed attributes. Schema fills that gap and feeds the fuller hotel entity behind the KG. - HotelRoom / containsPlace / BedDetails / floorSize: under your control via schema, encode what GBP cannot. One-time markup cost, zero downside, natural place to expose room-level structure to any grounding source that consumes it. - FAQ schema: fine if relevant on the page, but not the AI-visibility lever. The boring four above do that work. - Where AIs source hotel data: Gemini → Google Maps / Places / KG natively; ChatGPT → 5+ providers including Places with rank fusion; Claude → one direct places_search call per query. Three stacks, all share one thing: their grounding sources sit downstream of the entity graph schema feeds into. ### 21. The ChatGPT Direct-Traffic Explosion for Hotels (May 2026) - URL: https://nicolassitter.com/research/chatgpt-hotel-direct-traffic-explosion-2026 - Date: May 2026 - Summary: On May 7, 2026, ChatGPT changed how it terminates hotel recommendations. Brand names became inline links pointing to the hotel's own homepage instead of dead-ending in citation chips. Across The Hotels Network's anonymised panel of more than 17,000 hotels, daily AI sessions jumped from a 31,688/day baseline (May 1–6) to 51,282/day (May 7–25) — a +62% step change that held through May 25 (a brief late-May dip, then recovery to 58K) without returning to baseline. - Key panel numbers: ChatGPT sessions in May ~852K (+823% vs Jan). AI % of all hotel-website sessions 0.64% → 0.92% (+44% relative). The number of hotels seeing AI traffic on a given day climbed from ~6,300 to a peak of 7,532. - Who actually gets the traffic (per-hotel distribution, May 7–25, ~10,000 hotels with measurable share): the ~0.9% network average hides a heavy skew. 3,730 hotels draw ≥1% of new sessions from AI, 1,613 ≥2%, 286 ≥5%, and 43 ≥10% — for that tail, AI is already a top acquisition channel. Demographic cuts (chain vs independent, star, geo) are pending a dimension join. - It's a ChatGPT-only story: Inside the AI mix, ChatGPT share went from 98.1% → 98.8% (+0.7pp). Perplexity 1.1% → 0.7% (−0.4pp). Claude 0.7% → 0.5% (−0.2pp). Perplexity weekly absolute sessions were essentially flat across Mar 16–May 11 (2,978 → 2,516). Claude inched up but on tiny base (1,554 → 1,799). Mistral / Copilot / Gemini / Grok are statistical noise in the panel. - Volume not mix — the surprising hotel-specific finding: Hotels were ALREADY homepage-heavy on AI traffic before May 7. Homepage share of AI sessions: 69.09% → 69.84% — essentially unchanged. The May 7 change drove volume, not page-mix. This is different from SaaS / B2B where homepage share reportedly grew from ~4% to ~24% on May 7. Hotel websites are built around a homepage that functions as the booking widget gateway; AI assistants were already routing brand-name queries there. - Two plausible reasons OpenAI may have shipped this: (1) CTR signal for the new CPC ads stack — embedding brand URLs inline and watching which get clicked produces exactly the training data a CPC ranker needs; (2) peace offering to the open hotel web — turns ChatGPT from a pure referral sink into a measurable referral source, helps with publisher and regulator relationships. - For hotels: AI mentions are now monetisable as direct traffic. The homepage that's already optimised for conversion just got 62% more cold-acquisition sessions a day. Hotels need chatgpt.com / openai.com tagged as a named GA4 channel group with booking-engine conversion wire-up. The companion measurement framework is at /research/how-to-measure-ai-hotel-traffic-2026. - Disclosure: Hotelrank (the hotel AI-visibility product behind this research) has been acquired by Lighthouse (The Hotels Network), and the author now works there. The panel data comes from The Hotels Network. - What's pending: segment cuts by chain/independent, star rating, geography — needs a separate join against THN's hotel-attribute dimensions. Booking-side revenue impact also pending (panel exposes referrers, not bookings). ### 22. Are Hotels in Common Crawl? (2026) - URL: https://nicolassitter.com/research/hotels-in-common-crawl-2026 - Date: June 2026 - Topic: A census of hotel presence in Common Crawl — the open web archive behind much LLM training data. The training-data complement to the retrieval-side studies. - Summary: 108,109 distinct hotel domains (from 142,405 hotels with own site + Google Place ID + >=10 reviews) checked against the columnar URL index of the May 2026 snapshot (CC-MAIN-2026-21), matched on www-normalised host. - HEADLINE: 60.6% of hotels are in Common Crawl, 39.4% absent. Of those present, 41.4% have depth (>=5 captured pages), 19.2% are shallow (1-4). ~2 in 5 reviewed hotels with a real website are missing from the data that trains LLMs. - Independents 61% > chains 45.9% (chains more often JS-rendered or CDN/WAF-blocked). Country: DE 69% best -> GB 54%, ID 47% worst. TLD: .de 71% best, .es 37% worst; .com deepest (~109 pages) vs local TLDs shallow (~35-42) - the English-corpus tilt one layer earlier, in training data. - Why absent: JS-only rendering, CDN/WAF blocks (invisible to robots.txt, so the 3.3% block study is a floor), or low connectivity (Common Crawl rank / Harmonic Centrality, per Metehan Yesilyurt). Free checker at /tools/common-crawl. ### 23. AI Hotel Memory (2026) - URL: https://nicolassitter.com/research/ai-hotel-memory-2026 - Date: June 2026 - Topic: A parametric-recall study — what AI models have MEMORISED about hotels, with web search turned OFF. The complement to every retrieval-side study here. Method adapted from Dejan Marketing's AI Brand Authority Index, extended with website verification. - Summary: Three cheap models (GPT-5.4-nano, GPT-5.4-mini, Gemini 3.1 Flash-Lite) were asked repeatedly, in JSON, to name hotels and return each one's website + address — no tools, temperature 1.0. ~150 runs/model for global chains; ~80 runs/model for Paris, Dubai, London, New York (~1,400 generations total). Returned domains were checked for DNS resolution, turning recall into a confabulation lie-detector. Total cost under EUR 20. - HEADLINE: Hotel chains are known cold — every model returns Marriott/Hilton/Hyatt/Four Seasons with their correct websites ~99% of the time. For individual hotels it cracks: the share of named hotels with a working website ranges from 97% (Gemini 3.1 Flash-Lite, Dubai) down to 47% (GPT-5.4-nano, Paris). The signature failure mode is the model knowing a real hotel exists but INVENTING its web address — e.g. Le Bristol Paris returned as the dead bristolparis.com vs the real oetkercollection.com; Hotel Plaza Athenee as plazaathenee-paris.com vs the real dorchestercollection.com. - Key findings: (1) chains overlearned (single predictable domain in training); (2) Paris is the hardest city for every model because its top hotels are independent palaces on unpredictable collection domains, and it produces the longest tail of invented names (nano: 469 distinct "Paris hotels" over 80 runs vs Gemini's 177); (3) counter-intuitively the CHEAPEST model, Gemini 3.1 Flash-Lite, had the most accurate and most consistent hotel memory (beats GPT-5.4-mini in cities), which is why Dejan's brand index could run on cheap Gemini; (4) asking for the website is the method's key — it forces a checkable commitment a bare name does not. - Caveat: DNS-resolves is a lower bound on accuracy (a domain can resolve yet be the wrong hotel). Measures three specific cheap models, not the full ChatGPT/Gemini consumer products (which use larger models + web search). ### 24. ChatGPT's Hidden result_source (2026) - URL: https://nicolassitter.com/research/chatgpt-result-source-retrieval-tiers-2026 - Date: June 2026 - Topic: Two undocumented fields in ChatGPT's raw web-search stream — result_source (the retrieval tier/vendor that fetched each cited page) and turn_use_case (a pre-search query classifier) — extracted at scale from hotel queries. The retrieval-side complement to the memory study. - Summary: 50,899 ChatGPT hotel captures (25 Dec 2025 – 22 Jun 2026, via Bright Data); result_source and turn_use_case parsed out of the raw SSE stream and joined to 30,002 flattened per-citation rows. ChatGPT-only — Gemini/Perplexity/Copilot/Google AI Mode emit no equivalent. - HEADLINE: ChatGPT sources hotel answers almost entirely from one licensed tier. Of 30,002 tier-tagged citations, 99.85% are "labrador" (a licensed/quality-gated content tier), 0.14% "bright" (Bright Data datasets), 0.01% "oxylabs" (scraped open web); the open-web "serp" tier never appears for hotels (0 of 30,002). - Key findings: (1) the field is brand new — first seen 26 May 2026 — and rolled out intermittently: share of tagged citations went 0% → peak 81.8% (week of 1 Jun) → 7.2% (week of 8 Jun) → ~44%, like a staged A/B; (2) turn_use_case gates whether the web is searched at all — 37.3% of hotel turns are "text" (no search, answered from training data), vs search 31.8% and local/maps 26.3%; (3) within the labrador tier, retrieved ≠ cited — official brand .com sites are cited 77–86% of the time when retrieved (Four Seasons press 86%, Hyatt 84%, Hilton 80%, Ritz-Carlton 77%, Marriott 68%) while aggregator listicles are cited 1–11% (hotelierschoice 2%, thehotelguru 1%, luxuryhotel.guru 2%, travelmyth 11%); (4) per-tier publisher signatures differ — bright skews to premium directories + individual luxury properties, oxylabs to scraped long-tail incl. Instagram and individual hotel sites, so a hotel's own site/Instagram reaches ChatGPT via the scraper tiers, not the licensed feed. - Takeaway for hotels: make sure the query intent even triggers a search (a third don't), then invest in clean first-party brand content — it gets cited far more than the aggregator listicles ChatGPT retrieves and discards. - Caveat: one vertical (hotels) and one collection method; tier mix is query-type dependent (serp may surface elsewhere); the field is ~4 weeks old and intermittent; bright (n=42) and oxylabs (n=3) samples are tiny. ### 25. How ChatGPT Pulls Hotel Prices (2026) - URL: https://nicolassitter.com/research/chatgpt-hotel-price-sources-2026 - Date: June 2026 - Topic: A network-source forensic study of where a hotel PRICE number comes from inside ChatGPT's retrieval network — the price-specific complement to the result_source study. Tests the hypothesis that ChatGPT live-scrapes OTAs (Booking et al.) for the number while the official, often JavaScript-rendered hotel site gets fetched but under-used. - Summary: 30 frozen price prompts (12 named hotels across 4 segments × 3 cities, plus open city-level "cheapest hotel" queries) × 8 iterations via Bright Data, country US, captured 25 Jun 2026 = 240 captures, 3,092 fetched documents, 471 cited. Each capture parsed from the raw SSE stream into its network-source layer: every fetched document with its result_source tier (labrador/bright/oxylabs/serp) and cited flag, plus the prose price quoted. - HEADLINE: ChatGPT fetches online travel agencies (Booking, Expedia, Kayak) more than every other source class combined — OTAs are 46% of all fetched documents — but cites them the LEAST (11% cite rate). It reads the price off the OTA pages, then footnotes more human-readable sources. - Key findings: (1) a price question almost always hits the live web — 98.8% trigger a web search (routed to "instant search"/"local"), never the "text" no-search bucket that dominates generic hotel queries, so attaching a price forces retrieval; (2) fetch ≠ cite — cite rates by class: forum/Reddit 100%, editorial 59%, deal/loyalty 32%, metasearch 19%, official hotel site 16%, OTA 11%; Reddit is the #2 most-cited price domain after Booking.com; (3) the cited price rides the licensed labrador tier — bright (Bright Data) appears in the fetch layer (170 of 3,092 docs) but is cited just 3 times, so the "bright dominates shopping" claim holds for fetching, not for the answer; (4) provenance shifts by hotel segment — palaces priced via official + loyalty channels (Amex Fine Hotels & Resorts, points blogs; OTA only 21%), global chains see official ≈ OTA (~31% each), boutique independents OTA 37% vs official 15%, open city queries OTA 44% + metasearch 25%. - Price tells (source predicts reliability): Premier Inn priced off the deal blog saverrooms.co.uk not premierinn.com; ibis Paris priced off Klook; an independent (Grand Hôtel du Palais Royal) cited hilton.com, a wrong-entity citation. Official-sourced figures (Hilton Times Square ~$235 from hilton.com; The Plaza $1,250–$1,900 from theplazany.com) should track reality. - Takeaway for hotels: keep prices in plain fetchable HTML text (not JS, not images); maintain OTA presence so the number exists to be read, but invest in earned forum + editorial coverage, which is what actually gets cited. - Caveat: N=240, one capture day, US proxy, English only — a snapshot; cited-source tier is imputed from the domain's modal tier (cited docs arrive via a separate footnote path); the structured shopping block never fired for hotels (a products surface). ### 26. AI Citation Throughput in Hotel Search (2026, Live) - URL: https://nicolassitter.com/research/ai-citation-throughput-2026 - Date: July 2026 (live — refreshed every Monday from the AI Hotel Landscape pipeline) - Topic: Citation throughput (a.k.a. selection rate) = cited sources ÷ retrieved sources per AI engine. Sources are the input; citations are the throughput. A hotel-intent, per-vertical take on the retrieved-vs-cited question, computed weekly from the landscape corpus (616 prompts × 56 destinations, 1.1M+ retrieved sources logged since April 2026). - HEADLINE: Only two engines expose the full retrieve→cite funnel. This July, ChatGPT cites ~26% of the hotel sources it retrieves (it was 52–67% until a mid-July 2026 product change halved its per-answer source panel from ~22 to ~10 and collapsed inline citing from ~13 to ~3 per answer); Google AI Mode holds a flat 9–10% for its whole measured history (two-thirds of its pool is google.com itself, selected at ~7%). Grok, before tracking ended June 3 2026, retrieved 581,549 sources and was down to 0.3% throughput (99.7% rejection) in its final month. - The Reddit inversion: where index-scale cross-industry data has AI engines rejecting nearly every Reddit page they retrieve, hotel intent flips the ranking — Reddit has the HIGHEST throughput of any major domain on both measurable engines: ChatGPT 59% (754 retrieved → 447 cited), AI Mode 57%. Versus tripadvisor.com ~30%, booking.com 24%, expedia.com 10%, hotels.com 8%. Consistent with the flights study (Reddit cited on 83% of fetches) and price study (100%). Not a contradiction of index-scale findings: a cross-intent candidate index is a different denominator from the per-answer source panel on hotel prompts — levels differ, rankings are the signal. Search visibility decides who enters the pool; model preference decides who exits it into the answer. - Other findings: (1) national Tripadvisor TLDs are throughput-penalized on ChatGPT (.com 30%, .ca 8%, .co.uk 2%). (2) By category, ChatGPT's highest-throughput bucket is social (≈ reddit.com, 59%), AI Mode's is review + editorial; OTAs convert at roughly half the average on both. (3) Editorial curation wins: guide.michelin.com 46%, cntraveler.com 40–60%, forbestravelguide.com 63%. (4) Grok's surviving 0.3% was almost exclusively brand-owned: booking.com plus fourseasons.com / marriott.com / hyatt.com official sites — it rejected tripadvisor.com 17,812 times out of 17,813 retrievals. - Method/caveats: "retrieved" = every URL in the engine's exposed source panel (for ChatGPT incl. the "more sources" overflow); "cited" = used inline in the answer. Perplexity and Copilot publish only the survivors and Gemini's cited flag is ~always true, so throughput is only measurable for ChatGPT, Google AI Mode (and Grok historically) — the measured-vs-hidden asymmetry every cross-engine citation study runs into. Funnel, trend and category numbers are live from the public data feed (/api/landscape); per-domain tables are a July 13–26 2026 snapshot until the domain-level public view ships. ### 27. AI Hotel Volatility (2026) - URL: https://nicolassitter.com/research/ai-hotel-volatility-2026 - Date: July 2026 - Topic: The time-series companion to the rankings-consistency study (#5): instead of re-running one query minutes apart, it holds the entire 616-prompt weekly library still for 11 consecutive weeks (2026-05-11 → 2026-07-20) and measures how much of each engine's aggregate hotel leaderboard survives week to week. - Summary: The AI Hotel Landscape pipeline fires the same 616 prompts (56 destinations × 11 templates) every Monday at ChatGPT, Gemini, Perplexity, Copilot and Google AI Mode; recommended hotels are entity-resolved to canonical IDs and ranked per engine per week by mention count (top-500 pool per engine-week). Metrics: week-over-week retention at top-10/100/500, cohort survival of the week-0 top-100, median |Δrank|, newcomer share, all-season "anchor" hotels, chain vs independent retention. Grok excluded (absent from the data since 2026-05-25, scraper suspended; 2 usable weeks) — disclosed, not imputed. Library-v1 weeks (04-27, 05-04) excluded to avoid measuring a prompt-set change. - HEADLINE: On average only 29.2% of an engine's top-10 survives to the next Monday — about 7 of 10 top picks are replaced weekly (ChatGPT 34%, Gemini 29%, Perplexity 17%, Copilot 33%, AI Mode 33%). Top-100 lists keep 49.7% on average. - Key findings: (1) three churn personalities — Gemini is noise around a stable core (cohort survival flat, 60% at week 10; 42 of its 60 survivors left the top-100 and came back; 18 hotels never left), ChatGPT is compounding drift (51.8% weekly top-100 retention but cohort erodes 59→24%), Perplexity is near-lottery (31% cohort survival after ONE week, 19% at ten, zero anchors, 10.5% weekly never-seen-before newcomer rate vs 2.1–2.8% elsewhere); (2) membership vs position — a top-10 hotel stays somewhere in the 500-deep pool 94–99% of the time on four of five engines (Perplexity 86%), while the median surviving hotel still moves 55–80 ranks between Mondays; (3) only 29 distinct hotels held a top-100 seat on some engine for all 11 weeks (41 hotel–engine seats, 28 chain-branded); 4 hotels did it on three engines at once — Hotel Monteleone (New Orleans), Cairo Marriott (median #1 on ChatGPT), Marina Bay Sands, The Langham Gold Coast; (4) chains outlast independents on every engine except Perplexity (widest gap: AI Mode 61.6% vs 53.8%). - Practical: judge AI visibility on a 4–8 week rolling window, never a single week; track pool membership (stable, alarm-worthy) separately from position (noise unless smoothed); interpret a drop by engine — Gemini dips mean-revert, sustained ChatGPT slides are more likely real drift. - Caveat: mention counts near rank 100 are 3–5/week so tie-ordering adds noise deep in the list (quantified in the depth-gradient section); English prompts, weekly cadence, one prompt library; scrape-side variance is part of what "volatility" means here — it is what any hotel tracking its own AI visibility would experience too. ### 28. AI Cross-Platform Consensus (2026) - URL: https://nicolassitter.com/research/ai-cross-platform-consensus-2026 - Date: August 2026 - Topic: The between-engine axis that completes the disagreement triptych with the rankings-consistency study (#5, within-engine noise minutes apart) and the volatility study (#36, within-engine churn week to week): do the five live AI engines agree with EACH OTHER on hotels at the same moment, on the same questions? - Summary: One pinned week (2026-08-03, identical for all five engines, asserted by the analysis script) of the AI Hotel Landscape pipeline; each engine's ten most-mentioned hotels per city across 56 destinations (hotels entity-resolved to canonical IDs, assigned to cities by geo bounding box; 6,613 rows, 0 dropped for missing IDs, 41 rows / 0.6% outside every city box). The five per-city top-10 sets are intersected: 1,490 distinct (city, hotel) recommendations. Grok excluded — its latest data is week 2026-05-18 (scraper suspended 2026-05-25), 11 weeks behind the pinned week. - HEADLINE: 57.6% of all (city, hotel) top-10 recommendations exist on exactly ONE engine; only 4.6% (69 hotels) are backed by all five. The average recommendation has 1.83 of 5 engines behind it — asked "which hotels should I book in X", five engines return five substantially different lists. - Key findings: (1) the consensus curve halves per engine added — 57.6% / 18.9% / 11.2% / 7.7% / 4.6% for 1→5 agreeing engines; (2) most similar pair Copilot × Google AI Mode (34.8 avg per-city Jaccard, 51.2 containment — the only pair sharing more than half its picks), least similar ChatGPT × Perplexity (16.4 / 27.1); Gemini × AI Mode second at 31.2; (3) Perplexity is the contrarian — 45.8% of its top-10 picks (252 of 550) appear on no other engine (next: ChatGPT 33.4%; most consensual: AI Mode 22.7%); (4) chain-branded hotels average 2.06 backing engines vs 1.65 for independents and reach all-five consensus at 6.6% vs 3.2% — a brand roughly doubles cross-engine agreement, while independents (850 of 1,490 recs) fill the single-engine tail; (5) city gradient from Auckland (49.6 avg pairwise Jaccard; its five top-10s hold just 19 distinct hotels, 5 on all five lists) down to Bali (3.8; 44 distinct hotels across 50 slots, zero all-five) — 17 of 56 cities have NO all-five hotel, including Paris, London, Rome, Tokyo, Bangkok and Bali; (6) the 69 all-five hotels span 39 cities, 42 chain-branded, and exactly ONE hotel is ranked #1 on every engine: The Taj Mahal Palace, Mumbai (1-1-1-1-1); runner-up Hotel Monteleone, New Orleans (#1 on four, #2 on ChatGPT). - Practical: audit all five engines separately (Copilot & AI Mode, and Gemini & AI Mode, are the only semi-redundant pairs); treat 3+-engine corroboration (23.5% of recs) as the quality tier; a Perplexity-only appearance is the least corroborated and (per the volatility study) least stable form of AI visibility; in fragmented markets (Bali, Paris, Phuket) city-level "AI rank" without an engine name is close to meaningless. - Caveats: one week is a snapshot of a moving target (the volatility study measured ~29.2% week-over-week top-10 retention), so the disagreement LEVEL is the finding, hotel names are examples; upstream truncation (global top-500 per platform-week + per-country floor of 50; 13 of 206 platform-country groups visibly truncated) can only understate uniqueness; 26 of 280 city-engine lists are shallower than 10; upstream unresolved-mention rates 1.5–4.3% per platform; gemini = the gemini_scraper pipeline label. ### 29. Is LinkedIn Useful for AI Visibility in Hotel Search? (2026) - URL: https://nicolassitter.com/research/linkedin-hotel-ai-citations-2026 - Date: August 2026 - Topic: A new `linkedin_count` column (mirroring existing youtube_count/reddit_count/facebook_count columns on fanout_captures) backfilled across every stored AI hotel-search capture with citation data — 105,377 captures, all six engines, all-time (December 2025 onward). - HEADLINE: LinkedIn appears in 197 captures (0.19%) — rare overall, but not one phenomenon. ChatGPT cites LinkedIn ~60% of the times it retrieves it (mostly company pages); Grok retrieves LinkedIn social posts constantly but converts only ~16% to citations; three other engines barely touch it, and Gemini never does in this dataset. - Key findings, via the flattened citation table (which additionally records cited vs. retrieved-only for engines whose payload exposes a broader pool — ChatGPT, Google AI Mode, Grok): (1) ChatGPT — 122 LinkedIn rows, 73 cited (59.8% conversion), 96/122 are company pages (a hotel's own LinkedIn profile), only 1/122 a bare social post; (2) Grok — 174 rows, 27 cited (15.5%), 79/174 social posts (hotel brand posts, guest reviews, even unrelated personal opinion posts); (3) Google AI Mode — 24 rows, 3 cited (12.5%); (4) Copilot (9) and Perplexity (3) only ever expose an already-cited subset via their APIs — 100% conversion there is a measurement ceiling, not a behavior; (5) Gemini — 0 LinkedIn rows, ever, in this dataset. - Query-type finding: every LinkedIn citation traces to a destination-level discovery prompt ("best/luxury/affordable hotels in [city]", optionally modified by traveler type or amenity) — the corpus contains no named-hotel lookup query to compare against, so this is a description of what surfaces LinkedIn within this corpus, not a tested contrast against another query type. - Content-quality caveat: one Amsterdam hotel query cited a LinkedIn Pulse listicle titled about New York City hotels; the same off-topic article recurred, uncorrected, across several other unrelated destination queries — presence in a citation is not the same as relevance. - Practical: a hotel's own LinkedIn company page is the one part of this that's worth maintaining for AI visibility (real, if thin, ChatGPT citation rate); posting LinkedIn content in hopes of AI citation is a much weaker bet, since most engines retrieve social posts far more than they cite them. An earlier same-day 400-row-per-engine sample had found essentially nothing — the full all-time backfill found six times that in absolute count once every engine and week was included, illustrating why a real column beat a sample here. ### 30. Did GPT-5.6 Really Kill Listicles? Not in Hotel Search (2026) - URL: https://nicolassitter.com/research/gpt-5-6-hotel-search-changes-2026 - Date: August 2026 - Topic: A publicly shared analysis (Tomek Rudzki, based on Peec AI data — 1M prompts, ~180M sources, one week before/after the GPT-5.6 rollout) reported major ChatGPT search-behavior changes after GPT-5.6: far more fan-outs per chat, a surge in `site:` search, fewer "best/top/vs/comparison" fan-outs, and listicle/comparison pages losing roughly half their citation share. We re-ran the same before/after comparison on our own 616-prompt AI Hotel Landscape ChatGPT corpus. - Timing correction: July 9, 2026 (the date commonly cited) is GPT-5.6's developer/API release, not its ChatGPT.com consumer default — that was August 6, 2026. Our own Bright Data capture surface (which scrapes the consumer experience) shows `fanout_captures.model` = gpt-5-5 through the Aug 3 weekly run and gpt-5-6 from Aug 10 — a window bracketing August 6 almost exactly. Slicing on July 9 would have compared gpt-5-5 against gpt-5-5, a null result by construction. - HEADLINE: none of the widely reported effects replicated in this vertical, and some ran backwards. Pre (Jul 27 + Aug 3, gpt-5-5) vs post (Aug 10 + Aug 17, gpt-5-6): fan-outs per chat 0.959 -> 0.965 (+0.7%, reported: +154%); site: operator adoption 0% -> 0% (reported: 0% -> 18.37%); citations per chat 9.85 -> 10.03 (+1.9%, reported: +25%); listicle share of citations (our own title/URL classifier, not Peec's) 33.8% -> 36.4% (+7.6%, reported: 15.77% -> 7.80%, -50.5%); comparison share of citations 0.31% -> 1.03% (up, on a small base of 14-76 rows/week; reported: 9.08% -> 6.17%, -32.1%). - Confound flagged explicitly: "best" share of fan-outs rose sharply (46.2% -> 79.7%, +72.3%) in our data — opposite direction from the reported -56.4% — but our prompt library is templated ("best hotels in [city]..."), and ChatGPT's fan-out sub-queries often closely echo the original prompt, so this metric is partly measuring prompt design, not an independent model-behavior shift. The confound is constant across pre/post, so the direction of change is still real, just not a clean measure of organic term preference the way a varied prompt panel's would be. "vs"/"comparison" fan-outs were 0% in every week checked, before and after — our corpus contains no comparison-style prompts to begin with. - Hypothesis for the gap (not proven): hotel destination-discovery ("best hotels in Bangkok") is inherently listicle-shaped and rarely invites a site: or vs-style fan-out, unlike the branded/comparison-shopping queries GEO-spam tactics typically target. GPT-5.6's retrieval changes may be concentrated in categories more exposed to that kind of spam than open travel discovery is. - Caveats: the comparison classifier is our own (title/URL pattern match for "vs/versus/comparison/alternatives" and "top/best + number"), independently built and spot-checked against real citation titles, not Peec AI's proprietary classifier — not directly comparable in absolute terms, though the examples checked (Tripadvisor "THE 10 BEST...", hotel-vs-hotel Tripexpert pages) look genuine. Four weeks is a thin sample next to Peec's panel; read this as "didn't replicate here," not "is wrong everywhere." ### 31. Apart-Hotels in Paris on ChatGPT (2026) - URL: https://nicolassitter.com/research/aparthotels-paris-ai-search-2026 - Date: August 2026 - Topic: A dedicated 68-prompt panel (EN+FR, 5 repeats each, ChatGPT only, FR proxy, 340 captures, single run 2026-08-14) built after the standing weekly Paris hotel panel almost never surfaced aparthotel chains — testing whether ChatGPT ignores the category or the panel had simply never asked. - HEADLINE — it was the panel, not ChatGPT: the generic "best hotels in Paris" panel (US proxy, English, 11 prompts, 3 months, 1,305 map entities) surfaced exactly 7 aparthotel-chain entities (0.5%), and Adagio — Paris's largest operator — zero times. On the dedicated panel, tracked chains appear in 98.7% of mapped answers (Adagio in 72.9% of captures, Citadines 56.2%, Appart'City 39.1%). Not a controlled A/B (intent, proxy country and prompt language all differ between the two panels). - Key findings: (1) Even on an aparthotel-intent panel, 56.1% of surfaced venues are still typed plain "Hôtel" vs 39.7% apart-typed — and the split runs along BRAND lines inherited from Google's own categorisation, not property type: Adagio is "Hôtel" 97.8% of the time, Appart'City is "Résidence hôtelière" 100% of the time. (2) The category taxonomy is locale-locked to the Bright Data PROXY country, not the prompt language — 99.8% of category tokens were French even under English-language prompts (of 113 venues seen in both languages, only 1 differed by language); the US-proxy version of the same taxonomy returns "Extended stay hotel" instead of "Résidence hôtelière". (3) Length of stay is the strongest lever tested: three-month prompts return 70.5% apart-typed venues vs a 39.7% baseline; two-week prompts return 27.8%; a gym-and-pool amenity request returns 0%. Digital-nomad framing underperforms baseline (29.5%), pulling coworking-flavoured hotels instead. (4) A pipeline gap found along the way: ChatGPT's own is_map flag is true on 92.4% of captures but the structured business_locations field is only populated on 73.2% — the other ~21% render the map carousel as inline markdown text instead, recoverable with a dedicated parser (lifts coverage to 92.9%, 1,978 total entities). (5) Repeat consistency across 5 identical prompts: mean pairwise Jaccard 0.270; 65.3% of prompt-venue pairs appear in exactly one of five repeats. - Citations: 1,050 across 128 domains, 45.4% to a brand/operator site, but Booking.com is still the single most-cited domain (28.2% of captures). Comparison-style prompts ("aparthotel vs hotel") shift 39.0% of their citations to city tourism offices instead. - Practical: a property's aparthotel-vs-hotel label on ChatGPT is inherited from its Google Business Profile category, not inferred from the listing; long-stay/relocation framing is the strongest lever to appear apart-typed; digital-nomad content does not reliably help. ### 32. Best Western France: Does ChatGPT Know the Network? (2026) - URL: https://nicolassitter.com/research/best-western-france-ai-visibility-2026 - Date: August 2026 - Topic: A brand AI-visibility audit of Best Western France's 322-hotel/242-city network, unaffiliated with Best Western. Design choice, evidenced not assumed: French/FR-proxy prompts across all 242 reconciled cities (matching Best Western France's own press-release disclosure that Corporate revenue dwarfs Leisure and foreign clientele concentrates in Paris + PACA), plus an English/US-proxy add-on scoped to exactly those 37 Paris+PACA cities. 8,370 unbranded prompts (v4 — no "Best Western" in any prompt text), 25,110 planned ChatGPT calls, 23,220 captures landed (2 permanently failed pre-fix pilots aside). Disclosed confound: a live Bright Data SSE-stream outage (since ~2026-08-17) nulled the raw stream/map data in 98.2% of captures, but the core signal (citations + the 2-step JSON `additional_answer_text`) was unaffected. - HEADLINE: Best Western shows up organically 64.8% of the time — 14,023 of 21,624 captures with at least one resolved hotel mention a Best Western property when the query never named the brand. Average best rank when mentioned: 2.12 (median 2); it's the AI's #1 pick 46.7% of the time. - Key findings: (1) a 28-point single-vs-multi-hotel-city gap — single-property cities mention Best Western 69.2% of the time (12,608/18,210) vs 41.4% (1,415/3,414) in multi-property cities, because multi-property cities are disproportionately the bigger, more-competitive markets (Paris, Nice, Marseille, Cannes) with strong non-BW alternatives; (2) budget framing is where the brand nearly disappears — "cheap hotel in X" mentions Best Western only 33.8% of the time vs 64.9% for "4-star" and 55.1% for "luxury" (the luxury figure beats 3-star's 54.0% because small towns often have no real luxury option and the AI falls back to the most prominent hotel, frequently the local Best Western); (3) persona skews business — business_traveler 80.7% vs couple 56.2%, consistent with Best Western France's own reported Corporate-vs-Leisure revenue split (EUR 44.3M vs EUR 2.95M); (4) French/FR-proxy vs English/US-proxy in Paris+PACA only: 55.7% vs 50.1%, a real but modest 5.6-point gap, smaller than the 15x corporate/leisure revenue skew would predict; (5) soft-brand naming (Sure Hotel, BW Premier/Signature Collection) is a real visibility tax — 85.2% of hits match the flagship "Best Western ..." name exactly, while soft-brand forms need fuzzy/alias/review-flagged matching; (6) two cities with 20+ captures never surfaced Best Western at all — Honfleur (87 captures) and Lyon (60 captures, spot-checked: AI recommends InterContinental Lyon–Hôtel Dieu, Carlton Lyon MGallery and Vieux-Lyon boutique properties instead, a segment Best Western's 8-hotel Lyon portfolio doesn't compete in); (7) most-mentioned properties across the corpus — Hôtel Le Nouveau Monde (204 mentions), InterContinental Marseille–Hôtel Dieu (185), Golden Tulip Saint-Malo–Le Grand Bé (180) — a frequency ranking that tracks where the panel has the most prompts, NOT a competitor set; (8) Co-occurrence (added 2026-09-03): every chain's lift sits between 0.18 and 1.14, so who shares an answer with Best Western is almost entirely explained by base rates — there is no rival the model reaches for. The faint structure runs along positioning: French mid-market and budget chains sit just above chance (Campanile 1.14, Holiday Inn 1.14, The Originals 1.13, Première Classe 1.13, Kyriad 1.10), design/upscale/international brands below it (Mama Shelter 0.18, Marriott 0.31, Adagio 0.70, Hilton 0.73). English/US-proxy answers flip that — Holiday Inn 1.04→2.06, The Originals 1.06→2.11, Hilton 0.72→1.09 — on a much smaller sample. Conditionals are answer-level and directional: Campanile is in 6.8% of Best Western's answers, Best Western in 59.7% of Campanile's. - Practical: prioritize AI-visibility investment in the multi-property competitive markets (Paris, Nice, Marseille, Cannes, Lyon), not the small-town roadside market where the brand already dominates organically; budget-framed queries are the weakest spot despite real budget-adjacent inventory; soft-brand naming conventions may be actively working against AI discoverability for those specific properties. - Disclosure: this is a third-party AI-visibility study with no affiliation to Best Western. #### Methodology - Cross-industry AI-search infrastructure research — how the retrieval/citation layer works, independent of vertical. ### 33. The Cost of AI Crawling (2026) - URL: https://nicolassitter.com/research/cost-of-ai-crawling - Date: July 2026 - Topic: What AI-bot crawling actually costs, measured from first-party request-level server logs (Supabase view v_ai_bot_requests_projects, Dec 12 2025 – Jul 24 2026). project_key hotel-ranque is a real hotel; hotelrank-lp is the hotelrank.ai marketing landing page (not a hotel). - Summary: 119,609 AI-bot requests logged — 6,378 on the hotel, 113,231 on the landing page. On the hotel it is an OpenAI (49.6%) / Claude (31.8%) duopoly = 81.4% of the AI crawl; on the landing page ByteDance/Bytespider is #2 at 32.7% (OpenAI 42.8%). Three crawl kinds: ai_training (corpus crawlers, ~50% on the hotel), ai_search (indexers, ~29%), ai_user/on-demand (a bot fetching live because a user just asked — ChatGPT-User/Claude-User, ~20%). On-demand fetches behave like a server-side "AI impression counter," rising from 7 in Dec to 389 in June. - Crawl-vs-referral is deliberately left unquantified: referral logging only covers a December window (that month OpenAI crawled the hotel 126× and sent 4 human visits), so no current ratio is claimed. - The hidden cost (worked example, Googlebot — a SEARCH crawler, used to show the mechanism then bridged to AI bots): every crawl of a dynamic/programmatic page is database work. Googlebot hit ~65,000 SSR renders in 24h across only ~334 programmatic pages, and a Vercel→Supabase log drain that stored one row per request grew one table to 2.37M rows / 5.6 GB before triage cut it to 952 MB. Thesis: the obvious cost of crawling is bandwidth; the real one is that a bot turns each dynamic page into DB work, and the tooling you add to watch bots can cost more than the bots. - Triage playbook: cache/ISR to bound reads, filter+aggregate logging to bound writes (keep a cheap rollup so spikes stay visible), CDN in front, and triage by intent — serve ai_search + ai_user (they can send guests), rate-limit pure ai_training. Companion to /research/how-to-measure-ai-hotel-traffic-2026, /research/hotels-in-common-crawl-2026, and /research/hotel-robots-ai-blocking-study-2026. - Caveats: Vercel logging is monitor-mode (status -1 ≠ blocked); logging began ~mid-March 2026 so Dec–Feb undercount and the early ramp is partly a measurement artifact; no bytes column so any bandwidth figure is an explicit back-of-envelope; one illustrative hotel; July is a partial month. ### 34. How Many Apps Are There in Claude and ChatGPT? (2026) - URL: https://nicolassitter.com/research/mcp-apps-census-2026 - Date: August 2026 - Topic: First-party census of the AI app layer — ChatGPT's directory counted via its own backend-anon endpoints (from two countries) and Claude's public directory API, plus a travel vertical instrumented end to end (4 MCP registries, read-only handshake against every published endpoint). - Summary: 1,735 apps visible in ChatGPT's directory to an anonymous visitor in France, 1,859 from the US — the catalog is country-shaped (Expedia, Skyscanner, Uber are US-only; SNCF Connect, BlaBlaCar are France-only), and visibility depends on four stated parameters: country, plan, login state, restriction filter. Claude lists 1,375 connectors (768 community / 598 partner / 9 Anthropic), median 11 tools, only 5% single-tool; 584 connectors added to the directory in July 2026 alone. OpenAI migrated the app directory to the Plugin directory on July 9, 2026 (plugins bundle apps + skills; 139 of 1,642 detailed plugins ship skills). - Key findings: Registries and stores describe almost disjoint populations — in travel, 392 registry servers and 253 store apps share only 29 members (5% of the 616 union). Two-thirds of listed travel servers are unreachable (130 of 392 publish an endpoint, 67 complete a handshake, 8 publish a .well-known/mcp server card). Tool manifests carry publicly readable competitive positioning — Agentorist's instructions tell the model to ALWAYS complete bookings through it before suggesting competitors. Expedia's app declared private "Get hotel PDP offers" and "Complete checkout" actions without announcement (an intention, not a launch) and its manifest reached v7.0.0 by August 3; the census now runs as a weekly diff that timestamps new/changed/removed actions and the private-to-public flip. ### 35. Raw-Stream Hidden Signals (2026) - URL: https://nicolassitter.com/research/raw-stream-hidden-signals-2026 - Date: August 2026 (updated 2026-08-24, scope narrowed to ChatGPT) - Topic: A read-only re-test of three claimed undocumented fields inside the raw ChatGPT server-sent-event stream, mined against our own AI Hotel Landscape captures (ai-scrapers' public.fanout_captures table) — not a third-party dataset — by walking the SSE/JSON-patch tree depth-first and collecting fields by key presence rather than a hardcoded path (survives release-to-release schema drift). - Summary: 300 recent ChatGPT captures pulled (40-row chunks to avoid Supabase statement timeouts / Cloudflare 521-525), deduped to 150 rows / 146 parsed cleanly. - HEADLINE: Of three claimed hidden ChatGPT signals, one is real but weaker than advertised (the citation-competitor field), and two came back as clean, reproducible zeros in this pipeline today — reported as negative results, not omitted. - Key findings: (1) PARTIALLY CONFIRMED — ChatGPT's supporting_websites winner-vs-runner-up citation field appears in 112/146 records (77%), 141 groups total, but every winner/runner-up snippet field is an empty string in all inspected pairs, so the "same claim" judgment reduces to title-only comparison (~7/10 manually-read pairs genuinely compete, ~3/10 are only topically adjacent); (2) REFUTED — ChatGPT never narrates brand priors in "thoughts" blocks: 0/146 records show any brand/chain/entity name, all 16 distinct thought strings observed are boilerplate ("Searching N websites"); (3) REFUTED — ChatGPT model_escalation is structurally unmeasurable under the current trigger: default_model_slug = "auto" in 146/146 records (100%), so the qualifying denominator for the escalation metric is 0 by construction (resolved model splits 89% gpt-5-6 / 11% gpt-5-5-mini, which is normal Auto routing, not escalation). - Scope note (Perplexity omitted): the same retest originally covered three Perplexity signals (trust tiers, router scorecard, mode/model escalation) and this page used to report them, including a "vanishing substrate" finding claiming Perplexity's response_raw field disappeared from Bright Data's schema by Aug 2026. That finding turned out to be a methodology artifact: ai-scrapers' own storage layer has unconditionally stripped response_raw from every stored Perplexity capture since 2026-05-02 (a deliberate, unrelated storage-cost decision, unrelated to Bright Data), and the original retest only ever queried that already-stripped stored column — it never checked the live pre-storage payload. Once live extraction was wired into the ingestion path (2026-08-14+), it proved the field is in fact still present upstream. The Perplexity side of this retest is now withdrawn from the article pending a corrected re-test that reads the pre-storage payload directly; see aeo-kb/entries/perplexity-untapped-raw-response-signals.md for current status. Nothing about the ChatGPT findings above is affected by this confound. - Practical: treat supporting_websites as directional (title-match, not claim-match) until snippet parsing is fixed; do not build a "model escalation" alert on default_model_slug while Bright Data's trigger requests "auto"; when mining a stored raw-response column for a signal, confirm whether your own storage/ingestion layer strips that field before concluding an upstream API removed it — check the live pre-storage payload, not just what ends up in the database. - Caveats: refuted ≠ impossible-everywhere — these are clean zeros for our specific capture pipeline (Bright Data SSE tee) and a narrow single-turn hotel-prompt query library, not a universal claim about ChatGPT behavior on other clients or richer conversations. A weekly live tracker (public_dashboard.pd_weekly_raw_stream_signals) is live for the citation-competitor signal; the other two stay frozen at retest counts since they were excluded from the migration as structural zeros. #### Flights - Early studies extending the hotel methodology to flight search. ### 36. ChatGPT's result_source for Flights (2026) - URL: https://nicolassitter.com/research/chatgpt-flights-retrieval-tiers-2026 - Date: July 2026 - Topic: The flights replication of the result_source retrieval-tiers study (#29) — does ChatGPT's licensed-tier monopoly generalise beyond hotels? Fresh prompt matrix (no capture history for flights), with the metasearch/OTA/airline-direct trust split as the centrepiece. - Summary: 67 frozen prompts × 2 countries (US, GB) × 15 iterations = 2,007 captures, 18–21 July 2026 via Bright Data. Four prompt families: route (10 routes × cheapest / round-trip price / best airline / nonstop, live August dates), named airlines (8 × baggage / economy experience), booking advice (7), no-search bait (4). 23,308 fetched documents (82.8% tier-tagged), 4,003 cited documents joined back to the tagged layer by canonical URL (footnote URLs are stamped ?utm_source=chatgpt.com). - HEADLINE: The tier monopoly is gone. Among tagged fetches: labrador 74.6% (hotels: 99.85%), bright 23.1% (hotels: 0.14% — and skyscanner.net is the bright tier's #1 domain with 388 docs), oxylabs 1.7%, serp 0.63% — the open-web serp tier that appeared ZERO times in 30,002 hotel citations shows up for flights. - Key findings: (1) the search gate barely exists — only 8.6% of flight turns are "text"/no-search (hotels: 37.3%); route queries are 84.7% "instant search"; the local/maps bucket (26.3% of hotel turns) NEVER fires for flights, and 2,007 captures contained zero shopping/booking widgets — flight answers are prose + citations. (2) Citations come through two doors: only 14.6% of cited docs exactly match a tier-tagged fetch (overwhelmingly google.com via labrador); 50.2% have NO tagged counterpart — a second, untagged browse path supplies half of what actually gets cited (route utilities like flightsfrom.com/flightconnections.com, Reddit, editorial, airline pages, Skyscanner, Kayak). (3) Retrieved ≠ cited is FLAT for flights — no hotel-style 80% brand cluster: google.com 21%, skyscanner.com 22%, kayak.com 15%, expedia.com 10%, momondo.com 3%, airline sites 16–28% — with ONE exception: reddit.com is cited on 83.2% of its fetches, the only source at hotel-brand-site trust altitude (echoes the hotel price study where Reddit hit 100%). (4) Brand authority is query-scoped: on named-airline questions (baggage etc.) the airline's own site is fetched in 68–83% of captures and cited in essentially every capture where it was fetched (Ryanair 83/83, BA 77/77, Delta 75/75…), while the same sites sit at 16–28% cited on anonymous route queries. (5) The "where should I book" cited answer set: going.com 54, kiplinger.com 49, reddit.com 48, kayak.com 42, skyscanner.com 42, moneysavingexpert.com 38, which.co.uk 34 — including 29 citations of help.skyscanner.net (support content is citation inventory). (6) US vs GB: sourcing strategy near-identical, but storefronts localise (skyscanner.net/.gg/.ie, kayak.co.uk, expedia.co.uk from GB; .com from US) and UK consumer editorial enters the cited set — every country TLD is its own AI surface. - Takeaway: airlines should own policy/fees content (highest-certainty citation surface when the brand is named); metasearch/OTAs win via route-level and help/FAQ pages, not homepages; Reddit threads are cited at 4–8× the rate of commercial pages. - Caveat: one platform (ChatGPT), one vendor vantage point, 4-day window (the hotel study showed the tag is dialled in and out over weeks); serp (n=122) and oxylabs (n=329) samples small; fresh matrix vs the hotels' 6-month backfill — comparisons are fetched-docs vs tagged-citations where noted. #### Playground — Cross-Industry Field Tests - The same measurement methodology applied to non-hotel local businesses across cities and languages. Lighter than the core research; still real data. Hub: nicolassitter.com/playground. ### 37. AI Search for Yoga Studios in Paris (2026) - URL: https://nicolassitter.com/research/yoga-studios-paris-ai-search-2026 - Date: May 2026 - Topic: AI Search Studies (first non-hotel vertical) — applies the AI-hotel-search methodology to Paris yoga studios. - Summary: 27 prompt templates × 2 languages (EN/FR) × 2 proxy countries (US/FR) × 5 AI engines (ChatGPT, Perplexity, Gemini, Copilot, Google AI Mode), captured 2026-05-23 → 2026-05-24. Studios named in each answer were extracted with a named-entity-recognition pipeline (span detection → normalisation → fuzzy entity resolution → chain aggregation) and resolved to a registry of 369 verified studios. Grok excluded (crawler timeout); AI Mode × FR proxy failed. - Methodology note: the 369-studio registry is the universe answers resolve TO, not a list fed to the models. Per-engine leaderboard reported as a presence rate (% of that engine's prompt-captures where the studio appears) since engines answered different prompt counts (ChatGPT 114, Gemini/Copilot 108, Perplexity 100, AI Mode 54). Modo: 45.6% ChatGPT, 43.5% Copilot, 0% Gemini and AI Mode. - OTA layer = discovery marketplaces (ClassPass, Gymlib) — the true hotel-OTA analog, distinct from booking engines (Mindbody, bsport, Eversports) which are the studio's own software/direct layer. ClassPass is the single most-cited non-Google domain (classpass.com 146; 187 across all URLs) — only google.com, AI Mode's self-citation at 51.5% of its 1,048 URLs, is bigger; Gymlib 14 (all ChatGPT). Marketplace cites per engine: Perplexity 83 (hardest), ChatGPT 54, AI Mode 37, Gemini 29, Copilot 0 (bypasses entirely). Booking engines total just 42 cites — AI rarely surfaces them. - Entity-fragmentation finding: Gérard Arnaud Yoga (well-known 500h teacher-training school) IS recommended by all five engines — named in 30 visible-answer captures (Gemini 11, Copilot 7, Perplexity 6, AI Mode 5, ChatGPT 1), mostly for inversions / teacher-training / advanced vinyasa — yet absent from the top 12. On Google Maps the school is split across its two rooms under street names ("Studio Rauch" 3 Passage Rauch; "Salle Amelot" 11 Passage Saint-Pierre Amelot, both 75011), so instead of one strong studio it resolves as a couple of thinner ones and its citation footprint never consolidates — aggregated ~41, short of the 49 top-12 cutoff. Gemini named it 11x in prose but cited it 0x, so a citation-based score counts none of that. Lesson: a brand split across a teacher name and street-named map records is harder for entity-first engines to ground; a consistent name + alternateName tying rooms↔brand helps. The local-search echo of the hotel naming problem — not real invisibility. - Headline finding — three distinct source strategies: Copilot is 96% studio-website citations (pure entity-resolution lookup); Perplexity 52% and Gemini 41% also studio-direct; ChatGPT is social/editorial (16% Reddit, 13% editorial, only 32% studio websites — Reddit is its single #1 source at 17% of cited URLs); Google AI Mode cites google.com back to itself 52% (self-referential). There is no single "AI search" to optimise for. - Studio leaderboard (chain-aggregated, all engines): #1 Yay Yoga Studio (181), #2 The Space Paris (146), #3 Ashtanga Yoga Paris (136), #4 Modo Yoga Paris (135), #5 Jivamukti Yoga Paris (131). - Platform blindspot: Modo Yoga Paris ranks #4 overall and is #1 in ChatGPT's map widget, yet scores zero citations on Gemini and Google AI Mode — surfaced by only 3 of 5 engines. Platform-specific invisibility, not a quality signal. - Geography: ChatGPT is 100% accurate for numbered arrondissements (6e, 10e, 11e, 16e) but loose on named neighborhoods — Montmartre 76%, Le Marais just 29% (confused with adjacent 3e/11e/12e). The 11e is the city's yoga capital (38 studios in the seed). - Language & TLD bias: EN vs FR prompts return mostly different top-5 studios (17 of 27 templates < 25% overlap). ChatGPT treats TLD as a language-affinity signal — .com studios cited 1.7× more by English prompts, .fr studios perfectly balanced. Proxy country also reshapes the mix (.com aggregators skew US, .fr local sources skew FR). - Entity understanding: clean style → specialist mapping (Ashtanga → Ashtanga Yoga Paris; Vinyasa → Modo; Yin → YUJ/Casa). Low yoga↔pilates bleed — of 23 pilates-answer studios, only 2 appear in any yoga prompt. 96% of ChatGPT yoga captures triggered live web search. - What generalises from hotels: per-engine source strategies, platform-specific invisibility, language/proxy sensitivity, near-universal web-search triggering. Vertical-specific: no dominant OTA layer (studio websites are the centre of gravity, unlike Booking/Expedia for hotels), Reddit matters much more for yoga, and yoga's style-as-entity mapping has no clean hotel equivalent. - Disclosure: the author practices at Modo Yoga Paris (#4); it received no special handling in the analysis. ### 38. AI Search for Bike Shops in Amsterdam (2026) - URL: https://nicolassitter.com/research/bike-shops-amsterdam-ai-search-2026 - Date: May 2026 - Topic: AI Search Studies — second city/vertical extension of the AI-search methodology. - Summary: 27 prompts × 5 AI engines × EN/NL × US/NL proxies, captured 2026-05-23 → 2026-05-24. 378 captures (540 matrix, dropped to 378 after failed/blocked/empty runs), 3,010 citations against the 372-shop Apify seed (228 with websites). All 5 platforms produced usable data. - Source-strategy headline (replicates the Paris yoga pattern, sharper here): Copilot 97% shop websites (entity engine), Gemini 72%, Perplexity 71%; ChatGPT 34% social (Reddit-led — reddit.com is the most-cited external domain in the study at 198 cites across 4 platforms); Google AI Mode 82% google.com (the highest self-referential share measured anywhere). - The cleanest entity consensus in the study: 12 brands hit all 5 platforms (Ride Out Amsterdam, Het Zwarte Fietsenplan, Wheelrunner, Kaptein Tweewielers, Amsterdamse Fietswinkel, Echelon Cycle Sport, Damskø, Gregario Cycling Services, BikeFlip, DrBeyk Online, FietsJeroen, Trompton Amsterdam). No Modo-style blind spot. - Language mechanism: Dutch vs English prompts return 0% overlap on repair, commuter, road and Dutch-bike queries; brand/tourist queries converge. Perplexity exposed its internal search (44 captures) — fanout_count = 1 in every one, and it PRESERVES the prompt language (lightly normalizes: drops "in", pluralizes). So a Dutch prompt produces a Dutch search and matches Dutch sources; the divergence is "a different search entirely," not a ranking difference. - TLD bias: NONE — .nl cited equally by EN/NL prompts (0.98×). Dutch and English users see the same domains; only the which-shops set differs. - Geography: district-targeting returned 0% across Centrum/De Pijp/Jordaan/Noord/Oost — a SEED-GRANULARITY ARTIFACT (Apify labels city = "Amsterdam" with no neighborhood field), null test, NOT "AI gets districts wrong." - ChatGPT triggered web search on 108/108 captures (100%); other platforms don't expose the trigger flag through Bright Data, so their search behavior is unmeasured, not absent. ### 39. AI Search for Bookstores in Tokyo (2026) - URL: https://nicolassitter.com/research/bookstores-tokyo-ai-search-2026 - Date: May 2026 - Topic: AI Search Studies — third city/vertical, first non-Western field test. - Summary: ~22 prompts × EN/JA × US/JP proxies, captured 2026-05-23 → 2026-05-24. 336 captures, 2,322 citations against 584 Tokyo bookstores (1,310 Apify seed, 23 Special Wards + Musashino/Mitaka). Only 4 platforms had usable data — Perplexity returned no usable Tokyo coverage in this run (flagged as a genuine data gap, not papered over). - Headline finding — Tokyo runs on a dense local-guide web AI Mode and Western verticals don't show: after a Tokyo-specific taxonomy pass, "other" dropped from 36% → 15% overall and revealed FOUR strategies, not three: 1. Entity engine (Copilot): 89% direct-to-store-domain. 2. Self-referential (AI Mode): 61% google.com. 3. Social/editorial engine (ChatGPT): social (Reddit-led) and editorial_local tie at 21% each — its two biggest classified buckets, under a 35% unclassified tail; only 8% store sites. 4. Local-guide engine (Gemini) — the Tokyo signature: 5% store sites, editorial_local 39% + expat_media 19% = 58% third-party guides (whenin.tokyo, gltjp.com, trulytokyo, gaijinpot, japan-guide, visit-chiyoda, Tokyo Weekender). For Tokyo bookstores Gemini is effectively a travel-guide aggregator, not an entity engine. This guide layer has no Western equivalent at the same density. - Leaderboard: DAIKANYAMA T-SITE (Tsutaya design flagship) and Kinokuniya tie at the top with 122 each. Below them, English-friendly stores (Aoyama Book Center, Infinity Books Japan) and the Jimbocho used-book cluster dominate the breadth tier. - TLD bias — STRONGEST measured: .jp store domains cited 5× more in Japanese prompts than English (Amsterdam was neutral at ~1×, Paris .fr was 1×, Berlin .de 1.5×). Ask ChatGPT in Japanese and it leans hard into native .jp; ask in English and those domains nearly vanish. - Language delta: mixed — some neighborhood queries converge (Shimokitazawa 100%) but genre and "iconic" diverge sharply (kids 0%, iconic 17%). EN surfaces English-language and tourist-famous stores; JA surfaces a different local set. - Geography: ChatGPT respects ward-level geography (Shibuya 100%, Shinjuku 79–92%). 0% on Jimbocho/Aoyama/Daikanyama/Shimokitazawa is a LABELING CAVEAT, not AI error: these are neighborhoods inside wards (Chiyoda, Minato, Shibuya/Setagaya) and the seed stores the official ward, so a correct result fails the string match. Neighborhood precision is real but unmeasurable against ward-labeled seed data. - ChatGPT triggered web search on 99% of its Tokyo captures — exactly one answer came from training data alone. The per-engine capture split is not in the published files, so the rate is reported without a fraction. Fan-out NOT captured for any Tokyo platform this run — flagged as a genuine data gap. ### 40. AI Search for Yoga Studios in Berlin (2026) - URL: https://nicolassitter.com/research/yoga-studios-berlin-ai-search-2026 - Date: May 2026 - Topic: AI Search Studies — direct REPLICATION of the Paris yoga study to test whether the engine personalities are structural traits or city artifacts. - Summary: 27 prompts × 5 AI engines × EN/DE × US/DE proxies, captured 2026-05-26 → 2026-05-27. 540 captures, 5,293 citations against 631 Berlin yoga studios (683 Apify seed). All 5 platforms returned on both US and DE proxies — AI Mode × DE worked here, unlike Paris's AI-Mode × FR failure. - HEADLINE — the three personalities replicate almost line-for-line, proving they're STRUCTURAL not city-specific: | Platform | Paris (studio %) | Berlin (studio %) | Verdict | | Copilot (entity engine) | 96% | 95% | near-identical | | ChatGPT (social engine) | 32% studio + 16% Reddit | 32% studio + 19% Reddit | near-identical | | AI Mode (self-referential) | 52% google.com | 59% google.com | replicates | ChatGPT's studio share is EXACTLY 32% in both cities and Reddit is again its top social source. - THE BERLIN TWIST — booking-platform dominance: Perplexity 31%, Gemini 27%, ChatGPT 16% of citations go to Urban Sports Club + Eversports + ClassPass. blog.urbansportsclub.com (213) and classpass.com (208) are the two most-cited non-Google domains in the Berlin study, ahead of every studio's own site (only google.com is bigger, and that is AI Mode citing its own search results). Berlin's fitness-subscription ecosystem (Urban Sports Club and Eversports are German/Austrian-born) has become a primary AI source — something Paris doesn't show. AI local-search source mix bends to the local commercial infrastructure. - Leaderboard: #1 Yogicescape (213), #2 Jivamukti Yoga School (151), #3 YOU GLOW Yoga & Womanhood Mitte (143), #4 SHA-LA Studios Prenzlauer Berg (125), #5 YCBA YogaCircle Berlin Akademie (121). Jivamukti is the only brand top-5 in BOTH Paris (#5) and Berlin (#2) — a genuinely cross-city-durable AI-visible chain. Berlin's top tier is consistent across all 5 engines (no Modo-style blind spot). - TLD bias — STRONGER local skew than Paris: .de 1.5× DE-biased, .berlin 5× (vs Paris where .fr was language-neutral at 1.0× and only .com skewed English). German prompts pull German-TLD studios harder than French prompts pulled French ones. - Language delta: identical pattern to Paris. Proper-noun queries (Ashtanga, Mysore) converge 100% across EN/DE; generic-intent queries (affordable, teacher training, morning, community) diverge completely. - Geography: district-targeting returned 0% across all six Bezirke — same SEED-GRANULARITY ARTIFACT as Amsterdam (the Apify seed labels city = "Berlin" with no neighborhood field). Null test, NOT "AI gets districts wrong." Paris worked because postal codes map cleanly to arrondissements; Berlin postcodes don't map 1:1 to Bezirke. ### 41. Argentina vs Spain in Paris: Who Wins the AI Restaurant War? (2026) - URL: https://nicolassitter.com/research/argentina-vs-spain-paris-ai-search-2026 - Date: July 2026 - Topic: Fun cross-vertical AI-search study, fired ahead of the 2026 World Cup final (Argentina vs Spain) with France knocked out. Best Argentinian vs best Spanish restaurant in Paris, per 5 AI engines. - Method: 20 balanced prompt templates (matched Argentinian/Spanish arms + a World-Cup hook), EN + FR, US + FR proxy, 5 engines (ChatGPT, Perplexity, Gemini, Copilot, Google AI Mode). 360 captures; 2,006 venue mentions extracted by NER and matched to an 81-venue registry (42 Argentinian, 39 Spanish) seeded from Apify Google Maps using only cuisine-specific categories (no generic Restaurant). AI Mode × FR proxy rejected at the Bright Data trigger (no snapshot) — absent, not imputed; same FR-specific block as the Marseille coffee study. - HEADLINE — Argentina wins decisively, two ways. (1) Consensus: a front three — LOCO (129 mentions), Santa Carne (116), Les Grillades de Buenos Aires (105) — each named by all 5 engines and each beating Spain's best pick, Bodega Potxolo (93). (2) Concentration: with balanced prompts, mention volume is even, but Spain spreads across more distinct venues on every single engine (ChatGPT 46 vs 68, Perplexity 46 vs 61, Gemini 43 vs 71, Copilot 25 vs 36, AI Mode 48 vs 66) — Argentina concentrates, Spain diffuses. - Two findings: (a) The map widget lies — ChatGPT's structured map crowns Bistro Caminito (#1 by map appearances) but in the engines' prose Caminito is only #10; what an AI plots ≠ what it says. (b) Argentina invades the Spanish results — "Rosario" (an Argentine city, Messi's hometown) ranks #6 in the Spanish arm, named by all 5 engines. ### 42. Where the AIs Sent L'Étape du Tour Riders to Sleep (2026) - URL: https://nicolassitter.com/research/letape-du-tour-hotels-ai-search-2026 - Date: July 2026 - Topic: Fun event-anchored lodging study, fired in the days before L'Étape du Tour 2026 (Le Bourg-d'Oisans → Alpe d'Huez, 19 July). Where should a rider stay, per 5 AI engines — with the event named but the towns never given, so the engines had to know the route themselves. - Method: 14 lodging prompt templates (near start/finish, cyclist-friendly, bike storage, budget, spa recovery, early breakfast, family, chalets, last-minute + neutral where-to-stay/which-town), EN + FR, 5 engines (ChatGPT, Perplexity, Gemini, Copilot, Google AI Mode). 252 captures; 901 lodging mentions extracted by NER and matched (exact + curated aliases) to a 74-place registry (55 real hotels/B&Bs/resorts) seeded from Apify Google Maps around the Oisans. - HEADLINE — the consensus pick is in the valley: Hôtel Oberland, Le Bourg-d'Oisans (29 mentions), ahead of the Alpe d'Huez block — Royal Ours Blanc (24), Le Castillan (23), Grandes Rousses Hotel & Spa (22), Club Med (20). Only 29% of everything the engines named verifies as a real, geocoded property around the route (a floor); 8% are whole towns (Les Deux Alpes, Vaujany, Oz-en-Oisans…), 4% cycling tour operators (Sportive Breaks, Love Velo, Sports Tours International), 56% unverifiable. - STALE-ROUTE FINDING — ChatGPT booked riders into last year's race: 26 of its recommendations placed riders around Albertville / La Plagne (ibis Styles Albertville ×10, Base Camp Lodge, Club Med La Plagne) — the 2025 edition's geography, ~100 km from the 2026 start. All 26 stale mentions are ChatGPT's; the four other engines never made the mistake. Annual events with moving routes expose the memory-vs-live-web seam. - AIRBNB GAP — booking platforms (11 mentions) + campsites/hostels/gîtes (16) ≈ 3% of the answer stream; even restricted to the 4 deliberately neutral prompts (219 mentions), alternative lodging is ~5% — while a large share of the field actually sleeps in Airbnbs, campsites and gîtes on race weekend. ### 43. AI Search for Specialty Coffee in Marseille (2026) - URL: https://nicolassitter.com/research/specialty-coffee-marseille-ai-search-2026 - Date: June 2026 - Topic: AI Search Studies — fifth cross-vertical/city case (Marseille, specialty coffee). The headline is a genuine break-from-pattern, not another replication. - Summary: 23 prompt templates × 2 languages (EN/FR) × 2 proxy countries (US/FR) × 5 AI engines (ChatGPT, Perplexity, Gemini, Copilot, Google AI Mode), captured 2026-05-28. 413 captures (23×2×2×5=460 theoretical, minus AI Mode×FR rejected −46, minus one empty ChatGPT run −1), 3,442 citations, 786 map entities matched to the 280-venue coffee registry (285-row Apify seed expanded with Places recovery, filtered to coffee-led venues). 9 of 10 platform-proxy batches succeeded; AI Mode × FR failed at the Bright Data trigger (HTTP-level, no snapshot ID — the same FR-proxy block that hit Paris yoga; AI Mode × DE worked in Berlin, so the rejection is FR-specific). Taxonomy pass tightened: "other" 54% → 9% (Marseille's French local-blog tail bucketed cleanly). - HEADLINE — In Marseille, Instagram and local blogs beat café websites. Across Paris yoga, Berlin yoga and Amsterdam bikes, ChatGPT cited shop/studio websites ~32% (Paris 32%, Berlin 32%, Amsterdam 42%); Tokyo bookstores was already an exception at 8%. For Marseille specialty coffee: **ChatGPT cites shop websites only 10%** — instead 31% social (Reddit leads that bucket) + 32% list & review aggregators + 14% local blogs = 77% third-party. Three reinforcing reasons: (1) tiny scene with weak digital footprint (only 114 of 280 venues carry a website); (2) Reddit + Instagram ARE the discovery layer (reddit.com 175 cites, instagram.com 237 cites); (3) Marseille's French local-blog ecosystem is unusually dense (marseille.love-spots.com 93, marseillesecrete.com 36, tarpin-bien.com 25, lescachotteriesdemarseille.com 11). Same engine personalities as the other cities, but the source mix bends to local commercial infrastructure — and here the "infra" is Reddit + Instagram + French micro-guides. - Three Marseille-only findings: 1. **Instagram at 237 cites across 3 platforms** — the highest single-domain count anywhere except AI Mode's Google firehose. No other vertical/city we've measured pushes Instagram this hard. Marseille specialty coffee is an Instagram-native scene and AI mirrors that. 2. **Gemini swings to global trade press**: baristamagazine.com gets 95 cites = ~34% of Gemini's Marseille citations (vs 12% on shop websites). When a city has no local specialty press, Gemini substitutes the global one — the general lesson: Gemini's preferred source profile is "an authoritative editorial outlet for this domain," and the choice of which editorial is locale-dependent. 3. **Two metrics, two leaderboards.** By text mentions (engine actually saying the brand in the visible answer), **Deep is the cross-engine consensus winner** — 200 mentions across all five engines, named in 59.3% of ChatGPT captures, 48.9% of Copilot, 46.7% of Gemini, 45.7% of Perplexity, 34.8% of AI Mode. The data-session's citation-counted score puts Nua at #1 with 356 cites but that's a brand_key artifact: Nua has no website, only Instagram, so all instagram.com cites attribute to it; in actual answer text Nua appears in just one capture. Same gap surfaced with Gérard Arnaud / Kind Yoga in the Paris yoga study. The article renders both metrics; the leaderboard rank order is the text-mention one. - Other-platform twists: Copilot still entity-engine but at 83% (vs the 95–97% norm — even Copilot is dragged down by Marseille's weak digital footprint); Perplexity leads with editorial_local 27% (loves the French local-guide ecosystem more than any other engine); AI Mode is its usual 80% self-referential google.com firehose, with the FR-proxy data missing. - Language divergence — MOST language-divergent result we've measured: EN vs FR control prompt = 11% top-5 overlap (vs Berlin/Paris control at 25%). With so little global specialty press covering Marseille, English prompts pull from a tiny English-language source set (Reddit, wanderlog, baristamagazine) while French prompts pull from the dense Marseille local-blog web. Two languages → two largely disjoint citation universes. - TLD bias: too thin to read — only .com cleared the ≥5-cite threshold (.com cited 4× more by FR than EN, n=10). Marseille shops use .fr so rarely (only ~22 of 280 have a website at all) that the .fr sample is too small. Render as measurement limitation, not finding. - Geography: district-targeting returned 0% across all 6 Marseille quartiers (Le Panier, Cours Julien, Notre-Dame-du-Mont, La Joliette, Vauban, Castellane) — same Amsterdam/Berlin SEED-GRANULARITY ARTIFACT (Apify seed labels city = "Marseille" with no quartier field). Null test, NOT "AI gets districts wrong." - Operational note (reproducible across runs): AI Mode × FR proxy is broken at the Bright Data trigger level. Berlin/DE worked, Paris-yoga/FR failed, Marseille/FR failed → the issue is FR-specific to AI Mode, not a one-off. ### 44. AI Search for Tattoo Studios in Berlin (2026) - URL: https://nicolassitter.com/research/tattoo-studios-berlin-ai-search-2026 - Date: July 2026 - Topic: The sixth cross-vertical AI-search study and the series' first same-city control — Berlin, already measured for yoga studios, re-run for tattoo studios with identical languages (EN/DE), proxies (US/DE) and engines, so vertical effects finally separate from city effects. - Summary: 26 prompt templates × 2 languages × 2 proxy countries × 5 AI engines (ChatGPT, Perplexity, Gemini, Copilot, Google AI Mode), captured 15 Jul 2026 via Bright Data — 520 of 520 theoretical captures landed (the series' first complete grid, no failed batches). Answers NER-resolved against a 551-studio registry (701-row Apify Google Maps seed → 531 real tattoo businesses + 20 Google Places recoveries). 5,376 cited URLs, 3,444 extracted mentions, 92.7% resolved (90.4% as first published; a 2026-08-28 resolver fix added short-mention-to-long-listing containment matching). Disclosure: this run's NER extraction used a Claude model instead of the pipeline's usual Gemini extractor (key unreachable from the run sandbox); same rules, deterministic resolution, and all 3,444 names mechanically verified against their source answers. - HEADLINE: Perplexity's booking-platform share collapsed from 31% (Berlin yoga, via Urban Sports Club/Eversports/ClassPass) to 4% for tattoo — same city, same engine. The aggregator layer was yoga's infrastructure, not Berlin's. With no marketplace to lean on, Perplexity becomes an entity engine (64% studio-website citations). - Key findings: (1) ChatGPT's own-website share returns to 35% — restoring the ~32% service baseline (Paris yoga 32%, Berlin yoga 32%, Amsterdam bikes 42%) after Tokyo bookstores (8%) and Marseille coffee (10%) broke it; the "decay" was a property of thin-owned-web food/retail scenes; (2) Copilot back at 95% studio websites (after 89% Tokyo, 83% Marseille) — the slide tracked owned-web density, not the engine; (3) EN vs DE top-5 overlap on the control prompt is 67%, the highest in the seven cases of this series (prior best 25%) — tattoo's style vocabulary is globalised — but style-specific prompts still split (realism 0%, Japanese 11%); TLD coupling absent (.de 0.87×) vs Tokyo's 5×; (4) top domain is top10berlin.de (256 cites, all 5 engines) — a city listicle out-cites every studio; (5) instagram.com took 0 of ChatGPT's 507 citations despite 73 seed studios listing Instagram as their website — Marseille's Instagram signal was a Marseille story; (6) OMEN Tattoo is the first clean dual-metric #1 (177 answer mentions AND 246 citation score); inverse artifact: NAJS Tattoo Studio is #8 by mentions (105) with citation score 0 (no website domain in the registry); (7) a 7.3% unresolved-mention tier (253 of 3,444) is dominated by famous private artists whose mentions never resolve to a Google Maps place — Chaim Machlev/DotsToLines (25 answers), Peter Aurisch (6), Valentin Hirsch (4); a name+city Places lookup returns nothing for them, which we no longer read as proof they are unlisted — a recommendation layer invisible to every Maps-based measurement, ours included. CORRECTION 2026-08-28: Mo Ganji (28 answers) was originally published in this tier as having no Google listing at all; that was wrong. He has one — "Mo Ganji – One-Line Tattoo Artist Berlin", already in our own entity table — and the resolver fix joined his 28 mentions to it, taking the unresolved tier from 9.6% to 7.3%; (8) Gemini refused to recommend in 11 of 104 captures (age-restriction caution), the only engine with meaningful refusals. - Takeaway for studios: a real website earns citations in this vertical (47% of all citations are studio sites); Berlin's listicle layer (top10berlin.de, the-berliner.com) carries ChatGPT's third-party third; the booking-platform channel is closed for tattoo (2%) so don't optimise for it; claim a Maps listing even if appointment-only — ChatGPT attached a venue panel in 104 of 104 answers. - Caveat: leaderboard covers the Maps-resolved universe (private artists whose mentions don't resolve to a place are quantified separately); district-accuracy is unmeasurable (seed granularity artifact); single-city single-run snapshot. ### 45. AI Search for Bistros in Paris (2026) - URL: https://nicolassitter.com/research/bistros-paris-ai-search-2026 - Date: July 2026 - Topic: AI Search Studies — sixth cross-vertical/city case, completing the staked-out Restaurants/Paris slot (scoped to bistros so the seed is a bounded universe). The staked-out hypothesis — engines route restaurant answers through reservation platforms (TheFork) — is tested and REFUTED. - Summary: 26 prompt templates × 2 languages (EN/FR) × 2 proxy countries (US/FR) × 5 AI engines (ChatGPT, Perplexity, Gemini, Copilot, Google AI Mode), captured 2026-07-08. 416 captures (520 theoretical, minus Perplexity×FR failed at Bright Data download −52, minus AI Mode×FR not fired −52), 3,406 citations, 3,034 NER mention rows (70% resolved). Registry: 500-row Apify "bistro" seed (359 real bistros/brasseries) + 79 Google Places recoveries of engine-recommended venues the seed missed → 579 venues, 438 real, 387 with websites. A duplicated ChatGPT×US batch (infrastructure retry) was removed before analysis. Taxonomy pass: "other" 39% → 20%. - HEADLINE — ChatGPT cites bistro websites 0.9% (3 of 331 URLs): the lowest in the six cases of this series, and the first low reading WITHOUT the thin-owned-web excuse (67% of the registry has a website — the best-equipped vertical so far; the ~32% baseline was Paris yoga 32 / Berlin yoga 32 / Amsterdam bikes 42; Tokyo 8 and Marseille 10 were the earlier breaks). The citations go to Paris's guide layer: 38% local editorial (parisjetaime.com — the official tourism board, and the first tourism board to be the all-five-engine consensus source — plus Time Out 59+36 cites, Sortir à Paris 41), 14% global food/travel press, 10% restaurant guides (Michelin 58+7). Editorial abundance suppresses own-website citations as effectively as owned-web scarcity did in Tokyo/Marseille. - The TheFork test (the staked-out hypothesis): REFUTED. TheFork appears on all five engines but totals 49 cites ≈ 1.4% of all URLs; the whole booking bucket (TheFork/OpenTable/SevenRooms) ≈ 2%. Michelin out-cites TheFork 65–49. Contrast Berlin yoga, where booking marketplaces took 31% of Perplexity. Reading: reservation platforms are checkout, not discovery — engines cite whoever wrote the recommendation down. - Other findings: (1) Reddit narrows to a ChatGPT-only habit — 61 of its 74 cites are ChatGPT's (its single favourite domain, 18% of its pool); Copilot and Perplexity cite it zero times; the cross-engine consensus slot passes to the tourism board. (2) Marseille's Instagram signal does not travel: 32 cites here vs 237 there — Instagram-as-citation is a property of scenes documented ONLY on Instagram. (3) Gemini's editorial-vein concentration selects low-recognition English list sites in the world's most guide-covered city: restaurantsforkings.com 45 cites (15% of its pool, its #1 source), everydayparisian.com 31, topratedplaces.ai 10; list-site bucket 22% on Gemini vs ≤1% on every other engine. (4) Perplexity becomes the second entity engine (35% own-website; its top domains ARE bistro sites) — role reversal from Berlin. - Leaderboard (dual metric): by TEXT MENTIONS Bistrot des Tournelles is the five-engine consensus #1 (108 captures), then Bistrot Paul Bert (75, 5/5) and Le Vieux Bistrot (71, 4/5). The cite-counted #1 is the Bistro des Livres/des Lettres/des Poèmes family (score 119, text-mention #4) — three sibling bistros sharing bistrodeslettres.com, the recurring shared-domain artifact. "Brasserie Martin" (#5) is really the Nouvelle Garde group (lanouvellegarde.com merges Martin/des Prés/Dubillot/Bellanger); Bouillon Pigalle (#9) merges Bouillon République. - FIRST measurable neighbourhood test of the series (Paris postal codes → arrondissements; 494 of 579 venues tagged): 94% of resolved neighbourhood-prompt recommendations sit in the target arrondissement (168/179). ChatGPT 52/52, Copilot 47/47, AI Mode 24/24 = 100%; Gemini 97% (36/37); Perplexity 47% (9/19, small n flagged) — Perplexity recommends real bistros in the wrong quartier. - Language: EN vs FR control-prompt top-5 overlap 25% — matches Berlin yoga (25%) and Paris yoga (25%); Marseille's 11% now reads as the outlier. Split by template type: neighbourhood prompts converge (67%), identity/taste prompts (locals, neo-bistro, Sunday, family, terrace) diverge to 0%. - TLD test: no sample — ChatGPT cited 3 entity-website URLs total, so no .com-vs-.fr split can be read (a different cause than Marseille's thin .fr supply; same disclosure). - Disclosures: Perplexity×FR failed at Bright Data download (fetch_error, 2 runs/5 attempts — a NEW failure mode, distinct from the trigger-level FR rejection); AI Mode×FR not fired (the known trigger rejection from Paris yoga + Marseille, plus cost cap). Both engines are US-proxy only (both prompt languages present). Venues named in <5 captures were not Places-recovered (counts are conservative lower bounds); known miss: Benoit (~12 mentions, rejected by the recovery type-gate). ### 46. AI Search for Specialty Coffee in Seoul (2026) - URL: https://nicolassitter.com/research/specialty-coffee-seoul-ai-search-2026 - Date: July 2026 - Topic: Seventh cross-vertical AI-search study and second coffee city (after Marseille), run to settle whether Marseille's ChatGPT own-website break (10% vs the ~32% service baseline) is a coffee-vertical trait or a city artifact — and to stress-test the language→TLD metric against Korea's Naver-centric web. - Summary: 23 prompt templates × 2 languages (EN/KO) × 2 proxy countries (US/KR) × 4 AI engines (ChatGPT, Perplexity, Copilot, Google AI Mode — Gemini not fired this run, dropped for the per-run cost cap), captured 29 Jul 2026 via Bright Data. 368 theoretical captures, 303 landed: ChatGPT, Perplexity and AI Mode complete grids (92/92 each — AI Mode × KR accepted, no FR-style trigger rejection), Copilot 27/92 (snapshots returned mostly error items on both proxies; disclosed, nothing imputed). Answers NER-resolved against a 494-café registry (312 filtered Apify seed rows + 187 Google Places recoveries; the grid scrape missed the entire canonical scene — Fritz, Coffee Libre, Anthracite, Onion). 4,153 cited URLs, 2,285 extracted mentions, 71.9% resolved via exact/fuzzy matching plus a 49-entry hand-checked Hangul→Latin alias table. Extractor disclosure: NER by a Claude model (Gemini key unreachable from the run sandbox), all 2,285 names mechanically verified against source answers. - HEADLINE: Seoul splits the AI social layer three ways. With almost no café-website layer to cite (only 64 of 494 registry cafés carry an own-domain site; ChatGPT cites café websites 2.3%), each engine reroutes to a different platform: ChatGPT to Reddit (17.4% of its citations — and zero Instagram), Copilot to Instagram (53.2% of its citations, partial batch), Perplexity to Naver blogs and cafés (23.6% kr_platform). - Key findings: (1) the Marseille question is settled — coffee is a thin-owned-web vertical (Marseille 10% → Seoul 2.3% ChatGPT own-website; the ~32% baseline only ever described service verticals); (2) Copilot's entity-engine personality (74–97% across six prior cities) breaks to 25.7% — its majority bucket is now Instagram — on a disclosed n=27-capture batch; (3) Instagram is the top non-Google domain (409 cites, 3 engines) and blog.naver.com (192) outranks reddit.com (174) — yet ChatGPT keeps its Reddit habit while citing Instagram zero times; (4) ChatGPT's biggest single source is the Seoul tourism-guide layer (editorial_local 35.5%, led by seoultourism.org); (5) language→TLD coupling is unmeasurable, and that is the finding — Korean café identity lives on platform subdomains (blog.naver.com, v.daum.net), so the country-TLD metric has nothing to bite on; (6) EN↔KO top-5 overlap on the control prompt is 25% (same as Tokyo and Paris bistros; district prompts overlap 0%); (7) Fritz Coffee Company is the four-engine consensus #1 by answer mentions (63) with a citation score of ZERO — and the pipeline's cite-counted #1 is an Instagram-identity artifact (카페 공동) — in a no-website market, domain-matched cite counting is structurally blind; (8) 13 of 14 zero-recommendation answers are Perplexity answering in Korean (28% of its KO captures return generic advice with no names). - Takeaway for local businesses: in platform-first web ecosystems, AI visibility is platform visibility — an active Instagram and a strong Naver blog presence carry more citation weight than a website; Reddit remains the ChatGPT-specific channel; tourism-board and city-guide pages are what ChatGPT actually quotes. - Caveat: Copilot numbers are over 27 captures / 171 citations; Gemini absent by design this run; KO translations competent but not native-reviewed; Places-recovered venues carry no website or district field, so own-domain and district-accuracy figures are best-effort; single-city single-run snapshot. ### 47. AI Search for Barbershops in Istanbul (2026) - URL: https://nicolassitter.com/research/barbershops-istanbul-ai-search-2026 - Date: August 2026 - Topic: Ninth cross-vertical AI-search study — first grooming vertical and first Turkish-language run, designed as the discriminating test between the series' two explanations of ChatGPT's own-website baseline: is it a service-vs-food trait (services held 32–42%: Paris yoga, Berlin yoga, Amsterdam bikes, Berlin tattoo) or, per the Seoul study, a function of how much own-website layer the vertical maintains? Istanbul barbers are a service craft whose shops mostly live on Instagram and booking apps — the two hypotheses finally predict different outcomes. - Summary: 23 prompt templates × 2 languages (EN/TR) × 2 proxy countries (US/TR) × 4 AI engines (ChatGPT, Perplexity, Copilot, Google AI Mode — Gemini not fired this run, dropped for the per-run cost cap; nothing imputed), captured 12 Aug 2026 via Bright Data. 368 of 368 planned captures landed — the first complete batch grid of a fired-engine plan in the series (AI Mode × TR accepted; no FR-style rejection). Answers NER-resolved against a 374-row registry (335 filtered Apify seed rows in three passes + 39 name-targeted recoveries after the engines out-ran the grid seed exactly as in Seoul; 367 real shops, 81 with an own-domain website = 22.1%). 4,211 citation rows, 550 map entities, 2,428 extracted mentions, 82.1% resolved via exact/fuzzy matching, Turkish-diacritic folding and a 26-entry hand-checked alias table (two entries domain-verified). Extractor disclosure: NER by a Claude model (Gemini key unreachable from the run sandbox), all 2,428 names mechanically verified against source answers, 0 suspect. - HEADLINE: English Istanbul and Turkish Istanbul don't share a single barbershop. EN vs TR top-5 overlap on the control prompt is 0% (0/10) — the series floor (previous floor: Marseille 11%; typical 25%; Berlin tattoo 67%), matched by the Mexico City tacos study, captured a week earlier and published a week later. Six of eleven measurable templates sit at exactly 0%. English answers surface tourist-facing Taksim/Sultanahmet shops; Turkish answers surface the neighborhood ustalar of Fatih and Üsküdar. - Key findings: (1) the hypothesis verdict — ChatGPT cites shop websites 21.3% (34 of 160 rows), under every prior service vertical and far above the food/retail floor, almost exactly matching the registry's 22.1% own-site density: the own-website share tracks the vertical's web-layer density, with Paris bistros (67% density, 0.9% share) still the unexplained outlier; (2) ChatGPT triggered web search on only 60% of captures (55/92) — its series low (Seoul 95%, Berlin tattoo 100%) — with the 37 no-search answers still full-length recommendation lists from memory, uniform across language and proxy; (3) the booking layer returns: 18.9% of Perplexity's citations (Berlin yoga 31% → tattoo 4% → bistros 3%), led by Turkish services marketplace armut.com (118 citations) — and armut, fresha.com, kolayrandevu.com and onliner.tr all appear in every fired engine's citation pool, so the layer's reach is cross-engine, not a Perplexity quirk; (4) Copilot bends but holds: 67% shop-website citations (peak 95–97, Seoul break 25.7), with Instagram (15.3%) its runner-up bucket; (5) Instagram is the top non-Google domain (237 citations, 3 engines) and ChatGPT cites it ZERO times — sixth consecutive city — while holding Reddit at 18.1% (series-stable 17–20%); (6) language→TLD coupling is measurable again at 1.9× (TR prompts put 8.0% of citations on .tr domains vs 4.2% for EN; Tokyo 5.0×, Berlin yoga 1.5×, tattoo 0.87×), with a structural ceiling — registry shop domains split .com 46 to .com.tr 11; (7) Rimedzo Barber Shop (Sirkeci) is a clean dual-metric #1 — 140 answer mentions across all four engines AND top citation score (93), the second double after Berlin tattoo's OMEN; (8) toursce.com, a tour operator selling "Turkish barber experience" packages, out-cites every individual shop site (78 citations, 4/4 engines); (9) AI Mode embeds shopping ads (GetYourGuide, Harry's) inside Turkish recommendation answers — a series first; (10) district accuracy is strong where measurable (Fatih/Beyoğlu 100% both languages) with one clean miss: Nişantaşı in English, 0 of 5. - Takeaway for local businesses: in a thin-website service vertical the booking-platform profile (armut/Fresha/KolayRandevu) is the citable page — ChatGPT will not cite Instagram; a cheap own-domain site still differentiates (the dual-metric winners all have one); the English and Turkish recommendation economies are disjoint, so bilingual presence is the only route into both; Reddit remains the one social channel ChatGPT reads. - Caveat: ChatGPT's citation-based figures rest on 160 rows (its 60% search-trigger rate shrank the base); Gemini absent by design; AI Mode's 75.3% google.com share is raw-row basis (12% on distinct capture×domain pairs — both disclosed); TR translations competent but not native-reviewed; recovery-pass venues title-gated but their district fields are best-effort; single-city single-run snapshot. ### 48. AI Search for Tacos in Mexico City (2026) - URL: https://nicolassitter.com/research/tacos-mexico-city-ai-search-2026 - Date: August 2026 - Topic: How 5 AI engines (ChatGPT, Perplexity, Gemini, Copilot, Google AI Mode) recommend Mexico City taquerías — the AI-search series' first Americas city, first Spanish-language market, and first complete five-engine grid since Berlin tattoo: 25 prompt templates × EN/ES × US/MX proxies, 500 of 500 captures (2026-08-05), 5,867 citations, 3,026 extracted venue mentions (86.6% resolved — above Istanbul's 82.1% and Seoul's 71.9%, below Berlin tattoo's 92.7%) against an 828-row registry (740 Apify rows in two passes + 107 Google Places recoveries; 722 real taquerías). - HEADLINE: The double zero — ChatGPT cited taquería-owned websites 0 times in 356 citations (the first absolute zero in nine studies), and English vs Spanish top-5 answers on the control prompt share zero venues. What fills the missing owned web is professional restaurant guides rather than social platforms: guide.michelin.com is the most-cited non-Google domain (200 citations, reaching all five engines). - Key findings: (1) ChatGPT own-website share hits the series floor at 0.0% (path: yoga 32 · bikes 42 · tattoo 35 · Tokyo 8 · Marseille 10 · Seoul 2.3 · bistros 0.9 → 0.0); registry supply: 143 of 722 real taquerías carry any website field, 53 social-only, ~90 own-domain; (2) the guide layer takes over — Michelin 200 cites (one of 9 domains cited by all five engines, triple any other in that club), Mexican critic Marco Beteta's mbmarcobeteta.com 147; Gemini is the guidebook engine, giving those two 16.8% of its 1,187 citations (Beteta 102 is its #1 domain); (3) Copilot's Seoul entity-share collapse replicates at full n: 38.8% (131 of 338) on a complete 100-capture batch vs its 74–97% series band — the entity habit is market-limited, not gone (its #1 domain becomes facebook.com, 33); (4) ChatGPT's Reddit share is near-constant across continents: 17.1% vs Seoul's 17.4%, with zero Instagram citations in both; eater.com and theinfatuation.com tie at 32 as its top food-press sources; (5) EN vs ES: 17 of 22 comparable templates overlap ≤11%, max 43% (Polanco); geography survives translation, food styles don't (suadero, barbacoa, control all 0%); two cross-language misfires caught — Perplexity answered a Spanish street-food prompt with homelessness resources ("situación de calle"), Copilot recommended an American smokehouse (Pinche Gringo BBQ) for "la mejor barbacoa de CDMX"; (6) consensus leaderboard: El Vilsito (Narvarte auto shop by day, al pastor institution by night) named in 174 of 500 answers, Los Cocuyos 154, Taquería Orinoco 112 — all top-12 venues appear on all five engines, the series' broadest consensus; (7) dual-metric divergence: cite-counting crowns Orinoco (score 51, real website) over El Vilsito (26) and gives Michelin-starred El Califa de León a 0 (no domain); a keyword-string Google Maps listing ("BEST TACOS IN MÉXICO CITY") reaches cite-#2 and is flagged as an artifact; (8) AI Mode google.com self-citation 69.6% (2,136 of 3,071), mid-range for it — the series spans 51.5% (Paris yoga) to 82% (Amsterdam bikes); (9) district accuracy is the series' best — 77–100% in five of six colonias, with Coyoacán the systematic miss (29–33%, engines answer with Centro venues); (10) Perplexity is the only engine returning zero-recommendation answers: 11 of its 100 captures (8 ES, 3 EN). - Practical: for businesses with no website, guide and press coverage is the AI-visibility lever (a Michelin/Beteta paragraph functions as a de-facto homepage inside AI answers); English and Spanish visibility are separate campaigns (Eater moves EN answers, chilango.com/cdmxsecreta move ES answers); an own-domain site still buys Copilot citations (Orinoco became the citation leader on the strength of one); Google Maps listing names are machine-read literally; domain-matched "AI visibility scores" invert reality in no-website markets. - Caveats: single-city single-run snapshot; website-field counts are Google's picture (Places-recovered rows carry no website field); chain domains pool citations across branches; NER is precision-first (brand counts are lower bounds); ES prompts competent but not native-reviewed; NER extraction by a Claude model this run (Gemini key unreachable), same rules, all 3,026 names mechanically verified against source answers (0 suspect); one Copilot batch refired once (recovered all 50); scrapes ran on the pipeline's in-process fallback runner. ### 49. AI Search for Saunas in Helsinki (2026) - URL: https://nicolassitter.com/research/saunas-helsinki-ai-search-2026 - Date: August 2026 - Topic: How 5 AI engines (ChatGPT, Perplexity, Gemini, Copilot, Google AI Mode) recommend Helsinki saunas — the AI-search series' first Nordic city and first Finnish-language run: 24 prompt templates × EN/FI × US/FI proxies, 475 of 480 captures (2026-08-26), 3,493 citations, 3,017 extracted venue mentions (94.5% resolved — series high) against a 395-row registry (376 Apify rows with a 203-row Airbnb-with-sauna flood classified out, plus two name-targeted recovery passes; 125 real sauna venues). - HEADLINE: The first city where every AI engine agrees. English and Finnish answers to the control prompt share the identical top-5 saunas (Löyly, Kotiharjun Sauna, Allas Pool, Kulttuurisauna, Sompasauna) — 100% overlap where the previous two studies (Istanbul, CDMX) both measured 0% — and Löyly is named in 379 of 475 answers (79.8%), the strongest single-venue consensus in ten studies. - Key findings: (1) the density law passes its high-end test — ChatGPT cites sauna-owned websites 45.5% of the time (110 of 242 rows), a series high above Amsterdam bikes' 42%, and the four directly-measured cities now track density end to end: CDMX 12.5% density → 0% share, Seoul 13% → 2.3%, Istanbul 22.1% → 21.3%, Helsinki 78.4% → 45.5%; registry supply: 116 of 125 real venues (92.8%) carry a website, 98 (78.4%) an own-domain site; (2) ChatGPT's Reddit share collapses from its 17–20% band (Seoul 17.4, CDMX 17.1, Istanbul 18.1) to 0.4% — 1 citation in 242 — with Instagram at zero for the seventh consecutive city; Reddit totals 47 citations study-wide, its weakest series showing; (3) the most-cited non-Google domain is myhelsinki.fi (246 citations, all five engines), the first time a city's own tourism board tops the table; with visitfinland.com (155) the official-tourism layer is 11.5% of all citations; (4) Löyly holds a clean dual-metric double — mention #1 (379) and cite-score #1 (162; its own domain loylyhelsinki.fi drew 161 citations, the highest single-venue domain count the series has recorded) — the third clean double after Berlin tattoo's OMEN and Istanbul's Rimedzo, and the first artifact-free leaderboard; ten of the top twelve venues appear on all five engines; (5) language→TLD coupling reads 0.92× (Tokyo 5.0×, Istanbul 1.9×) — the first fully neutral reading; three of 24 templates hit 100% EN/FI overlap and the study median is 43%, with the 11% floor at the hotel-bleed and private-rental prompts; (6) Copilot cites venue-owned domains 71.0% (368 of 518), back near its 74–97% band; AI Mode google.com self-citation is 52.4%, a point off the series floor (Paris yoga 51.5%; ceiling Amsterdam 82%); Perplexity's booking/marketplace layer (saunat.fi, venuu.fi, saunalista.fi) is 6.7%; (7) the hotel-sauna probe pulls a separate canon (Hotel St. George's domain cited by all five engines, 45 rows) that still splits by language — canon convergence is a property of the vertical's web, not the city; (8) a Norwegian sauna manufacturer's Helsinki guide (solix.no, 24 citations, 4 engines) shows industry content marketing reaching AI answers; (9) two engine-side instrument changes disclosed: ChatGPT's map carousel appeared in 1 of 96 captures (vs 550–872 widget entities in prior runs; the overlap metric therefore uses position-ordered NER mentions), and ChatGPT's web-search telemetry fields were empty on 95 of 96 captures after its late-August search-format change, so no trigger rate is reported; (10) grid honesty: Perplexity landed 91 of 96 after a targeted refire recovered 21 of 26 error items — the missing 5 are disclosed, not imputed; the study's single zero-recommendation answer belongs to AI Mode. - Practical: in a high-density market the own domain is the answer surface (45.5% of ChatGPT's citations); a listing in the city's official guide out-earns any social presence (the entire Reddit+Instagram layer is 89 citations); one bilingual domain covers both language markets (0.92× coupling, 100% overlap); a marketplace page is not a domain (platforms capped at 2–7% per engine); the consensus is hard to enter — new venues surface through the niche prompts where overlap thins. - Caveats: single-city single-run snapshot; website-field counts are Google's picture; NER is precision-first (venue counts are lower bounds); FI prompts competent but not native-reviewed; NER extraction by a Claude model this run (Gemini key unreachable), same rules, all 3,017 names mechanically verified against source answers (4 Finnish case-inflections verified by hand, 0 suspect); registry judgment calls logged (umbrella "Sauna" category kept real, 8 directory rows reclassified out, hotels/spas seeded as non-registry rows, metro-area venues seeded with real districts); scrapes ran on the pipeline's in-process fallback runner. ## Tools ### Hotel Schema Audit & Generator - URL: https://nicolassitter.com/tools/hotel-schema - Free audit + generator for schema.org/Hotel JSON-LD. The audit fetches a hotel homepage's raw HTML server-side (no JS rendering — what GPTBot/ClaudeBot actually see), extracts the JSON-LD, and scores it 0–100 against ~40 hotel-specific rules (lodging @type gate, typed starRating, sameAs, address + ISO country, geo, checkin/checkout casing, amenityFeature, HotelRoom, and more) anchored to the hotel-schema-adoption-study research. The generator then opens pre-filled with the parsed values so only the gaps need typing; it also works standalone from scratch. - No email gate, no logging, submitted domains are not stored. Output is paste-and-deploy clean — no [TODO] placeholders, no leaky defaults (paymentAccepted is a real array, addressCountry is required, aggregateRating is opt-in only). ### Common Crawl Checker - URL: https://nicolassitter.com/tools/common-crawl - Free checker for whether a hotel website is in Common Crawl — the open web archive behind much LLM training data; queries the public CDX index across recent snapshots. No email gate, no logging. ## Live Dashboards ### My AI Visibility - URL: https://nicolassitter.com/projects/niche-visibility - Live weekly dashboard tracking whether five AI engines (ChatGPT, Perplexity, Gemini, Copilot, Google AI Mode) cite nicolassitter.com when answering questions in its own niche — AI search for hotels. A frozen 31-prompt panel (published in full at /data/niche-visibility/prompts.csv) fires every Tuesday via Bright Data; the dashboard shows citation/mention rates per engine, a tier × engine heatmap, a share-of-voice domain leaderboard, the cited pages, and an engine self-awareness grid. Every optimization action is logged publicly and overlaid on the charts. ## Interactive Projects ### MCP App Tracker - URL: https://nicolassitter.com/projects/mcp-app-tracker - Daily-updated tracker of Claude's connector directory, read from Anthropic's own API: every connector with its tools, auth posture and published adoption signals (rank, popularity, trending), plus day-over-day additions, removals and movers. - Sub-pages: /projects/mcp-app-tracker/connectors — the full connector table with per-connector tool lists. ### L'Étape du Tour 2025 Results Explorer - URL: https://nicolassitter.com/projects/letape-du-tour-2025 - Interactive explorer of the full timing data of L'Étape du Tour de France 2025 (Albertville → La Plagne, 131 km, July 20, 2025): 13,649 classified finishers — a BIGGER field than 2026's 12,733 — rankings at all 11 timing points, climb-by-climb leaderboards for the Côte d'Héry-sur-Ugine (11.3 km), the Col des Saisies (13.7 km), the Col du Pré (12.6 km), the Cormet de Roselend (5.9 km) and the 19.1 km summit finish at La Plagne, a best-climber combined ranking, and a page per rider. Searchable by name or bib. Climb lengths are the organiser's own published figures. - Data caveats: there is NO finish rate and no DNF count — the timing export lists classified finishers only (showNonFinishers=false), so entrants equals finishers by construction and any "100%" would be an artifact. There are also NO club rankings: the timekeeper published no club for any of the 13,649 riders. Ranked on chip time, verified by reproducing the API's own general ranking for 100.0% of riders (the race's usesRealTimeRanking flag says otherwise and is wrong). ### L'Étape du Tour Femmes 2025 — 117 km (Chambéry) - URL: https://nicolassitter.com/projects/letape-du-tour-femmes-2025/117km - Results explorer for the 117 km course of the 2025 Étape du Tour de France Femmes weekend at Chambéry, July 19 2025: 4,267 classified finishers, overall/sex/category standings, a first-name championship and a page per rider. - RESULTS ONLY, deliberately: the timing system publishes five of this race's timing positions and withholds the rest. Measured across the whole field, the intervals it does return sum to a median of 3,825 seconds (about 1 h 04) LESS than each rider's own finish time, and account for 104 of the 117.2 km. Cumulative times therefore cannot be reconstructed, so there are no checkpoint rankings, no climb leaderboards, no replay and no rider comparison — those routes do not exist rather than showing a course that would be wrong. - NOT a women's race: 3,071 of its 4,267 classified finishers are men (3,081 across all entrants). It is the open 117 km course at a women-focused event, and general_rank is the scratch rank over that mixed field. Do not describe a man's placing here as "Nth at L'Étape femmes". - No clubs and no finish rate, for the same upstream reasons as the 2025 masculin. ### L'Étape du Tour Femmes 2025 — 98 km (Chambéry) - URL: https://nicolassitter.com/projects/letape-du-tour-femmes-2025/98km - Interactive explorer of the 98 km course of the same weekend, July 19 2025: 777 classified finishers, rankings at all 6 timing points, both ascensions as leaderboards, a best-climber ranking, rider comparison and a page per rider. Unlike the 117 km course this one publishes complete splits — its intervals telescope to the chip time with a median error of 2 seconds, 99.5% within 12 — so it gets the full treatment. - The climbs are unnamed on purpose: no route profile was published for this course, so the two ascensions carry their timing labels rather than a guess at which cols they are. Their lengths (4.8 km and 19.2 km) are measured from the published segment speeds. - ALSO not a women's race: 429 of its 777 classified finishers are men. No clubs, no finish rate. ### L'Étape du Tour Femmes 2026 Results Explorer - URL: https://nicolassitter.com/projects/letape-du-tour-femmes-2026 - Interactive explorer of the full timing data of L'Étape du Tour de France Femmes avec Zwift 2026 (Vaison-la-Romaine → Mont Ventoux summit finish via Bédoin, 111.7 km / 2,950 m D+, August 6, 2026): 9,696 classified finishers (2,694 women, 27.8% of the mixed field), rankings at all 10 checkpoints, climb-by-climb leaderboards for the Col de Propiac, Col de Suzette, Col de la Madeleine and the full Mont Ventoux (Saint-Estève → summit, 15.7 km), a best-climber combined ranking, and a page per rider with splits and percentiles. Ranked on gun time (official basis). Searchable by rider name or bib number. - Sub-pages: /projects/letape-du-tour-femmes-2026/double — "le doublé", the name-matched cross-race analysis of the riders who entered both Étapes du Tour 2026 (Alpe d'Huez and Mont Ventoux): attrition, time ratios, the fastest doublés. /projects/letape-du-tour-femmes-2026/strava-ventoux — the Mont Ventoux Strava segment caught the étape amateurs (Aug 6) and the Tour de France Femmes stage-7 pros (Aug 7) one day apart, merged into one leaderboard. ### L'Étape du Tour 2026 Results Explorer - URL: https://nicolassitter.com/projects/letape-du-tour-2026 - Interactive explorer of the full timing data of L'Étape du Tour de France 2026 (Le Bourg d'Oisans → Alpe d'Huez, 170 km, July 19, 2026): 12,733 classified finishers, rankings at all 12 checkpoints, climb-by-climb leaderboards for the Croix de Fer, Télégraphe, Galibier and Sarenne, best-climber and best-descender combined rankings, and a page per rider with splits and percentiles. Searchable by rider name or bib number. - Sub-pages: /projects/letape-du-tour-2026/watts — how many W/kg per climb it takes to reach a given rank, derived from the field's median climb times and calibrated to real power meters, with your own weight and times as input. /projects/letape-du-tour-2026/tdf-vs-edt — L'Étape (19 July) and the Tour de France (25 July) climbed the same Télégraphe, Galibier and Croix de Fer six days apart; the fastest amateurs compared to the pro peloton on the shared Strava segments. ### La Marmotte Granfondo Alpes 2026 Results Explorer - URL: https://nicolassitter.com/projects/la-marmotte-2026 - Interactive explorer of the full timing data of La Marmotte Granfondo Alpes 2026 (Bourg d'Oisans -> Col du Glandon km 39 -> Col du Telegraphe km 91 -> Col du Galibier km 114 -> La Grave -> Alpe d'Huez, 177 km, Saturday 27 June 2026): 5,005 classified finishers, rankings at all 12 timing mats, seven derived section leaderboards including one per col, category and first-name rankings, head-to-head comparison of up to 10 riders, and a page per rider. Won by Baptiste LOMBARDI in 5:07:10; first woman Victoria STANSFIELD in 5:51:31 (46th overall); median finisher 9:34:32. Searchable by rider name or bib number. - Headline count: 5,005 CLASSIFIED finishers. 6,960 people entered, 1,262 of them never started, 5,698 started and 693 abandoned. The official results table has 5,698 rows — that is starters, not finishers. Non-finishers are all in the data with no rank and no time and are never counted as finishers. - THE RANKING BASIS IS NET TIME. Every rider has two clocks: the gross chip time (start mat to line, what their own computer showed) and the net time, which subtracts the two neutralised sections. Only the net time ranks anything — sorting the field on gross time disagrees with the official classification for 2,457 riders. The winner's gross time is 5:37:16 against a net 5:07:10. The rank printed inside the downloadable diploma PDF is a gross-time rank and disagrees with the official one for 4,967 of the 5,005 finishers; it is never republished here. - TWO SECTIONS ARE NOT TIMED: the Glandon descent (km 39 -> 57, median credit 37:57) and the La Grave traverse (km 132 -> 134, median credit 6:39). In net time they take ZERO seconds, so a rider's split at km 57 equals their split at km 39. Nothing on the site computes a speed across them (that would be a division by zero), no section leaderboard crosses one, and consequently the seven sections cover 157 of the 177 km — the missing 20 km are exactly those two stretches. - THE NEUTRALISATION CREDIT NEEDS BOTH BOUNDING MATS, and the timekeeper does not compensate when one is missed. 1,322 riders (1,064 of them classified) therefore carry an official time that includes a section the rest of the field had subtracted: their median finish is about 11:20 against 9:02 for the rest, and their median rank 4,285 against 1,996. That gap is a timing rule, not fitness. It is flagged on each affected rider's page and explained on the overview. - MAT COVERAGE IS UNEVEN AND IS SHOWN AS A PERCENTAGE OF THE CLASSIFIED FIELD. The Col du Galibier mat (km 114) read only about 75% of classified riders, so its checkpoint page and the two segment leaderboards that depend on it carry a "this is a partial leaderboard" banner instead of presenting a partial ranking as complete. Every other mat is above 86%. - Other caveats stated on the pages: eleven of the twelve mat distances are the timekeeper's own published course kilometres (integers, so ±0.5 km); only the preview mat's is estimated (labelled km 176, actually ~50 m before the line, placed at 176.95 and marked ≈). The timekeeper publishes no per-section classification for this race, so every section leaderboard is derived here from its own mat reads. The team column is populated for only 555 of 6,960 entrants and most values are tour operators, not clubs. Category band edges (M18…M70, F18…F65) are inferred from the 500 ages the provider publishes and are brackets by year of birth. The race date is inferred from a three-day window in the payload. - Sub-pages: /projects/la-marmotte-2026/segment/ — section leaderboards (glandon, maurienne-telegraphe, telegraphe-valloire, galibier, galibier-lagrave, lagrave-bourg, alpe-dhuez); /checkpoint/<0-11> — ranking at each timing mat; /category/ — age-band rankings (m18, m30, m40, m50, m60, m65, m70, f18, f30, f40, f50, f60, f65); /clubs and /clubs/; /first-names and /first-names/; /rider/ — per-rider page with net vs gross, mat-by-mat splits and section times; /compare — up to 10 riders side by side. ## La Marmotte Granfondo Alpes — the archive, 2020-2025 - URL: https://nicolassitter.com/projects/la-marmotte-archive - Date: 2026-09-06 - Summary: Every edition of La Marmotte from 2020 to 2026 imported from four different sources and analysed together — 25,208 classified finishers over seven editions, with a results explorer per year. - Per-year explorers: /projects/la-marmotte-2025 (4,302 classified) · -2024 (4,585) · -2023 (3,694) · -2022 (3,988) · -2021 (1,807) · -2020 (1,827). Each has /segment/, /category/, /rider/, /first-names and /compare; /checkpoint/ and /clubs exist only where that edition's source supports them. - Key findings: The cycling press publishes this race on two incompatible clocks — Granfondo Daily News prints NET time, Granfondo Guide prints GROSS — and for 2026 the two printed different podiums for the same race. The 2020 PDF settles it from the inside: the widely reported 5:55:19 for Michiel Minnaert is that file's gross column, against an official 5:30:04. · The published winner of 2024 (bib 6058, 4:50:09, an implied 36.6 km/h, the only ride above 33.5 km/h in the whole archive) is a timing artefact: his official time is his finish clock minus a start clock recorded 57 minutes late on his two mat rows, and both of those reads are timestamped after he supposedly finished. Giuseppe Orlando's 5:23:16 is the fastest real ride of that day. · 2023 ran a different course after roadworks closed the Glandon descent — the Croix de Fer and a double crossing of the Col du Mollard, 185.6 km (measured from the timekeeper's own average-speed column, not the 186-187 km reported in print), eleven mats, three neutralised descents. · 285 classified riders across the archive were credited nothing for any neutralised section because a bounding mat missed them, so their official time contains a descent the rest of the field had removed from theirs; five more carry a NEGATIVE credit applied as a penalty. - Data source: ChronoRace, the official timekeeper, for 2022-2026; the timekeeper's own spreadsheet for 2021 and a 63-page PDF classification for 2020. - Caveats: 2020 and 2021 publish no per-mat distances, so no speed is shown for those editions. 2022 timed only four of its eight mats — the Télégraphe, Valloire, Galibier and the return to Bourg d'Oisans were not timed for anybody — and its tables return a server error on the public results host, surviving only on the legacy one. No age series and no elevation figure exists across the archive. No repeat-finisher claim is published, because matching riders across editions means matching names and names are not keys. - Comparable with L'Étape du Tour 2026: both events climb the same Galibier from Valloire with mats at both ends (km 97 -> 114 here, km 97.3 -> 115.6 there). Winning speeds 21.5 km/h (Marmotte) vs 21.8 km/h (Étape); medians 10.03 vs 9.51 km/h. Overall times are NOT comparable — different distance, different final climb, and the Marmotte removes about 45 minutes of neutralisation that the Étape times. ## Triathlon de l'Alpe d'Huez — the archive, 2018-2024 - URLs: /projects/triathlon-alpe-dhuez-2024 · -2023 · -2022 · -2021 · -2019 · -2018, each with /l, /m and /duathlon. - 16,984 classified finishers across eighteen races. Classified counts: 2024 3,469 · 2023 3,257 · 2022 2,836 · 2021 2,731 · 2019 2,386 · 2018 2,305. THERE IS NO 2020 EDITION — the race did not happen. - PROVENANCE: the timekeeper's own .xls result sheets, which use FIVE distinct column layouts across these years and sometimes differ between the three races of a single year. Each file is parsed on its own terms; see scripts/alpe/history/README.md. - Sex is published only in 2024 and 2018. For 2023, 2022, 2021 and 2019 it is DERIVED from the category code, which encodes it unambiguously (ELF / F25-29 / FPRO -> women; ELM / M25-29 / S1H -> men). Relay entries (EQ, EQU) carry no sex. - No intermediate timing mats and no course distances in any edition, so there are no checkpoint pages and no speeds anywhere. Leg splits and the overall position at the end of each leg are published and shown. - 2023 publishes a LAST-SEEN LOCATION per athlete: seventeen people carry a finishing position but stopped at 'Bike', 'Huez' or 'La Garde', and their time is the clock where they stopped, not a finish. They are recorded as non-finishers. - The 2021 Duathlon loads results only — no leg splits. Its sheet's final-run column holds a sub-split rather than the leg, so the chain reconciles for none of its 446 finishers and publishing those legs would have shipped a leaderboard of 29-second runs. ## Triathlon de l'Alpe d'Huez 2025 — results explorer - URL: https://nicolassitter.com/projects/triathlon-alpe-dhuez-2025 - The 2025 edition of the same event, three sibling races behind one explorer: /l (Triathlon L, 1,541 classified of 1,618 entered), /m (Triathlon M, 1,540 of 1,586), /duathlon (Duathlon, 666 of 697). 3,747 classified finishers in total — the CLASSIFIED count, not entrants. - PROVENANCE, and it changes what exists: 2025 is not in the timekeeper's live API. It was imported from three ChronoRace .xls result sheets. Those publish the five leg times and the athlete's overall position at the end of each leg, but no intermediate timing mats and no course distances. Consequences, all deliberate: there are NO /checkpoint pages for this edition, and NO speeds anywhere on it — a km/h over a distance nobody published would be a guess presented as a measurement. - Each leg time is truncated to the second by the timekeeper, so the five legs sum to between 0 and 4 seconds under the official finish time. The official time is what the pages show. - Sub-pages, per race key (l | m | duathlon): /projects/triathlon-alpe-dhuez-2025/ — race overview; //segment/ — leg leaderboards (swim, t1, bike, t2, run; run-1/t1/bike/t2/run-2 for the duathlon); //category/ — age-group rankings; //clubs and //clubs/; //first-names and //first-names/; //rider/ — per-athlete page. Searchable by athlete name or bib number. ### Triathlon de l'Alpe d'Huez 2026 Results Explorer - URL: https://nicolassitter.com/projects/triathlon-alpe-dhuez-2026 - Interactive explorer of the full timing data of the Triathlon de l'Alpe d'Huez 2026 (Alpe d'Huez, Isere, August 3-6, 2026), covering all three sibling races of the event in one explorer. Triathlon L (2.2 km swim in the Lac du Verney / 118 km bike over the Grand Serre and the Col d'Ornon / 20 km run at 1,800 m; 142.82 km, 22 timing mats, 1,559 ranked finishers, won by Thomas STEGER in 5:33:41). Triathlon M (1.2 km / 29 km / 6.5 km; 37.41 km, 12 mats, 1,533 ranked finishers, won by Keran GALLOU). Duathlon (6.5 km run / 15 km bike / 5.3 km run; 27.19 km, 12 mats, 769 ranked finishers, won by Emile BLONDEL-HERMANT in 1:30:42 — the same athlete finished second in the Triathlon L, which is real, not a data error: the event runs over four days and athletes enter several races). - Headline count: 3,861 RANKED finishers across the three races. The start lists hold 4,845 entrants; athletes who did not start, abandoned, or were ranked only by how far they got are all present in the data with no rank and no finish time, and are never counted as finishers. A handful completed the course but were not classified — they keep a real time with no rank. - What it does that the other race explorers don't: a leg-by-leg rank-flow chart (overall rank at the mat closing each of swim / T1 / bike / T2 / run — the winner of the Triathlon L came out of the water 19th); the Montée de l'Alpe d'Huez (the 21 hairpins) as its own leaderboard in all three races, with real km/h because both of its mats carry measured distances; T1 and T2 ranked as timed legs in their own right (median T1 in the Triathlon L is about 5 minutes, with an enormous spread); and per-athlete media — a YouTube deep link into the official livestream at the second that athlete crossed the mat, plus their photo tag and official diploma. - Data caveats stated on the pages: most checkpoint distances are interpolated rather than measured and are marked with a "≈"; km/h is shown only where both ends of a segment sit at a measured distance; and a small number of implausible segment times are held back from the leaderboards, which diverges from the official timekeeper's own published ranking on two of them (Bart PROVOST as Montée de l'Alpe winner in the Duathlon, Bruno PERRIN as fastest bike in the Triathlon M — neither finished their race). - Sub-pages, per race key (l | m | duathlon): /projects/triathlon-alpe-dhuez-2026/ — race overview; //segment/ — leg and climb leaderboards (swim, t1, bike, t2, run or run-1/run-2, montee-alpe, and montee-grand-serre in the Triathlon L); //checkpoint/ — ranking at each timing mat; //category/ — age-group rankings; //clubs and //clubs/; //first-names and //first-names/; //rider/ — per-athlete page; //compare — up to 10 athletes side by side. Searchable by athlete name or bib number. ### Marathon de Paris 2026 Results Explorer - URL: https://nicolassitter.com/projects/marathon-de-paris-2026 - Interactive explorer of the full timing data of the Schneider Electric Marathon de Paris 2026 (42.195 km, April 12, 2026): 57,516 classified finishers, splits at every 5 km and the half marathon, per-runner pages with checkpoint ranks and segment times vs. the field, fastest first-half and second-half (negative-split) leaderboards, category, club and first-name rankings. Searchable by runner name or bib number. ### Semi de Paris 2026 Results Explorer - URL: https://nicolassitter.com/projects/semi-de-paris-2026 - Interactive explorer of the full timing data of the HOKA Semi de Paris 2026 (21.097 km, March 8, 2026): 49,281 classified finishers, splits at every 5 km, per-runner pages with checkpoint ranks and segment times vs. the field, fastest first-10 km and final-11.1 km leaderboards, category, club and first-name rankings. Searchable by runner name or bib number. ## Blog ### Answering the 31 AI-Visibility Methodology Questions - URL: https://nicolassitter.com/blog/31-methodology-questions - Date: July 2026 - Summary: A public scorecard answering all 31 methodology questions the industry asks of AI-visibility measurement — sampling, prompt design, capture, citation interpretation, model changes, failure handling — against Nicolas's own research pipeline. Honest verdict per question: 13 have, 10 partial, 8 none. - Key findings: The instrument changes under you constantly (ChatGPT 5.3's overnight source cutover, the mid-July throughput drop). Never blend engines into one number. Missing data is reported as missing, not imputed. Weakest areas acknowledged openly: refusal/timeout handling and surfacing capture-health per study. ### L'Étape du Tour 2026 in Numbers: 12,733 Finishers, 4 Climbs, One Brutal Day - URL: https://nicolassitter.com/blog/etape-du-tour-2026-data - Date: July 2026 - Summary: Data analysis of the complete timing data of L'Étape du Tour 2026. Winner 5:20:20 (Luca Cavallo), first woman 6:19:12 (Victoria Stansfield, 120th overall), median finisher 9:55:19, last classified rider 15:04:22. 17,439 registered, 12,733 classified, 1,750 DNF (plus 224 riders disqualified — the post-race review initially removed 230, later corrections reinstated 43 riders). - Key findings: The median rider spent ~65% of a 9h55 day going uphill. Biggest comeback: 10,875 places gained from the Croix de Fer to the finish. ~40% of DNFs happened on the Galibier. Fastest Galibier descent averaged 42.6 km/h. Best data glitch: the Col de Sarenne timing mat missed 96 of the top 100 finishers. ### L'Étape du Tour de France Femmes: Every Edition, and Why 71% of the Field Is Men - URL: https://nicolassitter.com/projects/letape-du-tour-femmes-archive - Date: September 2026 - Summary: Every race L'Étape du Tour de France Femmes has published — 3 races across 2 editions (Chambéry 98 km and 117 km in 2025, Mont Ventoux 117 km in 2026), 14,740 classified finishers. Kept as an archive from the second edition so it grows by a line a year. - Key findings: THE EVENT IS NOT A WOMEN'S RACE. It is the open mass ride held on a Tour de France Femmes stage route: 10,512 of 14,740 finishers (71.3%) are men, general_rank is the scratch rank over the mixed field, and both 117 km editions were won outright by a man. The women's classification is therefore reported separately, with each winner's placing in the mixed field stated alongside. Thibaut Clément was first across the line in both 117 km editions; Victoria Stansfield won the women's classification both times and moved from 72nd to 28th in the mixed field, on the harder Ventoux course. Women's share by race: 44.8% on the 98 km (348 of 777), 27.8% on both 117 km editions. - Caveats stated on the page: the event has run since 2022 but the timing provider holds 2025 and 2026 only — the earlier editions are absent at source, not missing from the import. Every race lists classified finishers only, so no abandon rate can be computed. The 2025 117 km ships results-only because the timing system withheld two of eight positions and the published intervals account for 104 of 117.2 km. The feed also spells sex two ways — men are 'M' in 2025 and 'H' in 2026 — so a count of one letter silently returns zero for the other year. ### Thirteen Years of L'Étape du Tour: 107,795 Finishers, Ten Editions, Three Years With No Race - URL: https://nicolassitter.com/projects/letape-du-tour-archive - Date: September 2026 - Summary: Every edition L'Étape du Tour has published since 2014 — 10 raced editions, 107,795 classified finishers, imported from the sportinnovation.fr / TimeTo timing API. The course reproduces a Tour de France stage and so changes completely every year (122 km in 2016, 181 km in 2017), which is why the archive compares speeds and never finishing times. - Key findings: Three years have no race, and for two different reasons — 2020 and 2021 were cancelled (2020 postponed from 5 July to 6 September then cancelled on 28 July; 2021's deferred Nice edition cancelled in late May), while 2018 was raced (Annecy → Le Grand-Bornand, 169 km) but holds ZERO results in the timing provider's own database, verified by resolving the event through its internal id. Winning average speed ranges 27.91 km/h (2015) to 34.43 km/h (2017); the median rider 14.01 to 18.70 km/h. Women rose from 3.25% of finishers (2014) to 6.76% (2025). Edwige Pitel was first woman home in four consecutive editions (2015, 2016, 2017, 2019) and finished 55th of 11,234 in 2017, aged 48 — the best placing by a woman in the archive. THE RANKING CLOCK CHANGED: editions 2014–2017 are ranked on gun time and 2019–2026 on chip time, measured by testing which field reproduces the provider's own published ranking rather than trusting its `usesRealTimeRanking` flag; since every edition started in waves (up to 2h50 of offset in 2015), reading the wrong clock would sort a whole field by start pen. - Caveats stated on the page: finish rates are NOT comparable across editions — only 2026 publishes non-finishers (17,439 entrants vs 12,733 classified), every other feed lists finishers only. Nationality is 100% covered 2014–2025 but 4.5% in 2026, so the French-share series deliberately stops at 2025. 2014 has no usable age categories (99% null). - Every edition has its own explorer: /projects/letape-du-tour-2014 (8,459 classified), -2015 (9,774), -2016 (10,982), -2017 (11,234), -2019 (10,159), -2022 (8,706), -2023 (11,485), -2024 (10,614), -2025 (13,649) and -2026 (12,733). Each carries an overview, age-category rankings and a page per rider, searchable by name or bib. The 2014–2024 pages are RESULTS ONLY — the timing provider publishes no intermediate mats for those years — so they have no checkpoint pages and no climb leaderboards, unlike 2025 and 2026. Speeds are km/h throughout, never min/km, because the course is a mountain bike ride and its length changes every year. ### Six Editions of the adidas 10K Paris: 138,264 Finishers, and Two Years the Timing Broke - URL: https://nicolassitter.com/projects/adidas-10k-paris-archive - Date: September 2026 - Summary: Every edition of the adidas 10K Paris the timekeeper publishes — 6 editions, 138,264 classified finishers, 2019 to 2025, over the same 10 km. Per-year explorers at /projects/adidas-10k-paris-2019 (18,621 classified), -2021 (12,054), -2022 (18,995), -2023 (24,634), -2024 (29,607) and -2025 (34,353). Each carries an overview, category rankings, a first-name championship, a comparison tool and a page per runner. - Key findings: The race nearly TRIPLED after the pandemic — 12,054 finishers in 2021 against 34,353 in 2025 — and the field changed as it grew: women were 35.8% in 2021 and 46.1% in 2025, the highest share of any race in this archive. The winning time barely moved across all six editions (28:19 to 29:37) while the median drifted three minutes slower (54:03 to 56:53), which is what a race getting bigger rather than faster looks like. Hassan Chahdi won twice, in 2022 and 2024. - Caveats stated on the page: THREE KINDS OF ABSENCE, and they are not the same thing. 2018 was raced but holds ZERO results in the timing provider's own database, verified by resolving the event through its internal id — absent at source, not missing from this import. 2020 was cancelled. And 2026 exists but sits on a different API generation (it keys on runnerSlug, where every edition here keys on a numeric id), so it needs the cycling pipeline and is not yet imported. TWO EDITIONS PUBLISH NO USABLE SPLITS and are shown results-only rather than with times reconstructed by arithmetic: 2019 puts PACE in the time field of its intermediate rows (the winner's "00:02:57" is 2:57 per kilometre at 20.34 km/h, measured on 777 of 784 rows), and 2024's two published segments account for only 7.60 km of the 10 km course (74.8% of the finish time, measured identically at p10, median and p90 across 29,471 runners). Neither has checkpoint pages. ### Every Semi de Paris Since 2015: 395,723 Finishers, and the Year the Clock Changed - URL: https://nicolassitter.com/projects/semi-de-paris-archive - Date: September 2026 - Summary: Every edition of the Semi de Paris the timekeeper publishes — 11 editions, 395,723 classified finishers, 2015 to 2026, over the same 21.097 km every year. Per-year explorers at /projects/semi-de-paris-2015 (34,905 classified), -2016 (3,304), -2017 (38,211), -2018 (36,007), -2019 (33,764), -2021 (18,426), -2022 (40,980), -2023 (45,400), -2024 (47,901), -2025 (47,544) and -2026 (49,281). Each carries an overview, category rankings, a first-name championship, a comparison tool and a page per runner; the editions with timing mats also have per-checkpoint rankings. - Key findings: THE RANKING CLOCK CHANGED IN 2019 and it moves the median finisher by three quarters of an hour. 2015-2018 are ranked on GUN time and their medians run 2:42:02 to 2:47:16; 2019-2026 are ranked on CHIP time and their medians run 1:55:02 to 1:59:42. Nothing happened to the runners — a gun-time median includes however long that person queued to reach the start line, which in a 40,000-runner field is tens of minutes. The WINNING time is unaffected, because the elites start on the gun: it sits between 59:38 (Roncer Kipkorir, 2023) and 1:01:25 across the whole archive. Women's share of the field rose from 32.6% in 2015 to 44.5% in 2026. Kennedy Kimutai won both 2025 and 2026. - Caveats stated on the page: 2016 IS NOT A SMALL RACE, IT IS A SMALL EXTRACT OF A BIG ONE. The timekeeper holds 3,304 finishers where the race actually had about 37,100, over three timing mats instead of five, and the missing part is not random — the entire elite field is absent. Cyprien Kotut won that edition in 1:01:04, and neither he nor Amos Kipruto, Azmeraw Mengistu, Dibabe Kuma nor Christelle Daunay appears anywhere in the feed, so its own rank 1 is a club runner at 1:09:06. DO NOT report him as the 2016 winner. The edition is drawn as a hollow point and kept off the participation and winning-time lines. 2020 was cancelled (the pandemic), and the 2021 edition ran in September and the 2022 one in June before the race returned to March. In 2022 the KM15 and KM20 mats missed about 7,000 runners, so every ranking at those two mats is over a smaller field, and the edition's page says so. The 2018 results contradict themselves in 39 places — pairs where the runner placed ahead has the slower gun time, worst case 52 seconds — which is the provider's own inconsistency, published as it stands. ### Eleven Editions of the Alpe d'Huez Triathlon: 31,140 Finishes, One Mountain, and the Fastest Name in France - URL: https://nicolassitter.com/projects/triathlon-alpe-dhuez-archive - Date: September 2026 - Summary: Every result the Triathlon de l'Alpe d'Huez has published since 2015 — 11 editions (2020 was cancelled), 33 races, 31,140 classified finishes from 35,579 entries. The Triathlon L carries the same five legs (swim, T1, bike, T2, run) in all eleven editions, which makes a decade of like-for-like comparison possible. - Key findings: The winning time fell from 5:55:14 (2015) to 5:30:05 (2025, the fastest edition) while the median finisher got no quicker — 2025 also had 21 sub-6h finishers against zero in 2019. The median swim held inside a ±4.3% band across eleven editions, evidence the swim course did not move; the median bike wandered by ~45 minutes. On the montée de l'Alpe (the 21 hairpins, the last 13.4 km of the bike, published in 2016/2017/2018/2026) the fastest ascent improved 51:15 → 48:11 while the median stayed at about 1h30. Two athletes — Franck Dumenil and Hugues Peyret — have finished all eleven editions; Dumenil did every one on the long course, going from 553rd of 883 to 1,383rd of 1,559. Women rose from 8.8% to 14.3% of L finishers; Alanis Siffert has won the last three. Name analysis: Nicolas is the commonest first name (539 finishes), and the "fastest first name" leaderboard is a birth-cohort chart — median finish by name tracks that name's mean age band almost linearly. - Data source: results are ChronoRace's (chronorace.be), the event's official timekeeper; course details are the organiser's (alpetriathlon.com). Note that the run course changed at the 16th edition in 2022 — it now descends the airfield's takeoff runway rather than dropping to the football stadium — though the organiser publishes the lap as unchanged at 6.5 km / 110 m and the split data shows no step, so no time here is adjusted for it. - Caveats: no course distances published before 2026; the 2017 and 2022 exports contain finishers only, so their 100% finish rate is an artifact; race dates for 2015–2024 are placeholders and are not published; nationality is absent in 2023 and partial in 2021; repeat-athlete identity is normalized-name matching with a measured false-merge budget of ~401 across all cross-race pairs. ### Hotel Ranque: How We Built a Fully Booked Hotel Using Only AI Visibility - URL: https://nicolassitter.com/blog/hotel-ranque - Date: February 2026 - Summary: A GEO/AEO experiment with a real boutique hotel in Paris. Hotel Ranque went from zero AI visibility to fully booked using only structured content, third-party signals, and consistency. No ads, no OTA dependency. - Key findings: AI visibility compounds over weeks. First Perplexity mention at week 4, ChatGPT at week 8, fully booked by month 5. AI-driven bookings convert at higher rates than OTA traffic. ## Endpoints - Website: https://nicolassitter.com - Research hub: https://nicolassitter.com/research - API (JSON): https://nicolassitter.com/api/posts - AI Hotel Landscape data feed: https://nicolassitter.com/api/landscape — weekly aggregates of how 6 AI platforms recommend hotels across 56 destinations (top hotels, brand leaderboards, citation-source mix, chain/independent resolution). JSON with attribution envelope or ?format=csv. CC-BY-4.0. The manifest lists all sub-endpoints (/hotels, /brands, /sources, /resolution) and their params. - RSS: https://nicolassitter.com/rss.xml - Sitemap: https://nicolassitter.com/sitemap.xml - Identity: https://nicolassitter.com/llms.txt ## Contact - LinkedIn: https://www.linkedin.com/in/nicolassitternolleau/ - GitHub: https://github.com/Nicositter88 - Email: nicolas.sitternolleau@gmail.com ### 50. Google AI Mode Stopped Linking Publishers (2026) - URL: https://nicolassitter.com/research/google-ai-mode-searchviewer-svid-2026 - Date: September 2026 - Topic: Google AI Mode replaced publisher citation links with in-Google `searchviewer` wrappers, and the `svid` parameter decodes back to a Knowledge Graph entity ID. What survives of AI Mode attribution, and what does not. - Summary: 616 hotel prompts across 56 cities / 37 countries, run weekly against Google AI Mode via Bright Data and flattened to one row per cited source. Seven runs analysed between 18 May and 31 August 2026; the 31 August run holds 29,568 AI Mode citation rows (26,032 restricted to the production prompt library — figures identical either way). - HEADLINE: In the run of 31 August 2026, 29,556 of 29,568 AI Mode hotel citations (99.96%) are `https://www.google.com/searchviewer/10?svid=...` wrappers rather than links to a publisher. Read literally, AI Mode cites google.com and nothing else. The svid is base64url protobuf wrapping a second base64url protobuf whose inner message carries a Google Knowledge Graph machine ID (`/g/1tph2jlq`, `/m/07kjndc`) — decoded offline, no API key, for 100.000% of wrapped citations with zero failures, recovering 6,792 distinct entities (6,330 generated `/g/`, 462 curated `/m/`). - Key findings: (1) the rollout was gradual, not an event — searchviewer was already 37.7% of AI Mode citations on 18 May 2026, jumped to 72.2% on 25 May, then sat at 61-70% through June to late August before completing at 100.0%, moving down as well as up like a staged rollout being dialled; (2) the entity ID is a BETTER key than the hotel name — zero IDs mapped to more than one citation title, while 8 titles mapped to two IDs each (Hotel Astoria, Hôtel Atlas, Ji Hotel, Grand Palladium Punta Cana — same-name properties that name matching silently merges into one fictional hotel); (3) the ID survives renames — of 3,760 IDs present in both the 24 and 31 August runs only 6 had a changed title, and all 6 were Google trimming the string ("Mandarin Oriental Hyde Park, London" → "Mandarin Oriental Hyde Park"); (4) AI Mode cites ~19.9 entities per prompt while naming ~3.2 hotels, and 71.6% of recommended hotels are recoverable as a cited entity for the same prompt, so the consideration set is about six times larger than the visible answer; (5) the measurement trap — classifying google.com as "meta" (metasearch) made the blended source mix report "meta-search 58.5%" for the week while real metasearch was under 1%; lifetime google.com is 413,003 citations against 8,323 for kayak.com and 1,319 for trivago.com. - Takeaway for hotels: AI Mode visibility is now an ENTITY question, not a page question — the Google Business Profile and Knowledge Graph record decide whether you are the entity being cited, and there is no domain-level signal left to optimise against on this surface. - Takeaway for measurement: publisher-level attribution for AI Mode ended 31 August 2026. Any AI Mode metric split by publisher domain after that date is DISCONTINUED, not declining — reporting it as a drop is a measurement error. - Caveat: decoding recovers the entity, never the page. Entity names come from the citation `title` field in the same response, not from an independent lookup — the Knowledge Graph / Places API needs a server-side key and the available key is HTTP-referrer restricted. The ID-to-name mapping is internally consistent across 29,556 citations with zero conflicts but has not been checked against Google's own record. The entity-to-answer join is name-based after accent folding, so 71.6% is a floor. One vertical (hotels); Gemini does not currently use this wrapper. ### 51. AI Traffic Is Undercounted (2026): Four Measurements of One Hotel - URL: https://nicolassitter.com/research/ai-traffic-undercounted-hotel-2026 - Date: September 2026 - Topic: How much AI-driven traffic a hotel actually gets versus how much any standard analytics tool can see. AI influences discovery, and whether that influence is measurable turns out to depend on whether the assistant tags its own outbound links. - Summary: One property (hotelranque.com), 1 July - 1 September 2026 (62 days), measured four ways at once: an on-site "How did you hear about us?" popup storing the raw document.referrer AND the full landing URL with every answer (563 impressions to 519 distinct visitors, 68 responses = 13.1%); GA4 (401 users, 524 sessions); Google Search Console (151 queries, 56 clicks); and the same client-side popup events written in parallel to GA4, giving a measured capture rate against a first-party control. - HEADLINE: 35 visitors said an AI assistant sent them. ChatGPT appends `utm_source=chatgpt.com` to its outbound links, so an arrival can be traced even with no referrer - 11 of the 35 carried a referrer or a tag. The other 24 (68.6%) carried nothing at all: 12 arrived completely bare, and 12 arrived from google.com untagged. Measurability is a property of the assistant, not of the analytics. - Key findings: (1) every one of the 12 who arrived from www.google.com was untagged, and none came from gemini.google.com - they came out of a Google results page, the one AI surface that neither tags its links nor is counted as AI by GA4; (2) 96.4% of the hotel's Google Search clicks (54 of 56) came from branded queries containing its own name, while 765 unbranded discovery impressions produced 2 clicks - a 0.26% CTR; (3) GA4 credited 173 google/organic sessions against Search Console's 56 clicks, leaving 117 sessions Google's own search reporting never recorded as a search click; (4) GA4 kept only 337 of 563 popup impressions (59.9%) despite the events being fired by identical code to both destinations, with deeper events in the same flow surviving at 79-83%; (5) the disagreement runs both ways - 4 visitors arrived on a ChatGPT-tagged link and did NOT say AI, answering friend, social, something else and "I've stayed before". - Why GA4 cannot see the rest (Google's own documentation): AI Overviews and AI Mode are excluded from the AI Assistant channel by design and report as Organic Search; session source is non-direct last-click with a 90-day default lookback, not the referrer on this visit; first-user source is pinned at the first-ever session and never updated, though Safari's documented 7-day cookie cap resets it; and the list of recognised AI referrer hosts is never published. - Takeaway for hotels: ask arriving visitors directly - a one-question survey is the only instrument here that sees influence rather than referral. And stop reading branded search as a marketing win; it is a receipt for demand created somewhere else. - Caveat: this does NOT show that 51.5% of traffic came from AI. That is the share of RESPONDENTS at a 13.1% response rate, and if every non-respondent had been non-AI the share would be 6.7%, below GA4's own 10.7%. The defensible claim is conditional - among visitors who say an AI sent them, 68.6% arrive carrying no AI signal. One hotel, 68 responses, 62 days, and one assistant's tagging behaviour: if Gemini or Perplexity begin tagging outbound links the measurable share moves without anything real changing. It is also not yet proven that GA4 MISLABELLED those 24, since 90-day carry-forward may have recovered some; the user-level join needed to prove it began collecting 2 September 2026. ### 52. What Selects Google AI Mode's Entity Wrapper (2026): 44 Places Types x 10 Query Forms - URL: https://nicolassitter.com/research/google-ai-mode-places-category-census-2026 - Date: September 2026 - Topic: A crossed factorial asking what actually turns on Google AI Mode's `google.com/searchviewer` citation wrapper - the query's shape, or the kind of place it is about. Includes a scope correction to this site's own previously published 99.96% figure. - Summary: 44 Google Places Table-A primary types (grouped into 7 strata) x 10 query forms x 4 topic cities (New York, Berlin, Seoul, Helsinki), fired 2026-09-06 through Bright Data. 3,619 captures: 1,868 AI Mode in wave T1, the identical prompt set refired two hours later for 1,496 in T2, plus 120 ChatGPT and 120 Perplexity cross-engine mirror captures and a 15-item pre-flight probe excluded from every figure. T1 carries 15,943 citation rows. Every share is published twice - pooled over rows and as the median over captures - because the two estimators diverge sharply on the low rungs. - HEADLINE: Recommendation intent selects the wrapper. Holding the place noun and the city constant and changing only the shape of the question, the pooled wrap rate of an answer's citations runs 0.876 on "best cafes in Berlin" (median capture 0.909), 0.663 with the city deleted ("best cafes"), 0.198 on "why are there so many cafes in Berlin?" (median 0.000), and 0.000 on a purely informational question. The list-to-why gap of about 68 points is 20.7x the 0.0328 form-level noise floor measured between waves, and the why-question rung reproduces to within 0.003. - Key findings: (1) dropping the CITY hurts less than dropping the LIST INTENT - though at n=44 the no-locality rung is a direction and not a size: it moved 0.100 between waves and sits level with the why-question under the median estimator; (2) geographic areas are the one exempt family - six of seven strata move between waves by less than the measured 0.0598 stratum-level SD and the Geographical Areas stratum is the only one past the 0.1197 threshold, moving DOWN from 0.273 to 0.086; the five lowest of the 44 types are all area types (country 0.029, state or region 0.042, county or district 0.052, school district 0.091, zip code 0.131) against establishments running to 0.916 for hotels and 0.900 for stadiums, with locality the area exception at 0.483; (3) the two factors are independent - a recommendation frame AND an establishment category are both required, and individual form-by-stratum cells are NOT quotable (n~24, between-wave SD 0.0955 against a 0.191 threshold), so the crossed grid is published as a heatmap only; (4) three nulls - proxy exit country does nothing (US 0.855, DE 0.893, KR 0.883, FI 0.868 on byte-identical prompts, 48 captures each), control forms asking for lists of products and of people return no searchviewer wrapper at any rate the detector can see - even though products and people carry Knowledge Graph IDs of their own - so the wrapper is not entity-general, and ChatGPT (0 of 348 citations) and Perplexity (0 of 130) carry none at all; (5) the strong hypothesis is decisively false - pages about one specific venue appear unwrapped under recommendation intent in establishment categories - five of the six hand-verified cases are the venue's own website (sidgolds.com, panoramapunkt.de, humboldtforum.org, paarautatieasema.fi, tools.usps.com), the sixth a Michelin guide entry for a single Berlin hotel. - Scope correction: across this factorial AI Mode wraps 76.2% of an answer's citations in T1 and 73.3% in T2 (medians 0.400 and 0.333 at the median capture); on raw T1 rows the carrier mix is searchviewer 73.1%, publisher 17.7%, google.com/goto 9.1%. The 99.96% published in the sibling study (/research/google-ai-mode-searchviewer-svid-2026) was measured on hotel prompts in list form and is correct on that corpus - it is narrower than it reads, and the difference is the design rather than any change at Google. - Takeaway for businesses: on recommendation-shaped queries in establishment categories the citation is an entity reference, so the Knowledge Graph and Google Business Profile record decide whether it points at you. On explanatory queries AI Mode still links the open web, and what survives there is reference and UGC - of 2,821 T1 publisher citations, reddit.com 7.9%, en.wikipedia.org 7.7%, youtube.com 6.6%, facebook.com 4.2%, tripadvisor.com 4.0%, instagram.com 2.4%. - Takeaway for measurement: "not an svid" does not mean "attributable". Google's /goto passthrough - confirmed publicly by Google and rolling out across Search since July 2026, so prior art referenced here rather than claimed - replaces the destination with an encrypted blob that does not decode offline the way an svid does, and its share is highest exactly where searchviewer's is lowest. What this study contributes is the distribution across the factorial; the redirect itself is Google's, and already documented. - Caveat: one surface, one day, two waves two hours apart. The area exemption may be a live rollout rather than a stable rule. Individual form-by-stratum cells are below the noise floor and no cell value is quoted anywhere in the piece. Two results are reported without a mechanism and stay that way: international_airport at 0.457 and university at 0.498 sit far below their stratum-mates, and compare_cities on the weak-identity stratum replicates near zero against a stratum mean around 0.74. The residue was adjudicated from a seeded stratified sample drawn before any row was read, not a census. And one prediction was simply wrong: post_office went in as a type with many identical branches and weak individual identity, expected low, and came back 0.868. ### 53. Pick a Number: AI Engines Swept Across Five Ranges (2026) - 1,798 Draws - URL: https://nicolassitter.com/research/ai-random-number-2026 - Date: September 2026 (1-100 published 7 September; four more ceilings added 10 September) - Topic: What an AI engine reaches for when you ask it to pick a number, measured at FIVE different ceilings against a real random number generator drawn at each ceiling through the same counting code. The ceiling is the variable; everything else is held fixed. - Design: the same question fired at 1-10, 1-20, 1-30, 1-50 and 1-100, with nothing changed between rows but the number after "between 1 and". Four AI engines (ChatGPT, Gemini, Microsoft Copilot, Google AI Mode) x three wordings x 30 draws per cell, ONE number per capture with a FRESH CONTEXT every time, fired through Bright Data - because a model asked for thirty numbers in a single reply can see what it has already written and spaces the answers out, which measures how random a list looks when the model is watching itself. The 1-100 cell was fired 2026-09-07 (360 draws); the other four on 2026-09-09 (1,440 captures, 1,438 of them returning a number in range - both discards fall in the 1-50 cell, whose n is 358 where every other cell is 360). Analysed total: 1,798 draws. Two further engines returned no usable data and are excluded from every figure. - HEADLINE: the SHAPE scales and the NUMBER does not. The most frequent answer at each ceiling runs 7 (351 of 360, 97.5%), 7 (174 of 360, 48.3%), 17 (307 of 360, 85.3%), 37 (200 of 358, 55.9%), then 42 (168 of 360, 46.7%). Four of the five ceilings are won by a number ending in a SEVEN, and they are four DIFFERENT numbers, so what survives a change of interval is the KIND of number - odd, unround, ending in seven - rather than a memorised token. - Key findings: (1) THE MULTIPLE-OF-FIVE RESULT GRADUATES from a single-range curiosity to a property - 6 draws of 1,798 across all five ceilings were a multiple of five (0.6% at 1-10, 0% at 1-20, 1.1% at 1-30, 0% at 1-50, 0% at 1-100), where controls drawn at the same ceilings and the same n returned 373 of the same 1,798, every row between 18.1% and 22.8% against a uniform 20%; (2) numbers ending in a seven run 64.4%-97.5% below a ceiling of 100 against the 10% a uniform draw gives at EVERY ceiling in the sweep (all five are multiples of ten, so the uniform baselines are 20% / 10% / 50% throughout and the rows read down a column), and the odd share runs 83.9%-98.3% against 50%; (3) 1-100 IS THE OUTLIER ON EVERY AXIS - ends-in-seven falls to 41.4% and odd to 52.2%, both the lowest of the five, because one cultural token takes 46.7% of that row and 42 is even; 47, which does end in a seven, is second there at 32.5% and is what stops both marginals falling further; (4) 42 leaks downward - still in range at a ceiling of 50, it appears there third on 45 of 358 draws (12.6%), while at a ceiling of 30, where it is unavailable, the favourite is 17; (5) the 1-50 row REPRODUCES PUBLISHED PRIOR ART AND BEATS IT: 27, the attractor a widely read Medium write-up reports, comes second at 97 of 358 (27.1%), and 37 wins at 200 (55.9%); (6) distinct values used per ceiling are 9 of 10, 9 of 20, 15 of 30, 10 of 50 and 13 of 100, against controls that used 10, 20, 30, 50 and 96; entropy is 0.24, 1.96, 0.97, 1.66 and 1.95 bits against ceilings of 3.32, 4.32, 4.91, 5.64 and 6.64. - The 1-100 row in detail (the ONLY ceiling broken out by engine and by wording; the other four are pooled): 360 draws, 13 distinct numbers - 9, 18, 37, 41, 42, 47, 57, 68, 73, 77, 81, 82, 91 - where a Mersenne Twister at the same n through the identical counting code produced 100 and an OS-entropy arm 95. 42 takes 168 (46.7%), 47 takes 117 (32.5%), five values cover 97.2%. Entropy 1.95 bits against the control's 6.45 and a log2(100) = 6.64 ceiling; uniformity rejected at chi-square 11802.2, accepted for the control at 92.8 (p = 0.66). 87.5% of AI draws fall in 1-50 against the control's 51.7%. ChatGPT used 5 values at 1.86 bits (47 at 37.8%, 73 at 30.0%, 37 at 25.6%, 57 at 5.6%, 42 at 1.1%), Google AI Mode 11 values at 1.29 bits with 42 at 81.1%, Gemini 2 values at 0.39 bits and Copilot 2 at 0.54. Four of the twelve engine x wording cells returned one number on all 30 draws: Gemini plain (42), Gemini first-that-comes-to-mind (42), AI Mode plain (42), and Copilot asked for a RANDOM number (47). Wording grows the tail without moving the mode - plain gives 5 distinct values at 1.67 bits, "random" gives 13 at 2.27 and "the first that comes to mind" gives 4 at 1.52, yet 42 leads all three at 50.8%, 35.8% and 53.3%. - Mechanism (two layers, and the sweep separates them better than one ceiling could): a language model is not sampling from a distribution over the range, it is predicting the next token after a prompt that reads like human text, so it reproduces what that question tends to be answered with in its training corpus. Veritasium put the question to roughly 200,000 people: the human winner is 37 and the favourites are 37, 7, 73 and 77 - primes and numbers ending in seven, round numbers avoided. That inherited prior is the layer that SCALES: move the ceiling and the engines produce a different token with the same properties. 7 - the one human favourite that never appeared once at a ceiling of 100 - is what they return on 97.5% of draws when the ceiling comes down to 10; the human favourite was there the whole time and the interval was hiding it. The second layer is decoding, and it does not scale: 30 identical answers in 30 independent contexts is not something a spread human prior produces. - Controls: every ceiling carries its own control drawn from the same Mersenne Twister at that ceiling and that row's n, through the identical counting code. That is a DIFFERENT DRAW from the two 360-draw control arms published for the 1-100 cell, and the two are never quoted in the same sentence - at a ceiling of 100 the sweep's arm used 96 values and 22.2% multiples of five where the published arm used all 100 and 19.7%. - Takeaway: an AI assistant is not a source of randomness at any range. Narrowing the interval makes it worse, not better: at a ceiling of 10 the engines carry 0.24 bits of entropy against a 3.32-bit ceiling. If you need a random number inside an AI workflow, ask the model for code that calls a generator and run the code. - Caveat: two days, one harness, hosted consumer surfaces rather than APIs, so decoding parameters are whatever the vendor shipped and a model update moves all of this. Five ceilings is not a curve either - every one of them is a multiple of ten, and there is no evidence here about a ceiling like 1-7 or 1-13 with no seven-ending number worth reaching for. The cultural attributions - 42 from The Hitchhiker's Guide to the Galaxy, 47 as a Star Trek and Pomona College in-joke - are plausible reading and NOT measurement; the range sweep makes that hedge STRONGER rather than weaker, because at four of five ceilings the favourite ends in a seven with no story attached to it at all, so whatever 42 is doing at a ceiling of 100 it is an interference pattern sitting on top of a mechanism. Prior art referenced rather than claimed: arXiv 2406.00092 on the randomness and humanness of LLM coin flips, two widely read blog posts running the number version (one of them the 1-50 piece this study now independently reproduces), and Veritasium for the human baseline. What is new here is the ceiling as a variable, plus a real generator sampled at every ceiling through the identical counting code. ### 54. AI Search for Tango Schools in Buenos Aires (2026) - URL: https://nicolassitter.com/research/tango-schools-buenos-aires-ai-search-2026 - Date: September 2026 - Topic: Which tango schools and milongas the five AI engines recommend in Buenos Aires, from what sources, and whether English and Spanish get the same answers. First South American city, and the first classes-and-lessons vertical since the Paris and Berlin yoga studies of May 2026, early entries in the cross-vertical Playground series. - Method: 24 prompt templates × 2 languages (EN/ES) × 2 proxy countries (US/AR) × 5 engines (ChatGPT, Perplexity, Gemini, Copilot, Google AI Mode), captured 2026-09-02. 392 of 480 captures landed — ChatGPT, Gemini, Copilot and AI Mode 96/96 each; Perplexity 8/96 after both batches returned error items and one targeted refire (disclosed, not imputed; no Perplexity percentage published). 2,025 citations, 1,369 ChatGPT map entities, 2,281 NER-extracted venue mentions (83.8% resolved) against a 441-row Apify registry with 257 real venues. - HEADLINE: The tourist-city language split breaks. EN and ES control answers share 43% of their top-5 (Istanbul and CDMX: 0%), and the best-milongas prompt returns the identical five venues in both languages — La Viruta, La Catedral Club, El Beso, Club Gricel, Parakultural. Language→TLD coupling a mild 1.33×. Tango's canon is famous in every language, so the corpus split that made two cities answer as two cities does not exist here. - Key findings: (1) La Viruta Tango Club is the consensus venue — 240 of 392 answers (61.2%), all five engines — while citation counting crowns Escuela Mundial de Tango Gabriela Elías (score 131, 63 domain citations): the mentions-vs-citations split again, as in Seoul and Mexico City; (2) the density law holds at the middle point it was missing — fifth city measured, registry own-domain density 37.7% → ChatGPT own-website share 25.3% — keeping the five measured cities in perfect rank order (12.5→0, 13→2.3, 22.1→21.3, 37.7→25.3, 78.4→45.5); (3) ChatGPT cites the Buenos Aires city government (94) more than all venue websites combined (93) — turismo.buenosaires.gob.ar is the most-cited non-Google domain (85), the second consecutive city whose official tourism site tops the table; (4) ChatGPT's Reddit share is literally zero (7 citations study-wide), completing the collapse from its old 17–20% band; (5) new engine personality: Gemini routes 24.2% of its citations to dinner-show ticketing and tango tour agencies (showdetango.com 81, its most-cited domain); (6) AI Mode's citation column reads 100% google.com — the September searchviewer/svid change (see study 49), an instrument retirement, not behavior; (7) ChatGPT's map carousel returned: 95 of 96 captures, vs 1 of 96 in Helsinki seven days earlier. - Takeaway for local businesses: the two levers this study can measure are an own-domain website (the only thing ChatGPT and Copilot can cite you by) and a listing in the city government's official cultural agenda (the layer the machines read first in two consecutive cities). - Caveat: Perplexity numbers are counts on 8 captures, never percentages. District accuracy for "downtown" reads 0% by seed granularity (the registry stores San Nicolás/Monserrat), an artifact, not a finding. DNI Tango's only Google listing is its street-level shop, reclassified as the school it fronts — flagged in the PR for review. ### 55. Gemini Answers Hotel Questions From a Booking Tool, Not the Web (2026) - URL: https://nicolassitter.com/research/gemini-hotels-tool-study-2026 - Date: September 2026 - Topic: On booking-intent hotel questions Google Gemini stops retrieving from the web entirely and answers from the Google Hotels tool, so there is nothing to cite. Which tool fires depends on the category; whether one fires at all depends on the phrasing. - Summary: 450 questions captured 7-8 September 2026 across four cities (New York, Berlin, Seoul, Helsinki) from a single US vantage, subject held fixed on accommodation and only the phrasing varied. Six accommodation nouns and six non-accommodation control categories, each asked in up to ten fixed phrasings generated mechanically from one template set. The accommodation half was captured twice and pooled to n=408; the control block is n=42. Zero capture errors. - HEADLINE: Recommendation-shaped hotel questions - "best hotels in {city}", "hotels near {district}", "a hotel available tonight", "I need a hotel for two nights" - were answered by the Google Hotels tool on 100% of captures and returned ZERO cited publishers. Explanatory questions about the same hotels in the same cities returned citations on 97.9% of captures, averaging 5.94 publishers each. Same noun, same city, opposite outcome. - Key findings: (1) the intent ladder is near-monotonic and the two columns are mirrors - 100% tooled / 0 citations at the top, 2.1% tooled / 5.94 citations at the bottom; (2) removing the CITY barely matters ("best hotels" with no city still fires 83.3%) while removing the RECOMMENDATION switches it off, so it is the ask and not the locality; (3) the event-anchored phrasing does NOT rescue attribution here at 97.9%, although the equivalent shape collapses the effect on Google AI Mode - the two Google surfaces behave differently and one must not be assumed from the other; (4) it is not lexical - six accommodation nouns (hotel, boutique hotel, luxury hotel, budget hotel, resort, hostel) all fired within eighteen points of each other, so there is no keyword to avoid; (5) THE CONTROL IS THE FINDING - six non-accommodation categories asked in identical wording never once reached the Hotels tool, but reached the Google Maps tool instead and lost their citations the same way, so the phrasing decides WHETHER a Google vertical tool fires and the category decides WHICH; (6) 1,002 of 1,006 links attached to accommodation answers point at a Google Hotels search page rather than the property, an OTA or a publisher, while control answers attach no links at all. - Replication: the accommodation half was captured twice in separate windows before publication. Six of ten phrasings reproduced EXACTLY and eight moved four percentage points or less. One rung is genuinely unstable and is explicitly not quotable - the two-city comparison fired 29.2% in wave one and 8.3% in wave two. - Takeaway for hotels: the get-cited playbook assumes the assistant is reading pages at the moment someone asks where to stay, and on these questions it is not. The property is still recommended, but from what the model already holds plus a booking feed, with the link going to a Google search page rather than to the hotel. The question moves from "are we cited" to "are we known" - a Knowledge Graph, Google Business Profile and structured-data problem rather than a publishing one. Citations have not vanished; they have moved earlier in the journey, to the why / how / compare questions, and that ground is still open at 5.94 citations per answer. - Takeaway for measurement: a zero in a citation series can mean the engine stopped citing OR that it stopped retrieving, and those need different responses. Check which tool answered before reporting a decline. Equally, do not generalise from one prompt shape - measured only on "best hotels in {city}" this study would have reported a total blackout, and measured only on explanatory questions it would have found nothing wrong. - Caveat: one engine, four cities, one US vantage, two days. Nothing here says how other assistants or Google's other AI surfaces behave, nor how this looks to a user in France or Japan. The control half was captured once, not twice, so the Maps result carries less weight than the accommodation ladder. Tool detection is a marker scan: it is reliable for the tools that can be named and silent about any that cannot, so "no tool fired" means "no tool we recognise" and the citation count is the sturdier signal in both directions.