Answering All 31 Methodology Questions for AI Visibility Platforms
David McSweeney published 31 questions asking anyone selling an AI-visibility number to substantiate it. I am not selling one — but I do build the thing he is questioning, and I publish the data, so the questions land on me anyway. Here is every one of them answered without a marketing department.
By Nicolas Sitter · read the original questions
Why answer at all
- •He is right that the burden of proof sits with whoever makes the positive claim. I make a version of that claim every time I publish a mention rate.
- •Most of these questions are answerable with data the vendors already have and do not publish. Run-to-run variance is a weekend of compute, not a research programme.
- •I have published the awkward version of several of these already — the 50.5% stability figure, the model cutover that broke my own series, the pipeline bug that scored real cafés as hallucinations. It is easier to answer honestly when the embarrassing numbers are already public.
- •And the failure mode he is warning about is real: a dashboard number, a trend line, and an agent quietly writing content briefs off a coin flip.
What is actually being measured (Q1–Q7)
The construct-validity block. This is where I do best on the mechanics and worst on the sampling.
What exactly is being measured — mention, citation, recommendation, sentiment, or a composite?
- •There is no composite. I never add mentions, citations and sentiment into one "visibility score", so there is no formula and no weighting to defend.
- •Three counters, kept apart: mention (the name appears in the answer text), citation (a link, and whose domain it points at), competition (every other domain the answer leaned on).
- •Sentiment and attributes are recorded as fields on each run, never folded into the headline number.
- •No inferred constructs. I do not model "trust", "authority", "influence" or "likelihood to buy" — none of those are observable in an answer surface, so I do not claim to measure them.
- •Honest caveat: "no formula" answers the question by refusing it. It means I have nothing to justify — and also that I offer no single number for a hotelier to put on a slide.
What population does the score claim to represent?
- •My panel describes my panel. It is a mention rate within a frozen set of prompts, not an estimate of what travellers encounter in a market.
- •Where I get closer than a keyword panel: the fan-out work reads what the engine itself expands one question into — 22,518 sub-queries from 5,000 prompts, 4.5 sub-queries per prompt on average. That is engine-observed, not user-observed.
- •That constrains how I write prompts. It does not validate the panel against real demand.
- •What I try to do about it in the labelling: call the output a "built statistic", say what it is a rate of, and keep it separate from server logs and bookings.
How were the tracked prompts shown to represent that population?
- •They were not. This is his strongest question and I cannot answer it.
- •Zero percent of my panels come from observed real user prompts. I have no access to prompt logs — nobody outside OpenAI, Google and Perplexity does.
- •What I do instead is bias reduction, not sampling: build from personas × location zoom, phrase every prompt as a need ("I'm a cyclist visiting Paris for a week"), and delete anything that smuggles the subject's amenity list in as keywords.
- •Weak supporting evidence that the method is not simply rigged: the same panel construction runs on verticals I have no stake in — Marseille coffee, Amsterdam bike shops, Tokyo bookstores, Berlin tattoo studios, Paris bistros — and produces losers as readily as winners.
- •That establishes the method is not self-flattering. It does not establish that it is representative. Different claim, and I should stop letting the first stand in for the second.
How many independent user intents does the portfolio contain?
- •I report raw capture counts — "252 captures", "2,006 mentions", "4,875 searches". Those are not independent observations and I have presented them as though the scale carried weight.
- •One seed prompt × EN/FR × five engines × repeats is heavily correlated by construction. The effective n is closer to the number of distinct seeds.
- •I have never computed an effective sample size adjusted for clustering. Fair hit.
- •Cheap fix I am adopting: report seeds / captures / mentions as three separate numbers on every study, so the independent unit is visible rather than buried in the biggest one.
How are prompt weights determined?
- •Everything is unweighted. Every prompt counts once — so there is no weighting model producing the result.
- •Prompts are grouped into four intent tiers — branded, persona, neighbourhood, generic — and reported per tier rather than blended.
- •Reporting the tiers separately is the poor version of "show weighted next to unweighted": you can see that generic queries are near-zero and branded ones are near-100 instead of averaging them into a flattering middle.
- •What is missing: no weight anywhere reflects commercial importance or real query frequency, because I do not know either. So the number is not "what matters most", it is "what I asked".
How do you prevent prompt selection from determining the result?
- •Yes, this is gameable, and anyone selling it should say so plainly. Adding branded prompts and dropping hard ones would move a score without anything changing in the world.
- •The discipline I publish: freeze the panel, version it, never edit an old prompt, only add in batches — and keep branded prompts in their own tier so they cannot inflate the discovery number.
- •Where I am weak: I do not cluster near-duplicate prompts. Ten paraphrases of one intent currently count as ten.
- •When a panel changes I split the series at that date rather than recalculating history. That preserves both readings but is not a recalculation, and a reader has to notice the split.
Does greater scale improve validity or only precision?
- •Only precision. He is right, and my own biggest numbers are the clearest illustration.
- •30,002 tagged citations and 4,875 live searches buy a very tight estimate of what my panel does. They say nothing about whether the panel resembles real demand.
- •Stated plainly: more runs of an unrepresentative panel is a more precise wrong number.
- •I should stop leading study headlines with capture counts, because volume reads as authority and here it is not.
Repeatability, uncertainty and cadence (Q8–Q12, Q18)
The block where I have the most real evidence — and one gap I have no excuse for.
Under what conditions is one run per prompt sufficient?
- •Never, and I have the repeated-run study to show it rather than an opinion.
- •4,000 queries across 8 cities × traveller tiers, with 37–77 repeated runs per city × tier cell.
- •Average overlap between two runs of the same prompt: 1.1 of the top 3 hotels. Average position-1 stability: 50.5%.
- •The variance is structured, not uniform: 96.1% position-1 stability for Berlin family hotels, 17.0% for London budget. Required run count is therefore a property of the prompt, not a constant.
- •What I publish as a working floor: 5 runs per prompt per engine per week, and never trust a single week — trust a three-week trend.
- •Gap I will close: I have not computed the run count at which each cell stabilises, even though 37–77 runs per cell is more than enough to bootstrap that curve. That study is owed and the data is already sitting there.
What uncertainty accompanies each score?
- •I publish point estimates. No confidence interval, no standard error. That is indefensible and it is the fix I am least proud of not having made.
- •The arithmetic is brutal: a 3-of-5 mention rate is 60% with a margin of roughly ±40 points. I have shown numbers like that on charts without the band.
- •My substitute has been a rule of thumb — "one green week is weather, three climbing weeks is a result". That is discipline dressed as statistics.
- •Concrete fix: Wilson interval computed on the actual run count, rendered on every rate chart. It is a day of work.
What uncertainty is excluded from the statistical model?
- •Everything except run-to-run variation — which is the small term.
- •Excluded: prompt selection, panel composition, engine choice and mix, entity resolution error, model drift, country routing, extraction rules.
- •So an interval I computed tomorrow would be a statement about my panel under fixed conditions, not about a market.
- •His framing of sampling-frame error is the right one: a narrow band around a biased panel is a confident wrong answer. Where I land: put the scope in the chart caption, not in a methodology footnote nobody opens.
Are prompt–platform observations statistically independent?
- •No, and I have direct evidence from my own data that treating them as independent is wrong.
- •When ChatGPT moved to 5.3 on 5 March 2026, the share of hotel answers grounded on user-generated platforms collapsed from 21% to 2% overnight, Reddit alone down 71%. Every configuration moved together.
- •That is one event. Counting it as hundreds of independent observations about hundreds of hotels would be a category error.
- •Same shape again in July: citation throughput on ChatGPT fell from roughly 60% to 26% after a product change, across the whole panel at once.
- •I do not currently model this correlation at all. I annotate the date and move on.
Why is daily cadence methodologically appropriate?
- •I agree with the criticism and I do not run daily. Weekly, on a fixed day.
- •The stated reason: engines churn their sources on roughly that order, and daily mostly buys noise stacked on top of run-to-run variance.
- •What I do not have is the measurement that would prove weekly is right rather than convenient — an autocorrelation curve, or the share of day-to-day movement that reverses with no intervention.
- •That volatility study is scoped and not published. Calling it "planned" is not the same as having run it, so: no evidence, just a defensible prior.
How reproducible are the results?
- •Test–retest is literally what the consistency study measures, so this one I can answer with a number instead of a promise.
- •Rerunning an identical prompt returns a list nobody has seen before between 45% and 97% of the time, depending on the market. London boutique: 62 unique lists out of 66 runs.
- •So the honest expectation to set for a customer is: under unchanged conditions, expect the list to differ. The stable thing is the rate across runs, never the run.
- •Reproducible outside my platform: the prompts are published, the engines are public, and captures come from the real answer surface rather than an API approximation. What cannot be reproduced is my exact captures — they are timestamped and country-routed.
The estimator and the execution environment (Q15–Q17)
What the method assumes about the world, and how far my logged-out execution sits from a real session.
What assumptions does each estimator require?
- •There is no estimator in the statistical sense. Counts and rates, no regression, no hierarchical model, no Bayesian shrinkage. Nothing to name because nothing is being fitted.
- •The one modelling step is entity matching: turning a name in an answer into a real venue, via fuzzy matching against a registry built independently of the engines.
- •Its load-bearing assumption is that the registry covers the real world. When that assumption breaks, real venues get scored as model hallucinations.
- •That is not hypothetical — it broke on me (see Q24), which is the honest version of "how was it tested": it failed, I caught it, I republished.
How do you account for personalisation and conversation history?
- •I measure the logged-out cold start and label it a floor, not the truth.
- •One axis I have actually bounded: 4,875 live ChatGPT searches for Paris across 5 languages × 5 countries. Roughly half the hotels change with language — but the sources barely move. It cites the same English web regardless; Japanese asked from Japan returned 0.2% Japanese sources.
- •Read: locale shifts the answer, not the grounding. That is a measured limit on one source of environment error, which is more than a shrug and much less than a general answer.
- •Execution mode matters too: 400 queries run in cached versus live mode on GPT-5.4 and 5.3 produced 83% different cited domains, with 13 cells sharing no domains at all.
- •What I have never tested, and structurally cannot: logged-in accounts with memory, multi-turn history, device, real behavioural signals. I do not have consented real sessions, and no vendor does either.
What does synthetic persona text actually validate?
- •Nothing about the person. Persona text is a prompt variable and I should describe it as one.
- •I have never tested "I'm a keen cyclist visiting Paris" against sessions from real cyclists, because I have no way to get them.
- •So the persona dimension buys phrasing coverage — it stops a panel being five rewordings of one sentence — and it does not reproduce account history, memory or location signals.
- •His point lands hardest against anyone selling persona targeting as audience simulation. It is prompt diversity with a costume on.
Counting rules: mentions, units, citations, failures (Q20–Q24)
The plumbing. This is where I have done the most work, and where I have publicly got something wrong.
How are differences between models handled?
- •I never blend engines into one number. Blending is averaging your Google rank with your Bing rank: it hides the only thing you would act on.
- •The engines are not the same instrument, and the gaps are enormous. ChatGPT cites venue websites around 32% of the time in some verticals and 0.9% for Paris bistros. Grok cited Reddit in 54.5% of runs and Facebook in 63.5% while it was still measurable. Gemini fans out roughly 7.7× more sub-queries than ChatGPT. Citation throughput: ChatGPT ~26%, Google AI Mode 9–10%, Grok 0.3%.
- •Averaging those into "AI visibility" would destroy the signal, not summarise it.
- •Missing platform data is reported as missing, not imputed. When Grok stopped being measurable I ended the series and said so on the chart rather than carrying a last-known value forward.
What counts as a meaningful mention, and how is entity resolution handled?
- •A mention is not a recommendation, and bare mention-counting lies. Being named as the noisy hotel trips the same counter as winning the category — which is why sentiment is a recorded field and why I say this out loud in the guide.
- •Position is recorded, so first-and-strongest is distinguishable from twentieth-in-a-list.
- •Entity resolution is genuinely hard and I have measured how hard: in the Étape lodging study only 29% of the lodging the engines named verified as a real property around the route, and ChatGPT kept recommending a town 100 km away because it was last year's start.
- •False positives are inspectable because raw captures are kept, and when the matching is wrong I republish. Aliases, parent companies and former names are handled by a hand-maintained registry — which does not scale and which I would not sell as a system.
What is the unit of counting?
- •Two units, kept apart and never summed: text mentions (per capture, deduplicated within an answer — appearing five times in one answer is one) and citations (per linked source).
- •When the two units disagree I publish both and explain the disagreement rather than picking the flattering one. In the Marseille coffee study the cite-counted leader and the text-mention leader were different venues, and the cite-counted one turned out to be an artifact.
- •Share-of-voice denominators come from an independently built registry of real venues — 74 places around the Étape route, 81 Paris restaurants, 579 Paris bistros — not from a competitor set someone chose to look good against.
- •Weak spot: position weighting. I record position but the headline rate treats a first recommendation and a twelfth mention as equal.
How are citations interpreted?
- •A citation is a linked source in the captured answer surface. Not an inferred source, not a plain-text URL, not "a page the model probably used".
- •The distinction he is pushing on — used versus cited — is measurable, and I built a study on it. Citation throughput is the share of retrieved sources an engine actually cites: ChatGPT around 26% (down from roughly 60% before a mid-July product change), Google AI Mode 9–10%, Grok 0.3%.
- •ChatGPT also tagged every retrieved page with an internal result_source field — until the tag vanished from captures in late July 2026 (0 of 616 in the July 27 scrape). While it was visible: across 30,002 hotel citations, 99.85% came from a single licensed tier — and 37.3% of hotel questions never touched the web at all.
- •Consequence for anyone counting citations as influence: most of what a model reads is discarded, and a third of answers are produced without reading anything. A citation-share metric is measuring the visible tail of a much larger process.
- •Citations are assigned to domains and consolidated per domain; syndicated duplicates are collapsed. Whether a source supports the recommendation or an incidental fact is not distinguished — that one is unsolved on my side.
How are failures and missing observations treated?
- •I have been bitten by exactly this and fixed it in public, which is the most useful thing I can offer here.
- •The failure: my text-mention pipeline classified real cafés as model hallucinations, because the seed registry under-covered the city. Genuine venues were being counted as errors — the metric was wrong in the direction that made the engines look worse.
- •The fix: a Places-recovery step to widen the registry before scoring, then a corrected leaderboard published over the original.
- •Where I am still weak: refusals, timeouts and malformed captures are currently dropped rather than counted as absence. That biases mention rates upward.
- •And I do not surface failure rates prominently enough for a reader to judge sample health. Adding a capture-health line to every study is on the list.
Audit trail, instrument changes, automation (Q25, Q26, Q28)
Can you check my work, can you see when the ruler changed, and can this thing act on its own.
Can customers inspect and reproduce the raw measurement?
- •Captured and retained per run: raw answer text, timestamp, engine, model version where the surface exposes it, country routing, and the citation list.
- •Published per study: the prompts, the counts, per-domain tables with CSV export, and the matching rules in prose.
- •Not published: the full raw capture corpus. The reasons are volume and scraping terms, not principle — which means it is a choice I can partly reverse and should.
- •Commitment: publish the frozen panel plus a sampled raw corpus, specifically so someone else can run their tool against the same input (see Q19).
How are methodology and model changes handled historically?
- •This is the question I would most like every vendor to answer, because the instrument changes under you constantly and it is visible if you look.
- •ChatGPT 5.3, 5 March 2026: user-generated sources collapsed from 21% of hotel answers to 2% overnight, Reddit down 71%.
- •Mid-July 2026 product change: ChatGPT citation throughput fell from roughly 60% to 26%.
- •The retrieval mix is not even stable across verticals — the licensed tier is 99.85% of hotel citations and 74.6% for flights.
- •What I do: annotate the date on the chart and treat the series as broken there rather than smoothing across it. A trend line that crosses a model cutover is two instruments drawn as one.
- •Where I am weak: when my own extraction rules change I split the series rather than recomputing history. Both readings survive, but the reader has to notice the split.
What safeguards exist before automated agents act?
- •Nothing in my stack edits a website off the back of a score. There is no automated action, so there is no automated action to guard.
- •The published safeguard is procedural: change one thing, log the date, and do not read a single week — one green week is weather, three climbing weeks is a result.
- •Honest asterisk: "a human is in the loop" is a safeguard that works at my scale and would not survive productisation. If I ever automated this, his question becomes the hard one and I do not currently have a confidence threshold to answer it with.
- •The failure mode he is describing is real and near: one changed answer triggering a content brief is a machine reacting to a coin flip.
Validation, causation and business outcomes (Q13, Q14, Q19, Q27, Q29–Q31)
The part that matters most and where the honest answer is thinnest — one property, one survey, no cross-vendor test.
How was the methodology validated?
- •One clean case, not a validation programme. Hotel Ranque was built with AI visibility as effectively the only channel — no OTA contracts, no ad budget — precisely so the two families of number could be checked against each other.
- •Order of events, as predicted: a long-tail Perplexity appearance around week four, consistent presence across the big engines by week twelve, generic "best boutique hotel in Paris" placements past week twenty. Enquiries followed the same curve.
- •Nothing was preregistered. Nothing has been independently replicated. No researcher without a stake has checked it. All three are fair criticisms.
- •And n=1 property. I would not generalise an industry from it, which is what the write-up says.
Does the validation use genuinely external data?
- •Two instruments that are external to my measurement system, which is the part of Q13 I can actually defend.
- •Guests: the booking form asked "how did you hear about us?" — 21 of 52 answers said AI search, against 27 for Google. Guests are not my pipeline. Self-reported, so a floor.
- •Server logs: 5,850 identified AI-bot requests over roughly seven months on that property, of which about 21% were on-demand fetches fired by a live question. That line went from under ten a month over winter to around 390 in June — the same shape as the recommendation ramp, measured by a completely different instrument.
- •Two independent families pointing the same way is the strongest evidence in this whole answer. It is still one property and 52 self-reported answers.
What agreement exists between vendors?
- •None. I have never run my panel against a commercial tool's panel on the same brands at the same time, and as far as I can tell nobody has published that test.
- •It is the single most useful study nobody in this space has run, and it would be cheap.
- •Standing offer, including to David: I will publish a frozen panel and a raw capture corpus, and I will co-run the comparison with any vendor willing to point their tool at the same input and publish the disagreement rate whichever way it falls.
- •The nearest thing I have is cross-engine disagreement, which is the same shape of problem: Perplexity gets Paris arrondissements right 47% of the time where other engines hit 94%, and ChatGPT's own map widget contradicts its own prose in the same answer.
What causal claims are permitted?
- •None. The rule I publish is "claim timing, not cause", and it predates this post.
- •The procedure: baseline both families for two to three weeks, change exactly one thing on a logged date, watch the per-engine built statistic, then check whether server logs and branded search moved in the same window.
- •No control group, no counterfactual, no holdout. Temporal sequence plus two independent instruments agreeing is as far as I go, and I label it correlation.
- •My own published verdict on how well it works is "a bit" — the engines churn sources weekly and much of what they cite sits on pages I influence but do not own.
What external outcome does the score predict?
- •The strongest evidence I have that AI answers move real money does not come from my score at all. When ChatGPT began embedding hotel links on 7 May 2026, daily AI referral sessions across a 17,000-hotel panel jumped 62% in a day and held — net-new, not reshuffled. 43 of those hotels now draw 10%+ of new sessions from AI.
- •That proves the channel is real. It does not prove my mention rate predicts it. Different claim, and he is right to separate them.
- •The regression I have not run: does last month's mention rate on a frozen panel predict next month's AI referral sessions, controlling for brand size and existing web prominence? I have panel data and a traffic panel. That study is owed.
- •A complication, honestly stated: every instrument in this channel errs low. Roughly 94% of assistant users loop back to Google before booking, so referral undercounts, which makes the relationship harder to detect rather than easier to fake.
Why should the metric be treated as a benchmark?
- •It should not, and I do not call it one.
- •The standing distinction on this site: real measures (server logs, referral sessions, branded search, bookings, the booking-form survey) are ground truth but lagging and undercounted. Built statistics (panel mention and citation rates) are synthetic but early and movable.
- •The stated job of the built statistic is leading indicator, not market position. It is early enough to act on and constructed enough to distrust.
- •That limitation is the headline of the guide, not a disclaimer under it — which is the specific thing he asks for in this question.
What would falsify the methodology?
- •Four results that would make me withdraw or heavily qualify the built statistic:
- •A. Panel mention rate shows no relationship to AI referral sessions or booking-form attribution across multiple properties over multiple months. That is the Q29 regression, and it can come back negative.
- •B. Two panels built independently by different people for the same hotel disagree on direction, not just level. Level differences are expected; direction disagreement means the instrument is measuring the panel-writer.
- •C. Run-to-run instability swamps between-hotel differences everywhere, not just in the most competitive cells. Right now the 17%-to-96% spread is the reason I believe there is signal; if it flattened toward noise across the board, there is nothing to measure.
- •D. Documented single-lever interventions repeatedly move the panel score with no movement in any real measure. A leading indicator that never leads anything is a vanity metric.
- •The fastest of these to run is the cross-vendor test in Q19, which is exactly why I would like someone to take that offer.
Where he is most right
- •Q3 — representativeness. Nobody in this industry has a sampling frame. We all write the exam we then sit.
- •Q4 and Q11 — independence. I have been quoting capture counts as though they were observations. They are not, and my own data on synchronised model changes proves it.
- •Q9 and Q10 — uncertainty. Publishing a rate from five runs with no interval is the kind of thing I would criticise in someone else's chart.
- •Q19 — cross-vendor agreement. The absence of this test is the loudest silence in the whole category.
- •Q29 — outcome prediction. "The score went up" predicting "the score will go up" is the trap, and I have not yet run the regression that would get me out of it.
My one pushback
- •The standard he sets — validated against ground truth, externally replicated — is one the alternatives also fail. Bookings do not tell you which channel produced them. Analytics files AI-created demand under branded search, because roughly 94% of assistant users return to Google before booking. Server logs prove a model read your page and say nothing about what it then recommended.
- •Every instrument in this channel errs in the same direction: low. That asymmetry matters. Reading the channel as bigger than it is costs you some wasted effort; reading it as smaller — which every default dashboard nudges you toward — costs you the channel.
- •So "directional, synthetic, and honest about both" is a real category, not a dodge. He concedes this himself near the end, which makes the actual disagreement much smaller than the tone of 31 questions suggests.
- •He is not wrong about any of the 31. He would only be wrong if the conclusion were "therefore do not measure". The conclusion I take is: measure, publish the method, name the uncertainty, and stop calling a panel rate a benchmark.
What I am changing because of this
An answer that costs nothing is not an answer. Five things I can do with data I already have:
- 1Confidence intervals on every rate chart
Wilson interval computed on the actual run count. Removes the worst offence in Q9.
- 2Report seeds, captures and mentions separately
So the independent unit is visible instead of buried inside the largest number. Q4.
- 3Publish the run-count stabilisation curve
I already have 37–77 repeated runs per cell. Bootstrapping how many runs a prompt needs before its rate settles is a re-analysis, not a new study. Q8.
- 4A capture-health line on every study
Refusals, timeouts and malformed captures counted and shown, instead of silently dropped. Q24.
- 5The cross-vendor test — open invitation
Publish a frozen panel plus a sampled raw corpus, and co-run the comparison with any vendor prepared to publish the disagreement rate whichever way it falls. Q19.
The underlying work
Does an AI Visibility Score Mean Anything?
Real measures vs built statistics — the framework most of these answers come from.
How to Prompt-Track AI Visibility
The panel method, the freeze rule, and the five runs per prompt.
Rankings Consistency Study
The repeated-run data: 50.5% position-1 stability, 17% to 96.1% by market.
Citation Throughput
Retrieved vs cited, measured weekly — the answer to Q23.