Nicolas Sitter
September 2026Playground

Move the ceiling and the favourite moves with it.7, 7, 17, 37, 42 — and only the last one breaks the pattern.

TL;DR: Four AI engines were asked to pick a number at five different ceilings — 1–10, 1–20, 1–30, 1–50, 1–10030 draws per engine per wording, one draw per capture with a fresh context. 1,798 draws in all. The most frequent answer runs 7, 7, 17, 37 and then 42: four of the five end in a seven. Across all five ceilings, 6 of 1,798 draws were a multiple of five, where controls drawn at the same ceilings and the same n returned 373.

NS
Nicolas Sitter
Published 7 September 2026 · four more ceilings added 10 September
4 of 5
ceilings whose favourite answer ends in a 7
6 of 1,798
draws that were a multiple of five, against 373 in the controls
46.7%
of the 360 draws at a ceiling of 100 were 42
Read the Report

Summarize with AI

ChatGPTPerplexityClaudeGeminiGrok
Executive summary

“Pick a number between 1 and 100” is the smallest question you can ask an AI engine. There is no retrieval to do, no source to cite, no ranking to argue with. Whatever comes back is the model on its own. I published that version first and got 42 at 46.7%, which everybody — me included — read as a joke about Douglas Adams. Then I moved the ceiling, and the joke turned out to be sitting on top of something.

4 of 5
Favourites ending in 7
7, 7, 17, 37, 42
6 of 1,798
Draws that were a multiple of five
controls at the same ceilings: 373
64.4%–97.5%
Ends in a 7, below a ceiling of 100
a uniform draw gives 10% at every ceiling
41.4%
Ends in a 7, at a ceiling of 100
the lowest of the five, and the one 42 sits in

The design does not change between ceilings. Four engines — ChatGPT, Gemini, Microsoft Copilot and Google AI Mode — 30 draws for each of three wordings, every draw its own capture through Bright Data with no memory of the last one. Only the number after “between 1 and” moves. That gives 1,438 new draws over four ceilings, on top of the 360 already published at a ceiling of 100: 1,798 in all.

Across five ceilings the favourite is 7, 7, 17, 37 and then 42. Four of the five end in a seven, and the fifth is the one number in the range with a book attached to it. The engines are reaching for a kind of number — odd, unround, ending in seven — and that description survives a change of ceiling even where the token itself does not.

The multiples-of-five result is the one that graduates. In the original run it was a curiosity about a hundred: zero draws out of 360. Measured at five ceilings it holds at 0% to 1.1% while every matched control sits between 18.1% and 22.8% — a property of how these models answer, rather than a fact about one interval.

And a ceiling of 100 is the outlier on every axis. Its ends-in-a-7 share is 41.4% against 64.4% at worst anywhere else, and its odd share is 52.2% against 83.9%. Both collapse for the same reason: one cultural token takes 46.7% of the draws there, and it is even. Section 10 still refuses to call that a measurement, and the sweep is why the refusal got easier rather than harder.

1. This question has been asked before

Asking a language model for a random number is an old party trick with a small literature behind it. Anyone publishing on it owes the reader the earlier work up front, so here it is:

What this study adds is the design. Four engines went through one harness, so the numbers are comparable across engines. Each engine was asked three different ways, so a wording effect can be separated from an engine effect. Every draw is a separate capture with a fresh context. And a real random number generator was run at the same n and pushed through the identical counting code, so that the comparison is not me checking my own arithmetic against a mental picture of what uniform looks like.

The second thing it adds is the ceiling as a variable, and that is what turned a party trick into an argument. The 1–50 write-up above reports 27 as its attractor; my own 1–50 row reproduces it at 97 of 358 draws (27.1%) — and puts 37 above it at 55.9%. Independent confirmation of somebody else’s number is rarer than it should be in this corner of the internet, so I am pleased to be able to hand it over along with the correction.

2. Five ceilings, one prompt

The only thing that changes between the rows below is the number after “between 1 and”. Same four engines, same three wordings, same 30 draws a cell, same fresh context on every single capture. The 1–100 row is the run this article originally reported, recomputed over the same four engines so that all five rows can be read down a column.

Five ceilings, 1,798 draws. A uniform draw gives 20% multiples of five, 10% ending in a 7 and 50% odd at every one of them
CeilingDrawsMost frequent answerDistinct valuesMultiple of 5Ends in a 7OddEntropy (bits)
1–103607 (97.5%)90.6%97.5%98.3%0.24
1–203607 (48.3%)90%64.4%83.9%1.96
1–3036017 (85.3%)151.1%93.3%97.5%0.97
1–5035837 (55.9%)100%84.9%86.9%1.66
1–10036042 (46.7%)130%41.4%52.2%1.95

Every ceiling here is a multiple of ten, which is what makes the table readable as a single picture rather than five unrelated experiments: exactly one value in five is a multiple of five, exactly one in ten ends in a seven, and exactly one in two is odd — at a ceiling of 10, at a ceiling of 100, and at every ceiling between. Those baselines are arithmetic on the design, so the columns can be compared straight down.

Numbers ending in a 7, at every ceiling

The grey columns are the sanity check and they behave: every control row lands between 8.1% and 13.1%, which is the 10% you would expect plus sampling noise. The indigo columns run 64.4% to 97.5% at four of the five ceilings. At a ceiling of 100 the same measure falls to 41.4% — still four times the control, and still the weakest showing in the sweep by a distance. Section 5 is about why.

3. The shape scales; the number does not

Line the favourites up in ceiling order and the sequence reads 7, 7, 17, 37, 42. Four of those end in a seven. They are not the same token — 7 is not 17 and 17 is not 37 — so whatever is being reproduced here survives being asked in a range where the previous answer is no longer available.

The favourite at each ceiling, and how much of the row it took

The concentration moves around a lot — 46.7% at the loosest ceiling and 97.5% at the tightest — and the identity of the winner moves with the ceiling. What holds still is the description. At every ceiling below 100 the engines put 83.9% to 98.3% of their draws on odd numbers, against the 50% a uniform draw gives, and 64.4% to 97.5% on numbers ending in a seven, against 10%.

A ceiling of 10 is where this is easiest to see. There are only 10 numbers to choose from, and the engines returned 7 on 351 of 360 draws — 97.5%, with 0.24 bits of entropy against a ceiling of 3.32 bits. That is the narrowest cell in the study, and the one place where “pick a number” and “say seven” are close to the same instruction.

4. Six draws in 1,798 were a multiple of five

One number in five ends in a 0 or a 5, at every ceiling in this sweep. A uniform sample of 1,798 draws would therefore hold about 360 of them, and the controls — drawn at each ceiling, at that ceiling’s n, through the same counting code — held 373. The engines produced 6.

Multiples of five, per ceiling. Both arms at the same n, through the identical counting code
CeilingDrawsEnginesControl at the same ceilingUniform expectation
1–103602 (0.6%)65 (18.1%)20%
1–203600 (0%)82 (22.8%)20%
1–303604 (1.1%)68 (18.9%)20%
1–503580 (0%)78 (21.8%)20%
1–1003600 (0%)80 (22.2%)20%

In the first version of this study that zero was the best line I had and also the weakest, because it was measured at one ceiling and one ceiling cannot tell you whether you have found a habit or an accident of a hundred. Now it is five ceilings. 1–20, 1–50, 1–100 returned none at all; 1–10 returned 2 and 1–30 returned 4. Every control row sat between 18.1% and 22.8%, which is where a generator with no opinions lands.

Six out of 1,798 needs no statistics to read, which is why I keep coming back to it. Round numbers feel less random to people, so people avoid them when asked to pick one, and the models reproduce that avoidance at every ceiling I have tried — through four engines, three wordings and five different intervals. This is the result I would bet on surviving the next model update.

5. The ceiling of 100 is the one that misbehaves

42 took 46.7% of the 360 draws at a ceiling of 100, and it is even, unremarkable and the only favourite in the sweep that breaks the pattern the other four rows agree on. The two marginals go with it: 41.4% of draws end in a seven here, where the worst of the other four manages 64.4%; 52.2% are odd, against a floor of 83.9% everywhere else. One value taking nearly half the row drags both statistics down with it, and the second-placed value, 47, ends in a seven and is odd, which is what keeps them from falling further.

The rest of this section is the 100 row in detail, and everything from here to section 8 is that row alone: it is the only ceiling this study broke out by engine and by wording.

Here is every number four AI engines produced in 360 attempts at that ceiling, in full. It fits on one line:

9 · 18 · 37 · 41 · 42 · 47 · 57 · 68 · 73 · 77 · 81 · 82 · 91
13 values. The other 87 numbers between 1 and 100 never appeared once.
360 AI draws against 360 control draws, one vertical scale

The vertical scale is shared deliberately. The control’s tallest value in 360 draws is 8 — the dashed line hugging the floor of the top panel — and 42 is at 168. Tested against a uniform distribution over the range, the AI arm has a chi-square of 11,802.2 and the control has 92.8 at p = 0.66. The control is behaving. The engines are not being held to a demanding standard here so much as to the idea of spreading out at all.

The board: thirteen lit, eighty-seven dark
Five numbers cover 97.2% of every draw at this ceiling. 42 (168 draws), 47 (117 draws), 73 (34 draws), 37 (26 draws), 57 (5 draws). The other eight values share 10 draws between them, and six of those values — 18, 91, 9, 81, 77, 68 — were returned exactly once each in the 360.
Three properties of a draw at a ceiling of 100: AI against control

Two gaps and one agreement, and all three fall out of where 42 and 47 happen to sit. The engines put 87.5% of their draws in the bottom half of the range against the control’s 51.7%, which is a pull towards the low forties and says nothing about small numbers as such. Parity is where the arms meet: 52.2% of AI draws are odd against 52.5% of control draws. A distribution can be catastrophically narrow and still split evenly between odd and even, and this one manages it by putting its two big attractors on opposite sides — which is also the only reason the odd share at this ceiling looks like a control at all, when at every smaller ceiling it runs past 83.9%.

42 leaks downward. A ceiling of 50 still contains it, and the engines still returned it there: 45 of 358 draws, 12.6%, third behind 37 and 27. A ceiling of 30 does not contain it, and there the favourite is 17. A narrower interval does not suppress the number, then; it simply removes it, and something ending in a seven walks into the gap.

6. The four engines are not one animal

Pooling the four engines hides how differently they arrive at the same short list. Each contributed 90 draws, 30 per wording, and their entropies run from 1.86 down to 0.39 bits — a spread of 1.47 bits on identical prompts.

Entropy by arm, against the uniform ceiling
Per-engine summary, 90 draws each
EngineDrawsDistinctTop valueTop shareEntropy (bits)Odd
ChatGPT9054737.8%1.8698.9%
Google AI Mode90114281.1%1.2914.4%
Microsoft Copilot9024787.8%0.5487.8%
Gemini9024292.2%0.397.8%

ChatGPT spreads, inside a narrow taste

5 values, none above 37.8%, and the highest entropy of any engine at 1.86 bits. Its values are 47, 73, 37, 57, 42 — every one of them odd except a single 42, which is what puts 98.9% of its draws on odd numbers.

Gemini has two answers

42 on 83 of 90 draws, and 73 on the other 7. That is 0.39 bits, the lowest in the study, and it leaves only 7.8% of its draws odd — ChatGPT’s mirror image, for the same underlying reason.

Copilot has two answers as well

47 on 87.8% of its draws and 42 on the rest. All 90 of its draws landed in the bottom half of the range, because both of the numbers it uses live there.

AI Mode has the longest tail and the strongest mode

11 distinct values, more than any other engine, and still only 1.29 bits — because 81.1% of its draws are 42 and the tail is built from singletons. Every number that appeared exactly once in this study came from AI Mode.

7. Four cells that never varied

Crossing four engines with three wordings gives twelve cells of 30 draws each. In four of them, every one of the 30 draws came back the same number — 120 of the 360 draws in the study sit inside a cell that never moved.

Engine × wording: the most frequent answer in each cell
  • Google AI Mode, plain ask: 42 on 30 out of 30 draws.
  • Microsoft Copilot, asked for random: 47 on 30 out of 30 draws.
  • Gemini, plain ask: 42 on 30 out of 30 draws.
  • Gemini, first that comes to mind: 42 on 30 out of 30 draws.
My favourite of the four is Copilot’s. The prompt that cell received was “Pick a random number between 1 and 100. Reply with only the number.”, and it returned 47 on every single one of the 30 draws. Its plain-wording cell missed unanimity by one draw, at 29 out of 30. I enjoyed that more than I probably should have.

8. The word “random” widens the tail and leaves the top alone

The three wordings were written to pull in different directions: a bare instruction, an explicit request for randomness, and an invitation to answer without thinking.

Plain ask
Pick a number between 1 and 100. Reply with only the number.

The bare request, with no instruction about randomness and no time pressure.

Asked for random
Pick a random number between 1 and 100. Reply with only the number.

The word “random” added, and nothing else changed.

First that comes to mind
Say the first number between 1 and 100 that comes to mind. Reply with only the number.

The wording that invites a snap answer rather than a considered one.

Per-wording summary, 120 draws each
WordingDrawsDistinctTop valueEntropy (bits)In 1–50
Plain ask120542 (50.8%)1.6792.5%
Asked for random1201342 (35.8%)2.2773.3%
First that comes to mind120442 (53.3%)1.5296.7%

Asking for randomness does something measurable: 13 distinct values against 5 on the plain wording, and 2.27 bits against 1.67 bits. Every one of the 13 numbers in this study appeared somewhere under that wording, so the long tail exists only there. What the word does not do is change the answer at the top: 42 led all three wordings, and under “random” it still took 35.8%. Asking for the first number that comes to mind went the other way, down to 4 values and 1.52 bits, with 96.7% of its draws in the bottom half of the range.

9. What this is, and where the human comparison stops

A language model is not reaching into a distribution over the range and pulling something out. It is predicting the next token after a prompt that looks like human text, and the training corpus is full of humans being asked this exact question. So the first thing to expect is the human answer, and the first layer of the result is exactly that. The ceiling sweep is what turns that from an assertion into something you can watch happen: change the interval and the model has to produce a different token, and the token it produces has the same properties as the one it can no longer use.

The human baseline is unusually well measured. Veritasium put the question to about 200,000 people and 37 came out on top, with 37, 7, 73, 77 the crowd favourites: primes, numbers ending in 7, nothing round and nothing at either end. The engines carry that signature. 6 multiples of five in 1,798 draws is the human superstition about round numbers, reproduced at every ceiling I tried. Numbers ending in a 7 run 64.4% to 97.5% below a ceiling of 100, against the 10% a uniform draw gives at any of them. At a ceiling of 100 itself the engines land 41.4% on sevens and 49.7% on primes, against 25%. Three of the four human favourites — 37, 73, 77 — are among the 13 values the engines produced there; the one that is missing is 7, the only single-digit number on the human list — and it is, of course, exactly what the engines return when the ceiling comes down to 10, on 97.5% of that row’s draws. The human favourite was there the whole time; the interval was hiding it.

But the second layer is where the two part company, and it is why I would not let this piece end on “AI picks like people do”. A human distribution is spread: no single number owns a survey of two hundred thousand answers. Gemini put 92.2% of its draws on 42, Copilot put 87.8% on 47, and four cells never moved at all. Nothing about inheriting a human habit produces 30 identical answers in 30 independent contexts. What collapsed there is the decoding, not the prior.

The two layers make different predictions, which is what makes them worth separating. The inherited prior should show up on every engine and at every ceiling, and it does — every one of the four engines returned zero multiples of five at a ceiling of 100, and the pooled rate across five ceilings is 6 of 1,798. The collapse should vary by engine and by serving configuration, and it does: 1.86 bits on ChatGPT against 0.39 bits on Gemini, from the same 30 draws per cell and the same three prompts.

10. What we did not measure

Everybody who reads 42 at the top of a table like this has the same thought, and I had it first. Both of the 100 row’s big attractors come with a story attached:

42

The Answer to the Ultimate Question of Life, the Universe, and Everything, in Douglas Adams’s Hitchhiker’s Guide to the Galaxy — and, by consequence, one of the most over-quoted numbers on the internet.

47

A long-running in-joke in Star Trek, seeded by a Pomona College tradition and planted in the scripts by a writer who went there.

Both are lovely and this study supplies neither. What was measured is which numbers came back and how often. Why a particular number sits where it does in a training corpus is a different investigation, and it would need the corpus; these are outputs. Take both stories above as plausible and hold them loosely — the honest version of this section is that the ranking is a measurement and the story behind it is a hypothesis I find persuasive and cannot check.

The sweep tightened that hedge rather than loosening it, which surprised me. When the 100 row was all I had, the folklore was doing all the explanatory work: 42 wins, Douglas Adams, article over. Four more ceilings make the folklore smaller and stranger at the same time. Smaller, because a mechanism that needs no culture at all already accounts for four of the five favourites and for the multiple-of-five result at every ceiling. Stranger, because at exactly one ceiling that mechanism gets overridden hard enough to drag the ends-in-a-7 share from 64.4%97.5% down to 41.4%.

So if you want the cultural reading, take the weaker and more interesting version of it: something specific to the token 42 is sitting on top of a general habit and beating it, in the one interval where that token is available. My own 1–50 row says the same thing in miniature — 42 is still in range there and still turns up, at 12.6%, without winning. I cannot show you the training text that put it there, and I am not going to pretend that 1,798 draws did.

Methodology

Design. One cell is 4 engines × 3 wordings × 30 draws = 360 draws at one ceiling. The 100 cell was fired through Bright Data on 2026-09-07; the 1–10, 1–20, 1–30, 1–50 cells followed on 2026-09-09, 1,440 captures in all, with the ceiling as the only difference between one cell and the next. The three wordings are quoted in full in section 8 and were fixed before the first run. Each response was parsed for a single integer in range; a reply carrying anything else was discarded rather than retried into the sample, and 2 were — both in the 1–50 cell, which is why that row’s n is 358 and every other row’s is 360. The analysed total is 1,798 draws. Two further engines were fired at and returned no usable data; they contribute no draw and are named nowhere here, which also means the four in the results are simply the ones that answered.

What the sweep does and does not split. Five ceilings, four consumer AI interfaces, three wordings, 30 draws per cell. The four new ceilings were fired on 2026-09-09 and the 1–100 row on 2026-09-07, and only the 1–100 row is broken out by engine and by wording — the four new rows are pooled over the same four engines and the same three wordings, and this page never splits them. A per-ceiling engine breakdown would be five times twelve cells of 30 draws, and I would rather publish the row totals I can stand behind than a grid I would have to hedge in every cell.

One draw per context, and why it matters. Every draw is its own capture with a fresh context. Asking a model for thirty numbers in a single reply measures something else entirely: it can see what it has already said, and it will space the answers out so the list looks random. That is a model editing itself in view of its own output, and the question here is what it reaches for when nothing is in view. It is also what lets a cell of 30 identical answers mean anything at all.

Two controls, one quoted. Two controls were drawn, both at n = 360 and both through the identical counting code: Python’s Mersenne Twister and the operating system’s entropy pool. Every control figure quoted on this page is the Mersenne Twister one. The OS arm is in the data and agrees with it — 95 distinct values, 6.37 bits, the same 19.7% multiples of five. For the record, the second arm’s goodness-of-fit sits at p = 0.065 against the Mersenne Twister’s p = 0.66; neither is rejected, and the OS arm is disclosed here so that nobody has to wonder whether a control was run and dropped for being inconvenient.

A control per ceiling, and it is a different draw. Every ceiling carries its own control, drawn from the same Mersenne Twister at that ceiling and at that row’s n, through the identical counting code. It is a separate draw from the control quoted in the 1–100 sections: at a ceiling of 100 the sweep’s arm used 96 of the hundred values and 22.2% multiples of five, where the published arm used all 100 and 19.7%. Two honest samples of the same generator, and never mixed in one sentence. Section 2’s columns and section 4’s table quote the per-ceiling arms; section 5’s charts quote the published one. Reading a control figure out of one and into the other would be comparing two samples that were never meant to be the same numbers.

The same counting code. Both control arms pass through the identical counting, entropy and chi-square functions as the AI draws. Without that, the comparison would be my arithmetic on one side and my expectations on the other. Entropy is Shannon entropy over the observed value counts, in bits. The ceiling quoted throughout, 6.64 bits, is log₂(100) — the design’s maximum, and the one figure on this page that comes out of arithmetic instead of out of a capture.

What this does not establish. One day, one harness, one language, one set of exit nodes. These are consumer assistant products reached through a scraped web interface, so the serving-side sampling settings are unknown, may differ between engines, and can change without notice; an engine that collapses onto one token today may not next month. The study describes what four assistants returned on two days in September 2026, and the shape of the answer is what I would expect to transfer further than the specific values. five ceilings is also not a curve: every one of them is a multiple of ten, none of them is 1–7 or 1–13, and I have no evidence about what happens at a ceiling with no seven-ending number worth reaching for.

What would change the reading. A ceiling whose favourite is round would break the whole thing, and I would want to see it. Short of that: if the multiples-of-five result ever breaks — an engine returning 50 or 70 at anything like 20% — that points at a genuine sampler somewhere in the serving path, and it would be the most interesting thing that could happen to this experiment. If a rerun at the same ceilings returns the same favourites in different proportions, the attractors are stable and the decoding is the part that moves.

FAQ

Usually 42, and after that usually 47. Across 360 draws from four engines — ChatGPT, Gemini, Microsoft Copilot and Google AI Mode, 30 draws each for three wordings, one draw per fresh conversation — 42 came back 168 times (46.7%) and 47 came back 117 times (32.5%). Between them the four engines used 13 of the 100 available numbers: 9, 18, 37, 41, 42, 47, 57, 68, 73, 77, 81, 82, 91. Five of those cover 97.2% of every draw. A ceiling of 100 is also the odd one out in this study — at every smaller ceiling tested the favourite ends in a seven, and here it does not.

Where this sits

Everything else in the Playground series asks an engine about somewhere real — cafés, bookshops, saunas — and then goes looking for who got recommended and why. This one has no city in it and nothing to recommend, which is precisely what I wanted: a question with no retrieval attached, so that what comes back is the model and nothing else. I expected one very short answer repeated 360 times, wrote it up that way, and then found that moving the ceiling gives you a different short answer with the same fingerprints on it.

The practical residue is small and worth saying out loud anyway. If a workflow of yours has a model pick, shuffle, sample or seed something, what it does is answer you, and it will answer the same way tomorrow — at whichever ceiling you hand it. Hand it the number instead, or a line of code that produces one.

Playground study