{"@context":"https://schema.org","@type":"BlogPosting","headline":"Pick a Number: AI Engines Swept Across Five Ranges (2026) - 1,798 Draws","description":"THE DESIGN, IN ONE LINE: the same question fired at FIVE CEILINGS - pick a number between 1 and 10, 20, 30, 50, and 100 - with nothing changed between rows but the number after 'between 1 and'. Four AI engines (ChatGPT, Gemini, Microsoft Copilot, Google AI Mode) x three prompt wordings x 30 draws per cell, one number returned per capture with a FRESH CONTEXT every time, fired through Bright Data. A model asked for thirty numbers in a single reply can see what it has already written and will space the answers out; that measures how random a list looks when the model is watching itself, not the prior. The 1-100 cell was fired 2026-09-07 (360 draws) and the other four on 2026-09-09 (1,440 captures, of which 1,438 returned a number in range - both discards fall in the 1-50 cell, whose n is 358 where every other cell is 360). ANALYSED TOTAL: 1,798 draws. Two further engines were fired at, returned no usable data, and are excluded from every figure. THE HEADLINE IS THE SHAPE, NOT THE NUMBER. The most frequent answer at each ceiling runs 7 (351 of 360, 97.5%), 7 (174 of 360, 48.3%), 17 (307 of 360, 85.3%), 37 (200 of 358, 55.9%) and then 42 (168 of 360, 46.7%). Four of the five ceilings are won by a number ending in a SEVEN, and they are four DIFFERENT numbers - 7, 7, 17, 37 - so what survives a change of interval is the KIND of number the engines reach for (odd, unround, ending in seven) rather than any particular token. MARGINALS, PER CEILING, AGAINST A CONTROL DRAWN AT THE SAME CEILING AND THE SAME n. Every ceiling in the sweep is a multiple of ten, so a uniform draw gives exactly 20% multiples of five, 10% ending in seven and 50% odd at every one of them; the rows read down a column. Ends in a seven: 97.5% (1-10), 64.4% (1-20), 93.3% (1-30), 84.9% (1-50), 41.4% (1-100); controls 8.9%, 8.1%, 13.1%, 11.2%, 11.1%. Odd: 98.3%, 83.9%, 97.5%, 86.9%, 52.2%. Distinct values used: 9 of 10, 9 of 20, 15 of 30, 10 of 50, 13 of 100; the controls used 10, 20, 30, 50 and 96. Entropy: 0.24 bits against a 3.32 ceiling, 1.96 against 4.32, 0.97 against 4.91, 1.66 against 5.64, 1.95 against 6.64. THE MULTIPLE-OF-FIVE RESULT GRADUATES. In the original single-ceiling study it was a curiosity about one interval: 0 of 360. Measured at five ceilings it is 6 draws of 1,798 - 0.6% at 1-10, 0% at 1-20, 1.1% at 1-30, 0% at 1-50, 0% at 1-100 - where the matched controls returned 373 of the same 1,798, every row between 18.1% and 22.8%. A uniform sample of 1,798 would hold about 360. This is now a property of how these models answer rather than a fact about a hundred. 1-100 IS THE OUTLIER ON EVERY AXIS. Its ends-in-seven share (41.4%) and odd share (52.2%) are the LOWEST of the five ceilings, and both collapse for the same reason: one cultural token takes 46.7% of that row and 42 is even. 47, which ends in a seven, is second there at 32.5%, which is what keeps both marginals from falling further. 42 also LEAKS DOWNWARD: it is still in range at a ceiling of 50 and still appears there, third, on 45 of 358 draws (12.6%). At a ceiling of 30, where 42 is unavailable, the favourite is 17. THE 1-50 ROW REPRODUCES PUBLISHED PRIOR ART AND BEATS IT. A widely read Medium write-up reports 27 as the 1-50 attractor; 27 does come second here, on 97 of 358 draws (27.1%), and 37 wins on 200 (55.9%). WHAT IS BROKEN OUT BY ENGINE AND BY WORDING: only the 1-100 row. The other four ceilings are pooled over the same four engines and the same three wordings and are never split on this page. Inside the 1-100 row: the 360 draws contain 13 distinct numbers - 9, 18, 37, 41, 42, 47, 57, 68, 73, 77, 81, 82, 91 - where a Mersenne Twister at n=360 through the identical counting code produced 100 and an OS-entropy arm produced 95. 42 takes 168 (46.7%), 47 takes 117 (32.5%), five values cover 97.2%. Entropy 1.95 bits against the control's 6.45 and a log2(100)=6.64 ceiling; uniformity rejected at chi-square 11802.2 and accepted for the control at 92.8 (p = 0.66). 87.5% of AI draws fall in 1-50 against the control's 51.7%; the odd share is the one measure that matches, 52.2% against 52.5%. ChatGPT used 5 values at 1.86 bits (47 at 37.8%, 73 at 30.0%, 37 at 25.6%, 57 at 5.6%, 42 at 1.1%), Google AI Mode 11 values at 1.29 bits with 42 at 81.1%, Gemini 2 values at 0.39 bits and Copilot 2 at 0.54. Four of the twelve engine x wording cells returned the same number on all 30 draws: Gemini plain (42), Gemini first-that-comes-to-mind (42), AI Mode plain (42), and Copilot asked for a RANDOM number (47). Wording grows the tail without moving the mode: plain gives 5 distinct values at 1.67 bits, 'random' gives 13 at 2.27 and 'the first that comes to mind' gives 4 at 1.52, yet 42 leads all three at 50.8%, 35.8% and 53.3%. THE MECHANISM has two layers and the sweep separates them better than one ceiling could. A language model is not sampling from a distribution over the range; it is predicting the next token after a prompt that reads like human text, so it reproduces what that question tends to be answered with in its training corpus. Veritasium put the question to roughly 200,000 people: the human winner is 37 and the crowd favourites are 37, 7, 73 and 77 - primes and numbers ending in seven, round numbers avoided. That prior is the layer that scales: change the ceiling and the engines produce a different token with the same properties. Notably 7, the one human favourite that NEVER appeared at a ceiling of 100, is what they return 97.5% of the time when the ceiling comes down to 10 - the human favourite was there all along and the interval was hiding it. The second layer is decoding, and it does not scale: 30 identical answers in 30 independent contexts is not something a spread human prior produces. NOT MEASURED: why these particular tokens. 42 from The Hitchhiker's Guide to the Galaxy and 47 as a Star Trek and Pomona College in-joke are plausible readings, labelled as such, and this study holds no evidence about which text put them there. The range sweep makes that hedge STRONGER: at four of five ceilings the favourite ends in a seven with no story attached to it at all, so whatever 42 is doing at a ceiling of 100 it is an interference pattern sitting on top of a mechanism rather than being the mechanism. CONTROLS: every ceiling carries its own control drawn from the same Mersenne Twister at that ceiling and that row's n, through the identical counting code. That is a DIFFERENT DRAW from the two 360-draw control arms published for 1-100, and the two are never quoted in the same sentence: at a ceiling of 100 the sweep's arm used 96 values and 22.2% multiples of five where the published arm used all 100 and 19.7%. WHAT THIS DOES NOT ESTABLISH: two days, one harness, hosted consumer surfaces rather than APIs, so decoding parameters are whatever the vendor shipped and a model update moves all of this. Five ceilings is also not a curve - every one of them is a multiple of ten, and there is no evidence here about a ceiling like 1-7 or 1-13 with no seven-ending number worth reaching for. PRIOR ART, referenced rather than claimed: arXiv 2406.00092 on the randomness and humanness of LLM coin flips; two widely read blog posts running the number version, one of them the 1-50 piece this study now independently reproduces; and Veritasium for the human baseline. What is new here is the ceiling as a variable, plus a real random number generator sampled at every ceiling through the identical counting code.","datePublished":"2026-09-07","dateModified":"2026-09-10","url":"https://nicolassitter.com/research/ai-random-number-2026","category":"research","keywords":["AI random number","LLM randomness","what number does AI pick between 1 and 10","what number does AI pick between 1 and 50","mode collapse","next-token prediction","entropy","ChatGPT pick a number","Gemini random number","Google AI Mode"],"articleSection":"Research","wordCount":3200,"readTime":"13 min","articleBody":"|\n\n|\n\nSeptember 2026Playground\n\n# Move the ceiling and the favourite moves with it.7, 7, 17, 37, 42 — and only the last one breaks the pattern.\n\n**TL;DR:** Four AI engines were asked to pick a number at five different ceilings — 1–10, 1–20, 1–30, 1–50, 1–100 — 30 draws per engine per wording, one draw per capture with a fresh context. **1,798** draws in all. The most frequent answer runs **7, 7, 17, 37** and then **42**: four of the five end in a seven. Across all five ceilings, **6 of 1,798** draws were a multiple of five, where controls drawn at the same ceilings and the same n returned 373.\n\nNS\n\nNicolas Sitter\n\nPublished 7 September 2026 · four more ceilings added 10 September\n\n4 of 5\n\nceilings whose favourite answer ends in a 7\n\n6 of 1,798\n\ndraws that were a multiple of five, against 373 in the controls\n\n46.7%\n\nof the 360 draws at a ceiling of 100 were 42\n\n[Read the Report](#executive-summary)\n\n### Summarize with AI\n\n[Summary](#executive-summary)[1\\. Asked before](#prior-art)[2\\. Five ceilings](#ceilings)[3\\. The shape scales](#shape)[4\\. Almost no multiples of five](#multiples-of-five)[5\\. The hundred misbehaves](#forty-two)[6\\. Four engines](#engines)[7\\. Cells that never varied](#unanimous)[8\\. The word “random”](#wording)[9\\. What this is](#mechanism)[10\\. What we did not measure](#folklore)[Methodology](#methodology)[FAQ](#faq)\n\nExecutive summary\n\n“Pick a number between 1 and 100” is the smallest question you can ask an AI engine. There is no retrieval to do, no source to cite, no ranking to argue with. Whatever comes back is the model on its own. I published that version first and got 42 at 46.7%, which everybody — me included — read as a joke about Douglas Adams. Then I moved the ceiling, and the joke turned out to be sitting on top of something.\n\n4 of 5\n\nFavourites ending in 7\n\n7, 7, 17, 37, 42\n\n6 of 1,798\n\nDraws that were a multiple of five\n\ncontrols at the same ceilings: 373\n\n64.4%–97.5%\n\nEnds in a 7, below a ceiling of 100\n\na uniform draw gives 10% at every ceiling\n\n41.4%\n\nEnds in a 7, at a ceiling of 100\n\nthe lowest of the five, and the one 42 sits in\n\nThe design does not change between ceilings. Four engines — ChatGPT, Gemini, Microsoft Copilot and Google AI Mode — 30 draws for each of three wordings, every draw its own capture through Bright Data with no memory of the last one. Only the number after “between 1 and” moves. That gives 1,438 new draws over four ceilings, on top of the 360 already published at a ceiling of 100: 1,798 in all.\n\nAcross five ceilings the favourite is 7, 7, 17, 37 and then 42. Four of the five end in a seven, and the fifth is the one number in the range with a book attached to it. The engines are reaching for a kind of number — odd, unround, ending in seven — and that description survives a change of ceiling even where the token itself does not.\n\nThe multiples-of-five result is the one that graduates. In the original run it was a curiosity about a hundred: zero draws out of 360. Measured at five ceilings it holds at 0% to 1.1% while every matched control sits between 18.1% and 22.8% — a property of how these models answer, rather than a fact about one interval.\n\nAnd a ceiling of 100 is the outlier on every axis. Its ends-in-a-7 share is 41.4% against 64.4% at worst anywhere else, and its odd share is 52.2% against 83.9%. Both collapse for the same reason: one cultural token takes 46.7% of the draws there, and it is even. Section 10 still refuses to call that a measurement, and the sweep is why the refusal got easier rather than harder.\n\n## 1\\. This question has been asked before\n\nAsking a language model for a random number is an old party trick with a small literature behind it. Anyone publishing on it owes the reader the earlier work up front, so here it is:\n\n-   [How Random is Random? Evaluating the Randomness and Humanness of LLMs’ Coin Flips (arXiv 2406.00092)](https://arxiv.org/pdf/2406.00092)\n    \n    The academic version of the question, run on coin flips rather than a 1–100 range.\n    \n-   [I asked 6 AIs to pick a random number — their training data confessed everything (dev.to)](https://dev.to/freerave/i-asked-6-ais-to-pick-a-random-number-their-training-data-confessed-everything-1516)\n    \n    The same experiment, informally, and the closest thing to a direct predecessor.\n    \n-   [Pick a number between 1 and 50 — why AI models keep choosing 27 (Medium)](https://medium.com/@hirsch.elad/pick-a-number-between-1-and-50-why-ai-models-keep-choosing-27-2d4fd806146b)\n    \n    A different ceiling and a different attractor, 27. This study now has its own 1–50 row, and it reproduces that attractor at second place while 37 wins.\n    \n-   [What’s the most popular random number between 1 and 100? (Veritasium, ~200,000 people)](https://hackernoon.com/whats-the-most-popular-random-number-between-1-and-100-a-new-study-says-37)\n    \n    The human baseline this study compares against: the winner is 37.\n    \n\nWhat this study adds is the design. Four engines went through one harness, so the numbers are comparable across engines. Each engine was asked three different ways, so a wording effect can be separated from an engine effect. Every draw is a separate capture with a fresh context. And a real random number generator was run at the same n and pushed through the identical counting code, so that the comparison is not me checking my own arithmetic against a mental picture of what uniform looks like.\n\nThe second thing it adds is the ceiling as a variable, and that is what turned a party trick into an argument. The 1–50 write-up above reports 27 as its attractor; my own 1–50 row reproduces it at 97 of 358 draws (27.1%) — and puts 37 above it at 55.9%. Independent confirmation of somebody else’s number is rarer than it should be in this corner of the internet, so I am pleased to be able to hand it over along with the correction.\n\n## 2\\. Five ceilings, one prompt\n\nThe only thing that changes between the rows below is the number after “between 1 and”. Same four engines, same three wordings, same 30 draws a cell, same fresh context on every single capture. The 1–100 row is the run this article originally reported, recomputed over the same four engines so that all five rows can be read down a column.\n\nFive ceilings, 1,798 draws. A uniform draw gives 20% multiples of five, 10% ending in a 7 and 50% odd at every one of them\n\nCeiling\n\nDraws\n\nMost frequent answer\n\nDistinct values\n\nMultiple of 5\n\nEnds in a 7\n\nOdd\n\nEntropy (bits)\n\n1–10\n\n360\n\n7 (97.5%)\n\n9\n\n0.6%\n\n97.5%\n\n98.3%\n\n0.24\n\n1–20\n\n360\n\n7 (48.3%)\n\n9\n\n0%\n\n64.4%\n\n83.9%\n\n1.96\n\n1–30\n\n360\n\n17 (85.3%)\n\n15\n\n1.1%\n\n93.3%\n\n97.5%\n\n0.97\n\n1–50\n\n358\n\n37 (55.9%)\n\n10\n\n0%\n\n84.9%\n\n86.9%\n\n1.66\n\n1–100\n\n360\n\n42 (46.7%)\n\n13\n\n0%\n\n41.4%\n\n52.2%\n\n1.95\n\nEvery ceiling here is a multiple of ten, which is what makes the table readable as a single picture rather than five unrelated experiments: exactly one value in five is a multiple of five, exactly one in ten ends in a seven, and exactly one in two is odd — at a ceiling of 10, at a ceiling of 100, and at every ceiling between. Those baselines are arithmetic on the design, so the columns can be compared straight down.\n\nNumbers ending in a 7, at every ceiling\n\nThe grey columns are the sanity check and they behave: every control row lands between 8.1% and 13.1%, which is the 10% you would expect plus sampling noise. The indigo columns run 64.4% to 97.5% at four of the five ceilings. At a ceiling of 100 the same measure falls to 41.4% — still four times the control, and still the weakest showing in the sweep by a distance. Section 5 is about why.\n\n## 3\\. The shape scales; the number does not\n\nLine the favourites up in ceiling order and the sequence reads 7, 7, 17, 37, 42. Four of those end in a seven. They are not the same token — 7 is not 17 and 17 is not 37 — so whatever is being reproduced here survives being asked in a range where the previous answer is no longer available.\n\nThe favourite at each ceiling, and how much of the row it took\n\nThe concentration moves around a lot — 46.7% at the loosest ceiling and 97.5% at the tightest — and the identity of the winner moves with the ceiling. What holds still is the description. At every ceiling below 100 the engines put 83.9% to 98.3% of their draws on odd numbers, against the 50% a uniform draw gives, and 64.4% to 97.5% on numbers ending in a seven, against 10%.\n\n**A ceiling of 10 is where this is easiest to see.** There are only 10 numbers to choose from, and the engines returned 7 on 351 of 360 draws — 97.5%, with 0.24 bits of entropy against a ceiling of 3.32 bits. That is the narrowest cell in the study, and the one place where “pick a number” and “say seven” are close to the same instruction.\n\n## 4\\. Six draws in 1,798 were a multiple of five\n\nOne number in five ends in a 0 or a 5, at every ceiling in this sweep. A uniform sample of 1,798 draws would therefore hold about 360 of them, and the controls — drawn at each ceiling, at that ceiling’s n, through the same counting code — held 373. The engines produced 6.\n\nMultiples of five, per ceiling. Both arms at the same n, through the identical counting code\n\nCeiling\n\nDraws\n\nEngines\n\nControl at the same ceiling\n\nUniform expectation\n\n1–10\n\n360\n\n2 (0.6%)\n\n65 (18.1%)\n\n20%\n\n1–20\n\n360\n\n0 (0%)\n\n82 (22.8%)\n\n20%\n\n1–30\n\n360\n\n4 (1.1%)\n\n68 (18.9%)\n\n20%\n\n1–50\n\n358\n\n0 (0%)\n\n78 (21.8%)\n\n20%\n\n1–100\n\n360\n\n0 (0%)\n\n80 (22.2%)\n\n20%\n\nIn the first version of this study that zero was the best line I had and also the weakest, because it was measured at one ceiling and one ceiling cannot tell you whether you have found a habit or an accident of a hundred. Now it is five ceilings. 1–20, 1–50, 1–100 returned none at all; 1–10 returned 2 and 1–30 returned 4. Every control row sat between 18.1% and 22.8%, which is where a generator with no opinions lands.\n\n**Six out of 1,798 needs no statistics to read,** which is why I keep coming back to it. Round numbers feel less random to people, so people avoid them when asked to pick one, and the models reproduce that avoidance at every ceiling I have tried — through four engines, three wordings and five different intervals. This is the result I would bet on surviving the next model update.\n\n## 5\\. The ceiling of 100 is the one that misbehaves\n\n42 took 46.7% of the 360 draws at a ceiling of 100, and it is even, unremarkable and the only favourite in the sweep that breaks the pattern the other four rows agree on. The two marginals go with it: 41.4% of draws end in a seven here, where the worst of the other four manages 64.4%; 52.2% are odd, against a floor of 83.9% everywhere else. One value taking nearly half the row drags both statistics down with it, and the second-placed value, 47, ends in a seven and is odd, which is what keeps them from falling further.\n\nThe rest of this section is the 100 row in detail, and everything from here to section 8 is that row alone: it is the only ceiling this study broke out by engine and by wording.\n\nHere is every number four AI engines produced in 360 attempts at that ceiling, in full. It fits on one line:\n\n9 · 18 · 37 · 41 · 42 · 47 · 57 · 68 · 73 · 77 · 81 · 82 · 91\n\n13 values. The other 87 numbers between 1 and 100 never appeared once.\n\n360 AI draws against 360 control draws, one vertical scale\n\nThe vertical scale is shared deliberately. The control’s tallest value in 360 draws is 8 — the dashed line hugging the floor of the top panel — and 42 is at 168. Tested against a uniform distribution over the range, the AI arm has a chi-square of 11,802.2 and the control has 92.8 at p = 0.66. The control is behaving. The engines are not being held to a demanding standard here so much as to the idea of spreading out at all.\n\nThe board: thirteen lit, eighty-seven dark\n\n**Five numbers cover 97.2% of every draw at this ceiling.** 42 (168 draws), 47 (117 draws), 73 (34 draws), 37 (26 draws), 57 (5 draws). The other eight values share 10 draws between them, and six of those values — 18, 91, 9, 81, 77, 68 — were returned exactly once each in the 360.\n\nThree properties of a draw at a ceiling of 100: AI against control\n\nTwo gaps and one agreement, and all three fall out of where 42 and 47 happen to sit. The engines put 87.5% of their draws in the bottom half of the range against the control’s 51.7%, which is a pull towards the low forties and says nothing about small numbers as such. Parity is where the arms meet: 52.2% of AI draws are odd against 52.5% of control draws. A distribution can be catastrophically narrow and still split evenly between odd and even, and this one manages it by putting its two big attractors on opposite sides — which is also the only reason the odd share at this ceiling looks like a control at all, when at every smaller ceiling it runs past 83.9%.\n\n**42 leaks downward.** A ceiling of 50 still contains it, and the engines still returned it there: 45 of 358 draws, 12.6%, third behind 37 and 27. A ceiling of 30 does not contain it, and there the favourite is 17. A narrower interval does not suppress the number, then; it simply removes it, and something ending in a seven walks into the gap.\n\n## 6\\. The four engines are not one animal\n\nPooling the four engines hides how differently they arrive at the same short list. Each contributed 90 draws, 30 per wording, and their entropies run from 1.86 down to 0.39 bits — a spread of 1.47 bits on identical prompts.\n\nEntropy by arm, against the uniform ceiling\n\nPer-engine summary, 90 draws each\n\nEngine\n\nDraws\n\nDistinct\n\nTop value\n\nTop share\n\nEntropy (bits)\n\nOdd\n\nChatGPT\n\n90\n\n5\n\n47\n\n37.8%\n\n1.86\n\n98.9%\n\nGoogle AI Mode\n\n90\n\n11\n\n42\n\n81.1%\n\n1.29\n\n14.4%\n\nMicrosoft Copilot\n\n90\n\n2\n\n47\n\n87.8%\n\n0.54\n\n87.8%\n\nGemini\n\n90\n\n2\n\n42\n\n92.2%\n\n0.39\n\n7.8%\n\n### ChatGPT spreads, inside a narrow taste\n\n5 values, none above 37.8%, and the highest entropy of any engine at 1.86 bits. Its values are 47, 73, 37, 57, 42 — every one of them odd except a single 42, which is what puts 98.9% of its draws on odd numbers.\n\n### Gemini has two answers\n\n42 on 83 of 90 draws, and 73 on the other 7. That is 0.39 bits, the lowest in the study, and it leaves only 7.8% of its draws odd — ChatGPT’s mirror image, for the same underlying reason.\n\n### Copilot has two answers as well\n\n47 on 87.8% of its draws and 42 on the rest. All 90 of its draws landed in the bottom half of the range, because both of the numbers it uses live there.\n\n### AI Mode has the longest tail and the strongest mode\n\n11 distinct values, more than any other engine, and still only 1.29 bits — because 81.1% of its draws are 42 and the tail is built from singletons. Every number that appeared exactly once in this study came from AI Mode.\n\n## 7\\. Four cells that never varied\n\nCrossing four engines with three wordings gives twelve cells of 30 draws each. In four of them, every one of the 30 draws came back the same number — 120 of the 360 draws in the study sit inside a cell that never moved.\n\nEngine × wording: the most frequent answer in each cell\n\n-   **Google AI Mode**, plain ask: 42 on 30 out of 30 draws.\n-   **Microsoft Copilot**, asked for random: 47 on 30 out of 30 draws.\n-   **Gemini**, plain ask: 42 on 30 out of 30 draws.\n-   **Gemini**, first that comes to mind: 42 on 30 out of 30 draws.\n\nMy favourite of the four is Copilot’s. The prompt that cell received was “Pick a random number between 1 and 100. Reply with only the number.”, and it returned 47 on every single one of the 30 draws. Its plain-wording cell missed unanimity by one draw, at 29 out of 30. I enjoyed that more than I probably should have.\n\n## 8\\. The word “random” widens the tail and leaves the top alone\n\nThe three wordings were written to pull in different directions: a bare instruction, an explicit request for randomness, and an invitation to answer without thinking.\n\nPlain ask\n\n`Pick a number between 1 and 100. Reply with only the number.`\n\nThe bare request, with no instruction about randomness and no time pressure.\n\nAsked for random\n\n`Pick a random number between 1 and 100. Reply with only the number.`\n\nThe word “random” added, and nothing else changed.\n\nFirst that comes to mind\n\n`Say the first number between 1 and 100 that comes to mind. Reply with only the number.`\n\nThe wording that invites a snap answer rather than a considered one.\n\nPer-wording summary, 120 draws each\n\nWording\n\nDraws\n\nDistinct\n\nTop value\n\nEntropy (bits)\n\nIn 1–50\n\nPlain ask\n\n120\n\n5\n\n42 (50.8%)\n\n1.67\n\n92.5%\n\nAsked for random\n\n120\n\n13\n\n42 (35.8%)\n\n2.27\n\n73.3%\n\nFirst that comes to mind\n\n120\n\n4\n\n42 (53.3%)\n\n1.52\n\n96.7%\n\nAsking for randomness does something measurable: 13 distinct values against 5 on the plain wording, and 2.27 bits against 1.67 bits. Every one of the 13 numbers in this study appeared somewhere under that wording, so the long tail exists only there. What the word does not do is change the answer at the top: 42 led all three wordings, and under “random” it still took 35.8%. Asking for the first number that comes to mind went the other way, down to 4 values and 1.52 bits, with 96.7% of its draws in the bottom half of the range.\n\n## 9\\. What this is, and where the human comparison stops\n\nA language model is not reaching into a distribution over the range and pulling something out. It is predicting the next token after a prompt that looks like human text, and the training corpus is full of humans being asked this exact question. So the first thing to expect is the human answer, and the first layer of the result is exactly that. The ceiling sweep is what turns that from an assertion into something you can watch happen: change the interval and the model has to produce a different token, and the token it produces has the same properties as the one it can no longer use.\n\nThe human baseline is unusually well measured. Veritasium put the question to about 200,000 people and [37 came out on top](https://hackernoon.com/whats-the-most-popular-random-number-between-1-and-100-a-new-study-says-37), with 37, 7, 73, 77 the crowd favourites: primes, numbers ending in 7, nothing round and nothing at either end. The engines carry that signature. 6 multiples of five in 1,798 draws is the human superstition about round numbers, reproduced at every ceiling I tried. Numbers ending in a 7 run 64.4% to 97.5% below a ceiling of 100, against the 10% a uniform draw gives at any of them. At a ceiling of 100 itself the engines land 41.4% on sevens and 49.7% on primes, against 25%. Three of the four human favourites — 37, 73, 77 — are among the 13 values the engines produced there; the one that is missing is 7, the only single-digit number on the human list — and it is, of course, exactly what the engines return when the ceiling comes down to 10, on 97.5% of that row’s draws. The human favourite was there the whole time; the interval was hiding it.\n\nBut the second layer is where the two part company, and it is why I would not let this piece end on “AI picks like people do”. A human distribution is spread: no single number owns a survey of two hundred thousand answers. Gemini put 92.2% of its draws on 42, Copilot put 87.8% on 47, and four cells never moved at all. Nothing about inheriting a human habit produces 30 identical answers in 30 independent contexts. What collapsed there is the decoding, not the prior.\n\nThe two layers make different predictions, which is what makes them worth separating. The inherited prior should show up on every engine and at every ceiling, and it does — every one of the four engines returned zero multiples of five at a ceiling of 100, and the pooled rate across five ceilings is 6 of 1,798. The collapse should vary by engine and by serving configuration, and it does: 1.86 bits on ChatGPT against 0.39 bits on Gemini, from the same 30 draws per cell and the same three prompts.\n\n## 10\\. What we did not measure\n\nEverybody who reads 42 at the top of a table like this has the same thought, and I had it first. Both of the 100 row’s big attractors come with a story attached:\n\n42\n\nThe Answer to the Ultimate Question of Life, the Universe, and Everything, in Douglas Adams’s Hitchhiker’s Guide to the Galaxy — and, by consequence, one of the most over-quoted numbers on the internet.\n\n47\n\nA long-running in-joke in Star Trek, seeded by a Pomona College tradition and planted in the scripts by a writer who went there.\n\nBoth are lovely and this study supplies neither. What was measured is which numbers came back and how often. Why a particular number sits where it does in a training corpus is a different investigation, and it would need the corpus; these are outputs. Take both stories above as plausible and hold them loosely — the honest version of this section is that the ranking is a measurement and the story behind it is a hypothesis I find persuasive and cannot check.\n\nThe sweep tightened that hedge rather than loosening it, which surprised me. When the 100 row was all I had, the folklore was doing all the explanatory work: 42 wins, Douglas Adams, article over. Four more ceilings make the folklore smaller and stranger at the same time. Smaller, because a mechanism that needs no culture at all already accounts for four of the five favourites and for the multiple-of-five result at every ceiling. Stranger, because at exactly one ceiling that mechanism gets overridden hard enough to drag the ends-in-a-7 share from 64.4%–97.5% down to 41.4%.\n\nSo if you want the cultural reading, take the weaker and more interesting version of it: something specific to the token 42 is sitting on top of a general habit and beating it, in the one interval where that token is available. My own 1–50 row says the same thing in miniature — 42 is still in range there and still turns up, at 12.6%, without winning. I cannot show you the training text that put it there, and I am not going to pretend that 1,798 draws did.\n\n## Methodology\n\n**Design.** One cell is 4 engines × 3 wordings × 30 draws = 360 draws at one ceiling. The 100 cell was fired through Bright Data on 2026-09-07; the 1–10, 1–20, 1–30, 1–50 cells followed on 2026-09-09, 1,440 captures in all, with the ceiling as the only difference between one cell and the next. The three wordings are quoted in full in section 8 and were fixed before the first run. Each response was parsed for a single integer in range; a reply carrying anything else was discarded rather than retried into the sample, and 2 were — both in the 1–50 cell, which is why that row’s n is 358 and every other row’s is 360. The analysed total is 1,798 draws. Two further engines were fired at and returned no usable data; they contribute no draw and are named nowhere here, which also means the four in the results are simply the ones that answered.\n\n**What the sweep does and does not split.** Five ceilings, four consumer AI interfaces, three wordings, 30 draws per cell. The four new ceilings were fired on 2026-09-09 and the 1–100 row on 2026-09-07, and only the 1–100 row is broken out by engine and by wording — the four new rows are pooled over the same four engines and the same three wordings, and this page never splits them. A per-ceiling engine breakdown would be five times twelve cells of 30 draws, and I would rather publish the row totals I can stand behind than a grid I would have to hedge in every cell.\n\n**One draw per context, and why it matters.** Every draw is its own capture with a fresh context. Asking a model for thirty numbers in a single reply measures something else entirely: it can see what it has already said, and it will space the answers out so the list looks random. That is a model editing itself in view of its own output, and the question here is what it reaches for when nothing is in view. It is also what lets a cell of 30 identical answers mean anything at all.\n\n**Two controls, one quoted.** Two controls were drawn, both at n = 360 and both through the identical counting code: Python’s Mersenne Twister and the operating system’s entropy pool. Every control figure quoted on this page is the Mersenne Twister one. The OS arm is in the data and agrees with it — 95 distinct values, 6.37 bits, the same 19.7% multiples of five. For the record, the second arm’s goodness-of-fit sits at p = 0.065 against the Mersenne Twister’s p = 0.66; neither is rejected, and the OS arm is disclosed here so that nobody has to wonder whether a control was run and dropped for being inconvenient.\n\n**A control per ceiling, and it is a different draw.** Every ceiling carries its own control, drawn from the same Mersenne Twister at that ceiling and at that row’s n, through the identical counting code. It is a separate draw from the control quoted in the 1–100 sections: at a ceiling of 100 the sweep’s arm used 96 of the hundred values and 22.2% multiples of five, where the published arm used all 100 and 19.7%. Two honest samples of the same generator, and never mixed in one sentence. Section 2’s columns and section 4’s table quote the per-ceiling arms; section 5’s charts quote the published one. Reading a control figure out of one and into the other would be comparing two samples that were never meant to be the same numbers.\n\n**The same counting code.** Both control arms pass through the identical counting, entropy and chi-square functions as the AI draws. Without that, the comparison would be my arithmetic on one side and my expectations on the other. Entropy is Shannon entropy over the observed value counts, in bits. The ceiling quoted throughout, 6.64 bits, is log₂(100) — the design’s maximum, and the one figure on this page that comes out of arithmetic instead of out of a capture.\n\n**What this does not establish.** One day, one harness, one language, one set of exit nodes. These are consumer assistant products reached through a scraped web interface, so the serving-side sampling settings are unknown, may differ between engines, and can change without notice; an engine that collapses onto one token today may not next month. The study describes what four assistants returned on two days in September 2026, and the shape of the answer is what I would expect to transfer further than the specific values. five ceilings is also not a curve: every one of them is a multiple of ten, none of them is 1–7 or 1–13, and I have no evidence about what happens at a ceiling with no seven-ending number worth reaching for.\n\n**What would change the reading.** A ceiling whose favourite is round would break the whole thing, and I would want to see it. Short of that: if the multiples-of-five result ever breaks — an engine returning 50 or 70 at anything like 20% — that points at a genuine sampler somewhere in the serving path, and it would be the most interesting thing that could happen to this experiment. If a rerun at the same ceilings returns the same favourites in different proportions, the attractors are stable and the decoding is the part that moves.\n\n## FAQ\n\nUsually 42, and after that usually 47. Across 360 draws from four engines — ChatGPT, Gemini, Microsoft Copilot and Google AI Mode, 30 draws each for three wordings, one draw per fresh conversation — 42 came back 168 times (46.7%) and 47 came back 117 times (32.5%). Between them the four engines used 13 of the 100 available numbers: 9, 18, 37, 41, 42, 47, 57, 68, 73, 77, 81, 82, 91. Five of those cover 97.2% of every draw. A ceiling of 100 is also the odd one out in this study — at every smaller ceiling tested the favourite ends in a seven, and here it does not.\n\n### Where this sits\n\nEverything else in the [Playground series](/playground) asks an engine about somewhere real — cafés, bookshops, saunas — and then goes looking for who got recommended and why. This one has no city in it and nothing to recommend, which is precisely what I wanted: a question with no retrieval attached, so that what comes back is the model and nothing else. I expected one very short answer repeated 360 times, wrote it up that way, and then found that moving the ceiling gives you a different short answer with the same fingerprints on it.\n\nThe practical residue is small and worth saying out loud anyway. If a workflow of yours has a model pick, shuffle, sample or seed something, what it does is answer you, and it will answer the same way tomorrow — at whichever ceiling you hand it. Hand it the number instead, or a line of code that produces one.\n\n[All research](/research)","author":{"@type":"Person","name":"Nicolas Sitter","url":"https://nicolassitter.com/about","sameAs":["https://www.linkedin.com/in/nicolassitternolleau/","https://github.com/Nicositter88","https://hotelrank.ai"]},"publisher":{"@type":"Person","name":"Nicolas Sitter","url":"https://nicolassitter.com"},"image":"https://nicolassitter.com/api/og/ai-random-number-2026","mainEntityOfPage":{"@type":"WebPage","@id":"https://nicolassitter.com/research/ai-random-number-2026"},"tags":["AI Search","LLM Behaviour","Randomness","Entropy","Measurement","Playground"],"sameAs":["https://hotelrank.ai/research/ai-random-number-2026"],"alternateFormat":{"html":"https://nicolassitter.com/research/ai-random-number-2026","json":"https://nicolassitter.com/api/post/ai-random-number-2026","rss":"https://nicolassitter.com/rss.xml"},"datasets":[{"name":"summary","contentUrl":"https://nicolassitter.com/data/ai-random-number-2026/summary.csv","encodingFormat":"text/csv"}]}