Can AI guess which number will come up in the lottery?
Some media outlets in Costa Rica use that question to attract readers. At Caja de Arena, we set out to experiment with the premise and these are our results.
.jpg)
No, end of story. Still, we did run the test anyway, but...
Why ask the obvious?
A few days ago we published in Sandbox an experiment titled "The country that asks the machine for the winning number" where we analyzed 9,025 notes from the Costa Rican press published between January and September 2026 and extracted via www.mihaiku.com. Of the 1,472 that mentioned artificial intelligence, 61 were lottery and football predictions. This group turned out to be the most anchored in the country, with national actors in 60.7% of cases, and the least questioned: none of those notes was classified with a negative tone.
The general conclusion of the analysis was that the press transcribes more than it questions. This experiment asks the question that was missing from that framework: whether the machine really guesses. The answer is known before you start, but knowing it is not the same as proving it with the country's official data and with the notes that actually circulated. That's why we measured it.
The experiment has three parts. First, we contrasted each number that the press attributed to an AI with the official draw result. Second, we built 32 prediction methods, from frequency counts to a neural network, and pitted them against chance. Third, we let the winning model predict 2026 on the same dates the press did.
The weekly oracle
We catalogued 50 notes that published numbers attributed to an artificial intelligence for drawings by the Social Protection Board (JPS), between December 2024 and September 2026. La Teja published 46 of them, NCR Noticias three and El Financiero one. Since late 2025, La Teja turned the format into an almost fixed section for National Lottery Sundays and Chances on Tuesdays and Fridays.
See the data
The notes rarely say which artificial intelligence they used. Only one of the 46 from La Teja names it in the body of the text; in the others, ChatGPT or Gemini appear only in the image credit. Almost all close with the legend "Note made with the help of AI": the tool that supposedly chooses the number also writes the note that announces it.
One hundred six numbers, one hit
The 50 notes published 106 numbers. We contrasted all of them with the official first prize of each draw. There was one hit. A person choosing numbers completely at random, in the same draws and in the same quantity, would have hit on average 1.06 times. The press hit exactly what chance hits.
See the data
The only hit deserves context. On November 18, 2025, a note published fifteen different numbers for the same Chances draw, and one of them was 32, which came out that night. Covering fifteen of a hundred numbers gives a 15% probability of hitting any one. None of the 42 combinations of number and series published, which is what pays the jackpot, matched the result.
The fingerprint of the machine
The published numbers are not distributed as chance would distribute them. The 27 appeared in seven different notes and series 318 in five. We simulated 10,000 catalogs in which each note chooses at random, and such a repetition appears with a probability of 0.3% for the 27 and 0.01% for the 318. Forty of the hundred possible numbers were never published.
See the data
That concentration says nothing about the draws; it says something about the tool. A text generator that "chooses at random" repeats its own preferences, and writing that consults it every week accumulates them. The combination 47-318 was published for two different draws, in December 2025 and in May 2026. The 47, moreover, did not come out in any ordinary draw between 2019 and 2025.
What the headline promises
Of the 50 notes, 36 clarify in the body that the combination has no more probability than any other. That clarification coexists with justifications that contradict it. A June note states that no calculation can anticipate the result and, further on, that "numbers in the same family tend to maintain a certain inertia." Others invoke statistical patterns, rarely repeated combinations or numerology.
See the data
Fourteen notes do not clarify anywhere that the result depends on chance, and five headlines directly suggest that artificial intelligence knows what will come out. The distance between headline and body explains the neutral tone that CA-001 found in this framework: the note does not state anything false literally, but the headline reader receives the promise without the clarification.
Thirty-two ways to try
The second part of the experiment abandons the press and tests the question seriously. We used the official results published by JPS on its open data portal: 247 ordinary National Lottery draws and 6,325 Nuevos Tiempos draws between 2019 and 2025 (Social Protection Board [JPS], 2026). Each game was analyzed separately.
We built 32 prediction methods in nine families. Among them are popular intuitions, such as hot numbers, cold ones or the repetition of the last result; statistical methods, such as smoothed frequencies and Markov chains; and machine learning, with logistic regression and a neural network. Each method assigns a probability to each of the hundred numbers using only previous draws.
See the data
The rules were set in writing before touching the data. The oldest 60% of the draws were used to train, the next 20% to choose the best method and the most recent 20% was locked away for a single final examination. The metric was the advantage in log loss against chance: the more probability the model assigns to the number that actually comes out.
Nuevos Tiempos: chance won the competition
In Nuevos Tiempos, with 1,265 validation draws, none of the other 31 methods beat random chance. Neither the neural network nor the hot numbers managed to assign more probability to the actual outcome than a uniform distribution. The competition's winner was, therefore, random chance itself, and in the final exam it tied with it by definition.
This result carries weight because the test had power. With 1,265 draws in the final exam, a 14% bias in the frequency of some number would have been detectable. None appeared. The preregistered verdict is that there is no evidence of predictability: seven years of official data contain no pattern these methods could exploit.
National Lottery: a winner that random chance manufactures
In National Lottery, the frequency method won the competition with a small advantage over random chance. In the final exam, with 50 draws, it maintained a minimal advantage whose confidence interval included zero. To know if that meant anything, we repeated the entire procedure, with the 32 methods and the same competition, on 1,000 simulated lotteries.
See the data
In 87% of those fair lotteries, some method beat random chance during the competition. Only 20% continued winning in the final exam. A winner like the real one appeared in one of every ten fair lotteries. With 247 draws it is not possible to distinguish a small bias from noise: the preregistered verdict is that the data are insufficient.
The cold numbers were already predicted
The pattern factory also operates on real data. In the 247 ordinary draws, ten numbers never came out and the 60 came out eight times, more than triple what was expected. It looks like a signal. However, a perfectly fair lottery with 247 draws produces on average 8.4 missing numbers, and the observed distribution fits what random chance predicts.
See the data
The 60 that never came out
The last test put the winning model against the press on the same ground. Without modifying it, we asked it to predict the 32 ordinary draws of 2026, from January 4 to September 27, each one with the information available before the date. None of those draws had participated in training, the competition, or the final exam.
See the data
The model recommended 60 as the most probable number on the 32 dates, for being the most frequent of the previous period. The 60 did not come out even once. In the 18 draws for which La Teja published a number, the press got zero right and so did the model. Its overall advantage over random chance was minimal and indistinguishable from noise: one of every seven random sequences equals it.
How much would have to be expected
The underlying reason is arithmetic. A lottery with one hundred possible numbers requires an enormous amount of draws to reveal a small bias. Detecting that a number comes out 3% more than it should requires about 27,500 test draws: 528 years of National Lottery at a rate of one draw per week, or 30 years of Nuevos Tiempos.
See the data
The second question
No artificial intelligence, neither the press's nor ours, guessed the lottery. The useful finding is another: the procedure that produces headlines—consulting many times and keeping what seems to work—manufactures convincing patterns from pure chance. A media outlet that asks the second question doesn't need more technology, but to separate the data with which it chooses from the data with which it verifies.
Methodology
Press releases. The catalog combines 39 notes found in the CA-001 corpus through a documented filter and 11 notes found by web search. Each number was transcribed from the literal text of the note. Official results come from the JPS open data historical reports, from the official lists of each draw, and from a compiled 2026 results verified against them.
Test against random chance. For the press, an exact null distribution was used that treats the draw, and not the note, as an independent unit. For the models, a synthetic twin was used that repeats the full selection on fair lotteries, so that the p-value discounts the search among 32 methods. Repetitions between notes were evaluated with 10,000 simulations.
Controls. The protocol, data cuts, and verdict rule were versioned before training. The final exam was executed once behind a lock that records each access. An error in three secondary metrics forced a second authorized reading, which verified that the primary metric and verdicts did not change. Both readings are in the record.
Limitations. The catalog does not include all published notes, but only those found in the corpus and search. The parameters of each game were verified by observation of official data and not against regulations. The 2026 comparison between press and model is descriptive: with 18 draws, no difference would be distinguishable. This analysis does not recommend numbers or gambling strategies.
Reproduction. Code, data, protocol, and decision record are in the experiment repository. The figures in this text are regenerated with uv run jps-lottery build-report. The data for each figure are in the folder datos/ and the list of the 50 notes analyzed, with their links, in the 50 notes analyzed.
References
Ávalos Elizondo, M. (2026, September 14). The country that asks the machine for the winning number (Experiment CA-001). Caja de Arena. https://www.cajadearena.com/experimentos/el-pais-que-le-pregunta-a-la-maquina-por-el-numero-ganador
Junta de Protección Social. (2026). Reports of favored numbers 2019–2025 from National Lottery, extraordinary draws, Popular Lottery, and Nuevos Tiempos [Data sets]. https://www.jps.go.cr/transparencia/datos-abiertos
Source: own analysis with data from the Junta de Protección Social and the CA-001 corpus.
Technical details
- CA-007
- Experiment
- Completed
- 9 min
- Python · Vega-Lite
- The press published 106 numbers and got one right; chance would have gotten 1.06. None of the 32 methods we tested beat chance.