How our lottery statistics are calculated
The arithmetic is reproducible, but its meaning depends on what was counted. These notes explain the sample, the comparison model, and the safeguards against turning ordinary variation into a prediction.
- Written and reviewed by:
- Reysa Technologies
- Published:
- Last reviewed:
Define the sample before counting
Every analysis begins with a filtered set of imported draws. Scheme, date range, prize rank, matching type, and major-prize filters can change that set. The page shows the active period so two readers can reproduce the same selection.
Consolation duplicates are excluded where the question is about independently drawn endings. Entries without a location remain valid for digit analysis but cannot support a district comparison. We do not silently fill missing fields with guesses.
Observed frequency is a count, not a forecast
A digit-frequency table counts how often each digit occurs overall or in each position. An ending-frequency table counts the final four digits as one value. Percentages divide those observed counts by the number of eligible observations in the selected sample.
High and low rows are descriptive labels within that sample. Independent random draws do not remember the table. A frequently observed ending is not made more likely, and an unseen ending is not due to appear.
Expected counts come from a comparison model
For single digits, the simple model gives each digit one chance in ten at each position. For four-digit endings it gives each of the 10,000 possible endings an equal chance. Expected counts are the sample size multiplied by those model probabilities.
The model is a reference line, not proof about the draw machinery. Prize structures can contribute different numbers of observations per draw, and changes in schemes or archived coverage can alter the sample. We state these limits beside the comparison.
What the statistical tests ask
A chi-square calculation asks whether category counts differ from an equal-frequency model by more than ordinary sampling variation might suggest. Serial correlation and autocorrelation ask whether neighbouring observations move together in a consistent linear way.
A test result is sensitive to sample size, ordering, and repeated testing. It does not identify a cause, prove manipulation, or predict the next number. A small p-value is a reason to inspect data and assumptions, not a winning strategy.
Limits and reproducibility
Results are only as complete as the imported archive. PDF layout changes, scanning errors, missing source documents, and later official corrections can affect a calculation. Generated snapshots are invalidated after imports or corrections and rebuilt from the stored records.
Filters, sample sizes, update times, and data-quality status are displayed so readers can challenge the work. CSV exports support independent checking, but anyone republishing a conclusion should also preserve the filters and retrieval date.