Methodology
How results enter the archive, what we do with them, and where the data has limits.
1. Collection
We collect only completed, officially published draws. Future dates are never imported. Requests are spaced apart, identify this application, and respect robots.txt. Access controls and CAPTCHAs are never bypassed; an administrator uploads the published document when automated access is unavailable.
2. Storage and provenance
Each source document is stored outside the public web root with its original filename, type, size, download time, and SHA-256 checksum. The checksum prevents duplicate imports and later proves which file a record came from. Extracted text is retained for audit.
3. Parsing
Text from PDF, HTML, or plain-text sources is read by rules that identify the draw, date, venue, prizes, matching rules, and winning numbers. Parser version 1.0.0 is saved with every draw.
The parser does not guess. Ambiguous values create warnings, unclassified lines are retained, and low-confidence results are marked for human review.
4. Number handling
Ticket numbers remain text everywhere. This preserves leading zeros: 000056 and 56 are different tickets. We keep the printed value, a normalised value, its digits, and the last four digits.
Prizes are matched either to the whole ticket, including its series, or to the last four digits. The archive records which rule each prize tier used.
5. Analysis
Statistics are calculated in chunks and cached until data changes. Standard tests include chi-square, Shannon entropy, the Wald–Wolfowitz runs test, serial correlation, autocorrelation, and Wilson score intervals. Small samples are clearly flagged.
6. What this archive cannot tell you
For important limitations, see the Full disclaimer.
7. Known limitations
- Coverage is incomplete. The archive contains imported draws, not necessarily every draw ever held. Current coverage runs from 07-08-2017 to 05-09-2026.
- Location data is partial. Districts are published only for some prize tiers, so geographic analysis covers a biased subset.
- Source layouts change. Older result documents may parse less cleanly than recent ones.
- Image-only PDFs are skipped because the importer does not use OCR.
- This is not an official source. Always verify against the published official result.