Building a daily historical deduction game sounds simple until you realize how broken historical data actually is on the open web.
When we started Dawn of Time, our premise was strict: every figure must have five concrete, defensible clues. A birth year, an exact birthplace coordinate on a world map, an exact death place coordinate, a single primary role, and an archival portrait. If any of those five facts cannot be defended with primary or peer-reviewed records, the figure cannot enter the game.
Here is what happened when we tried to extract that data from public knowledge bases, why we discarded an entire 509-figure legacy dataset, and how we built a clean pipeline that serves 500 verified figures with zero database latency.
The centroid trap: why raw Wikidata coordinates fail gameplay
The earliest prototype of the game attempted to query Wikidata for coordinates using simple SPARQL queries.
That failed immediately. When historical figures lack a documented birthplace, automated data pipelines frequently fall back to the coordinate of the modern country or historical empire.
For example, an automated scraper might place an ancient philosopher in the geographic center of modern Turkey, or place an early monarch in the geometric centroid of Kazakhstan.
In a geographic deduction game where a single pin tells the player where someone lived, a country centroid is a lie. If a player sees a pin in rural central Spain, they assume the figure was born in that specific valley, not that an algorithm lazily defaulted to the geographic center of the Iberian Peninsula.
We established a non-negotiable rule: zero synthetic centroids.
Every coordinate in our dataset must represent a defensible, documented locality: a specific city, town, village, fortress, or monastery. If a historical figure is only known to have been born "somewhere in Gaul" or "somewhere in Ancient Egypt," that figure was disqualified. Quality and fairness beat artificial quota.
From 1,600 candidates to 500 verified figures
We began with a raw discovery pool of 1,600 unique human QIDs gathered from Pantheon notability indexes and Wikidata sitelink counts.
To make the game fair, a candidate had to survive six progressive filters:
- Human verification (`P31 = Q5`): Discarding mythological, fictional, and composite figures.
- Defensible birth and death years: Figures with disputed centuries or unknown lifespans were removed. If historical consensus accepts a circa year (such as Confucius at c. 551 BCE), we preserved the explicit circa qualifier.
- Locality-level coordinates: Both birth and death coordinates had to point to verified historical settlements.
- Taxonomic role mapping: Every figure had to fit into one of sixteen reviewed historical roles (Ruler, Military leader, Scientist, Philosopher, etc.).
- Authentic archival likeness: The candidate required a verified public domain or Creative Commons portrait. AI reconstructions and stock-photo recreations were banned.
- Cultural recognizability: We measured 90-day English Wikipedia pageviews to ensure figures possessed enough cultural footprint to make daily play fair.
Out of 1,600 candidates, 650 reached the shortlist. A manual review pass then filtered out ambiguous entries, leaving exactly 500 production-ready figures.
What we learned about portrait licensing
Image licensing on the web is fraught with subtle traps.
Many web games simply hotlink images from search engines or use low-resolution thumbnails without attribution. Because Dawn of Time is committed to museum-grade curation, every single image had to be mirrored locally with verifiable provenance.
Our final breakdown across the 500 figures includes:
- Public domain: 394 portraits (78.8%)
- Creative Commons licenses: 106 portraits (21.2%), including CC BY-SA 4.0, CC BY 3.0, and CC0
When a player solves a puzzle, the dossier reveals the exact license, creator credit, and Wikimedia Commons source page. But before the result is revealed, clue five serves only the raw image through an opaque storage key. We intentionally strip filenames, captions, and metadata from the payload so inspection tools cannot spoil the answer.
Serving 0ms SSR without database latency
Our original backend design called Supabase PostgreSQL on every request. While Supabase is reliable, fetching schedule rows and clue snapshots over the network added between 80ms and 220ms of round-trip time on cold starts.
For a daily web game, 200ms is perceptible lag.
Because our 365-day schedule and 500-figure corpus are frozen and verified before release, we pre-bundled the static schedule and target snapshots into server memory.
When a player opens the homepage:
- Next.js server components read the current UTC date.
- The puzzle target is retrieved from in-memory cache in under 1 millisecond.
- The server renders only Clue 1 (the birth year) directly into the initial HTML.
- Subsequent clues are fetched client-side per stage as the player progresses.
This gives the user an instant first paint while ensuring clues 2 through 5 and the answer identity never leak in the initial document source.
Dispatches from the cartographer
Curating historical data taught us that history is rarely neat. Scholars disagree on dates, borders shift across empires, and artistic depictions change over centuries.
If you spot an anomaly in any puzzle, or if you are a history educator using Dawn of Time in your lectures, please send a note to our curator desk at ahsanriaz8000@gmail.com. We treat historical accuracy as an ongoing discipline.