Genealogy Genie — Documentation
Public documentation for Genealogy Genie, including how we measure the quality of our AI research
and what those measurements currently say.
AI Quality Scorecard
→ Current scorecard
Every build, an offline judge inspects the research output the AI actually produced — the work log
it wrote, the sources it cited, the people it identified, the insights it drew — and grades it
against one shared rubric. The scorecard is that result, published unedited.
It is written to be checkable rather than reassuring:
- One rubric for every category. The same seven questions are asked of each surface, so the
grades are comparable to each other. Where a question does not apply, the cell reads NA and says
why, rather than being quietly dropped.
- Grades and certainty are separate. A letter says what we observed; a certainty rating says how
much weight it can bear. A result from a small sample is marked as such instead of being presented
as settled.
- What we borrow is named, and so is what we do not implement. The rubric is modelled on
published evaluation work. Where our check is weaker than the paper it is modelled on, the
scorecard says so in that row.
- Our own measurement failures are published too. When a category is ungraded because our
instrument is wrong rather than the AI, it appears as NA with the reason attached.
The card carries its own method section — read that before the grades.
A note on what this is not
It is not a marketing figure or a rounded-up headline. Grades move down as readily as up, categories
drop out when we cannot measure them honestly, and corrections are made in place rather than
quietly. If a number here matters to you, the definition behind it is the one written in the
scorecard — not the one in whatever standard it is modelled on.