The numbers, including the bad ones.
This page is generated from the engine's own calibration artefacts. It is not marketing copy and it is not hand-maintained. When the measurement is poor, the poor measurement is what appears here.
live from the engine
Measured against known answers
2026-06-02 · n=3 fixtures · run 97 days ago
2/3
Did we identify the right human behind the business at all.
0/3
Was the contact anchored to the subject's actual metro.
0/1
Did the top-ranked number equal the known-correct cell.
0/3
Phone and email both present and above the confidence gate.
Read this before quoting the numbers above. With <10 graded cases and/or <3 known cells, these are FRAMEWORK baseline numbers, not an accuracy verdict. Add Rob's confirmed cells to make it real.
The label loop is the whole product.
| Signal | Now | Why it matters |
|---|---|---|
| Hit findings | 19,403 | The engine works. It has generated real output across real cases. |
| Ground-truth labels | 862 | Each one closes the loop a little further. EM, the Bayes-net CPTs and conformal coverage all feed on these. |
| Labeled as subject | 97 | Labelled as the subject — ALL machine-derived, none confirmed by a human. Feeds the Fellegi-Sunter EM re-fit. |
| Labeled as not subject | 765 | Labelled as not the subject; 1 of 862 labels in total came from a human. Without these, precision cannot be measured at all. |
| Measured precision | 11% | Measured against the labeled set. |
| Human-confirmed positives | 0 of 97 | Every positive above was machine-derived. 1 label in this corpus came from a person, and it was a negative. A precision resting entirely on labels the system generated for itself is circular, and it is labelled auto for that reason. |
How often the face matcher calls two strangers the same person.
0 of 3,569 impostor pairs
At the operating threshold of 0.6308. The closest pair of genuinely different people in the corpus sits at 0.7421, and the threshold is set below it on purpose.
not measurable on this corpus
every subject holds one photograph; 1 genuine pair(s) exist in the corpus, which cannot support a true-match rate. A number here would be invented, so there is not one.
across 85 distinct subjects
Small enough that this is a statement about a matcher, not about a product. It is published at this size rather than held back until it flatters.
Read this before quoting it. The corpus is published obituary photographs — the subjects are deceased, which is the right place to measure a biometric matcher before pointing one at a living person. It is recomputed from the index on every page build rather than stored, so it cannot drift from what it describes. Face evidence is not wired into any trace today; this measures the matcher, not the product.
A number came back. Was it theirs?
200 people across 5 runs, drawn at random from a state voter roll of 9.2 million registrants, traced with that roll switched off, then scored against what the state holds. A government source, independent of every broker this pipeline reads. Both runs are pooled — publishing the better of the two would be cherry-picking.
Retrieval is not the problem.
Anywhere in our candidate set. Spread across runs: 25–40%.
It was our TOP answer. Ranking, not retrieval, is now the bottleneck.
Per subject, free tier, one at a time.
A miss is weaker evidence than a hit. The roll holds ONE number per registrant, given to a county on some past date and often a landline, so coverage here is a floor rather than a verdict. Whether our other answers are live cells cannot be settled offline — libphonenumber returns FIXED_LINE_OR_MOBILE for essentially every US number, so it takes a carrier lookup.
What it claimed, against what happened.
Every finding carries a stated confidence band. This is the only question that matters about those bands: when the engine said strong, was it?
The curve is one point, n=536. 536 of 862 labels sit in a band that stated a probability. The rest carry no claim, so there is nothing about them to check — they are listed below for completeness, not as evidence.
| Stated band | Claimed | Observed | n | Gap |
|---|---|---|---|---|
| MEDIUM | 60% | 18% | 536 | 42 pts off — under-delivered |
| EXCLUDED | — | 0% | 325 | no stated prior for this band |
| (NONE) | — | 100% | 1 | no stated prior for this band |
This curve is not yet evidence. The bands are populated almost entirely by machine-derived labels, and the largest band that states a probability holds 536 observations — too few to distinguish a well-calibrated engine from a lucky one. It is published because a curve you can check beats a number you cannot. Brier 0.3242, scored against the stated confidence band (STRONG/MEDIUM/LEAD priors).
The corpus, stated plainly
planner estimate, ±a few percent
Owner name + address, farmed from public sources. Ours outright.
exact count
Small, and real. The retrieval path is reconnected — this number is proof it works, not proof of scale.
exact count
Same path, same caveat. Growing only from real traces, never seeded.
Self-hosted and live. Fetched from the orchestrator health endpoint on this page load.
✓ phoneinfoga ✓ spiderfoot ✓ maigret ✓ sherlock ✓ holehe ✓ webcheck ✓ theharvester ✓ redis