The certificate

The numbers, including the bad ones.

This page is generated from the engine's own calibration artefacts. It is not marketing copy and it is not hand-maintained. When the measurement is poor, the poor measurement is what appears here.

live from the engine

Last calibration run

Measured against known answers

2026-06-02 · n=3 fixtures · run 97 days ago

67%
Owner resolved
uncalibrated

2/3

Did we identify the right human behind the business at all.

0%
Metro matched
uncalibrated

0/3

Was the contact anchored to the subject's actual metro.

0%
Exact cell
uncalibrated

0/1

Did the top-ranked number equal the known-correct cell.

0%
Qualified
uncalibrated

0/3

Phone and email both present and above the confidence gate.

Read this before quoting the numbers above. With <10 graded cases and/or <3 known cells, these are FRAMEWORK baseline numbers, not an accuracy verdict. Add Rob's confirmed cells to make it real.

What closes the gap

The label loop is the whole product.

SignalNowWhy it matters
Hit findings19,403The engine works. It has generated real output across real cases.
Ground-truth labels862Each one closes the loop a little further. EM, the Bayes-net CPTs and conformal coverage all feed on these.
Labeled as subject97Labelled as the subject — ALL machine-derived, none confirmed by a human. Feeds the Fellegi-Sunter EM re-fit.
Labeled as not subject765Labelled as not the subject; 1 of 862 labels in total came from a human. Without these, precision cannot be measured at all.
Measured precision11%Measured against the labeled set.
Human-confirmed positives0 of 97Every positive above was machine-derived. 1 label in this corpus came from a person, and it was a negative. A precision resting entirely on labels the system generated for itself is circular, and it is labelled auto for that reason.
Measured, not asserted

How often the face matcher calls two strangers the same person.

0
False-match rate

0 of 3,569 impostor pairs

At the operating threshold of 0.6308. The closest pair of genuinely different people in the corpus sits at 0.7421, and the threshold is set below it on purpose.

True-match rate

not measurable on this corpus

every subject holds one photograph; 1 genuine pair(s) exist in the corpus, which cannot support a true-match rate. A number here would be invented, so there is not one.

85
Faces indexed

across 85 distinct subjects

Small enough that this is a statement about a matcher, not about a product. It is published at this size rather than held back until it flatters.

Read this before quoting it. The corpus is published obituary photographs — the subjects are deceased, which is the right place to measure a biometric matcher before pointing one at a living person. It is recomputed from the index on every page build rather than stored, so it cannot drift from what it describes. Face evidence is not wired into any trace today; this measures the matcher, not the product.

The right-party rate

A number came back. Was it theirs?

200 people across 5 runs, drawn at random from a state voter roll of 9.2 million registrants, traced with that roll switched off, then scored against what the state holds. A government source, independent of every broker this pipeline reads. Both runs are pooled — publishing the better of the two would be cherry-picking.

94%
Produced a phone

Retrieval is not the problem.

32%
Right number found at all

Anywhere in our candidate set. Spread across runs: 25–40%.

13%
Right-party rate

It was our TOP answer. Ranking, not retrieval, is now the bottleneck.

56.7s
Median trace

Per subject, free tier, one at a time.

A miss is weaker evidence than a hit. The roll holds ONE number per registrant, given to a county on some past date and often a landline, so coverage here is a floor rather than a verdict. Whether our other answers are live cells cannot be settled offline — libphonenumber returns FIXED_LINE_OR_MOBILE for essentially every US number, so it takes a carrier lookup.

The reliability curve

What it claimed, against what happened.

Every finding carries a stated confidence band. This is the only question that matters about those bands: when the engine said strong, was it?

The curve is one point, n=536. 536 of 862 labels sit in a band that stated a probability. The rest carry no claim, so there is nothing about them to check — they are listed below for completeness, not as evidence.

Stated bandClaimedObservednGap
MEDIUM60%18%53642 pts off — under-delivered
EXCLUDED0%325no stated prior for this band
(NONE)100%1no stated prior for this band

This curve is not yet evidence. The bands are populated almost entirely by machine-derived labels, and the largest band that states a probability holds 536 observations — too few to distinguish a well-calibrated engine from a lucky one. It is published because a curve you can check beats a number you cannot. Brier 0.3242, scored against the stated confidence band (STRONG/MEDIUM/LEAD priors).

Owned data

The corpus, stated plainly

48,561,526
Parcel records

planner estimate, ±a few percent

Owner name + address, farmed from public sources. Ours outright.

138
Phones in corpus

exact count

Small, and real. The retrieval path is reconnected — this number is proof it works, not proof of scale.

16
Emails in corpus

exact count

Same path, same caveat. Growing only from real traces, never seeded.

8 / 8
OSINT tools healthy

Self-hosted and live. Fetched from the orchestrator health endpoint on this page load.

✓ phoneinfoga ✓ spiderfoot ✓ maigret ✓ sherlock ✓ holehe ✓ webcheck ✓ theharvester ✓ redis