What does The Matched-Pair Test: Dr. Jonah Tebaa on the Bias an AI Dashboard Hides mean in practice?
Dr. Jonah Tebaa shows that aggregate AI hiring dashboards hide significant sub-group disparities behind stable overall metrics. In a composite case he presents, an overall pass rate obscured an adverse-impact disparity ratio of 0.42, where identical resumes differed only by candidate origin names, yielding a 30 percent shortlist rate for Group A versus 12.5 percent for Group B. To uncover proxy penalties, Tebaa uses the four-step Matched-Pair Audit and the four-fifths rule threshold to systematically test whether automated screening systems treat identical profiles equally.
A hiring dashboard can report a healthy pass rate and still be blind to a 30-percent-versus-12.5-percent gap sitting one layer beneath it. That distinction — between what a system reports in aggregate and what it does to any one group within that aggregate — is the starting point for a method Dr. Jonah Tebaa uses in his work: the matched-pair audit.
An Illustrative Case, Not a Case File
Tebaa walks through the method using a composite scenario, not a named client engagement, and he is careful to flag it as such. The case is illustrative: a mid-size professional-services firm, roughly 1,200 applications a month, an AI resume screener cutting that volume to a shortlist of about 220 candidates — an 18 percent pass-through rate. On paper, the system looks well-tuned. Pass rate is stable. Time-to-shortlist has improved. Recruiter satisfaction is up.
Tebaa's point is that none of those figures answer the only question he considers load-bearing: whether the system treats otherwise-identical candidates the same way. A dashboard organized around throughput cannot surface that on its own. It has to be tested for, deliberately.
What Forty Matched Pairs Found
The test he describes is built on 40 matched pairs of resumes — identical stated degree, identical years of experience, identical skills — varying only the name at the top, chosen to signal a different national or regional origin. All 80 resumes go through the live system in the same week, against the same open requisitions, not a sandbox environment. The design is not original to Tebaa; it is borrowed from paired testing in civil-rights enforcement, where the US Department of Justice has since 1992 run a Fair Housing Testing Program that uses trained testers who pose as prospective renters, borrowers, or patrons for the purpose of gathering information about how a provider actually behaves.
In the illustrative case, Group A was shortlisted 12 times out of 40, a 30 percent rate. Group B was shortlisted 5 times out of 40, a 12.5 percent rate — a disparity ratio of 0.42. Tebaa uses the four-fifths rule as his working threshold, a figure borrowed from US Equal Employment Opportunity Commission adverse-impact guidance stating that a selection rate below 80 percent of the comparison group's rate warrants investigation, since a rate under four-fifths (or eighty percent) of the rate for the group with the highest rate will generally be regarded by the Federal enforcement agencies as evidence of adverse impact. He is explicit that this is a practical bright line rather than a codified legal standard across MENA jurisdictions, and does not claim otherwise. That distinction — treating a useful heuristic as a heuristic, not as borrowed legal authority — is itself part of how he frames rigor.
A Proxy Nobody Coded For
The obvious assumption is that the model is reading the name field directly. In the case Tebaa describes, it was not. A blind human review of resumes sitting in the borderline-score band traced the gap to a phrasing pattern common in CVs written by non-native English speakers — a pattern the model had learned to penalize independent of the candidate's actual English-proficiency score, which was already present elsewhere in the application. No one had coded a rule for national origin. The model had inferred a proxy on its own, and the dashboard had no mechanism to notice.
The Order That Makes It an Audit
The remediation, in Tebaa's account, was narrow: remove the phrasing feature's independent weight and let the existing, genuine proficiency score carry that signal instead. Re-running the same 40 pairs produced a shortlist rate of 30 percent versus 26 percent, a disparity ratio of 0.87 — inside the threshold set before the retest.
Tebaa places particular weight on the sequence in which the evidence was produced, not just its content. The file he describes contains four elements: a methodology memo written and dated before any results existed; a timestamped raw pre-remediation shortlist log; the 0.8 threshold decision, signed off before anyone saw an outcome; and the retest results themselves. He frames that ordering — bar set first, results examined second — as the feature that distinguishes an audit from a post-hoc rationalization. A threshold chosen after the numbers are known, in his framing, is not a threshold at all.
The Method, and Where Else It Applies
Tebaa names the approach the Matched-Pair Audit and describes it as four repeatable steps:
- Build. Construct pairs identical on every stated qualification, differing only in the one attribute suspected of driving the outcome.
- Run blind. Push the pairs through the live system, under real conditions, rather than a test environment.
- Set the bar first. Fix the pass/fail threshold before any result is examined.
- Document either way. Record the test, the threshold, and the outcome regardless of what it shows — a clean result counts as evidence too.
He extends the method past hiring by pointing to any system that scores, screens, or triages people or cases: credit scoring, claims triage, fraud flagging, support-ticket routing among them. In his framing, the pairs change with the domain; the discipline of the four steps does not.
Underlying the specific case is a broader position: for Tebaa, the risk in automated screening systems rarely announces itself. It sits inside a number that looks defensible until it is tested against a matched pair, and the difference between a system that is fair and one that merely looks efficient is whether anyone went looking before being asked to.