What does The Denominator Problem: What a 70 Percent AI Claim Missed mean in practice?
A claimed 70 percent AI cost reduction missed the expensive rework on ambiguous cases excluded from the denominator, according to Dr. Jonah Tebaa. Measuring only 7,000 clean invoices at $1.35 ignored 3,000 flagged items requiring human correction at $6.20 each against the original $4.50 manual baseline. To establish defensible figures before making staffing decisions, Dr. Jonah Tebaa advocates calculating a blended figure across the entire workload, revealing the true overall efficiency gain of 37.7 percent.
When a finance team celebrates a 70 percent cost reduction from a new AI rollout, Dr. Jonah Tebaa's first question is not whether the number is impressive. It is what the number was divided by.
In his work advising executives on AI performance claims, Dr. Tebaa has found that the most common error is not in the technology. It is in the arithmetic used to describe it, and specifically in what gets left out of the denominator before a percentage is calculated.
A familiar shape of good news
He points to a composite example drawn from a pattern he sees recur across industries: a company processing 10,000 invoices a month. Fully manual coding costs $4.50 per invoice, $45,000 a month in total. After introducing AI-assisted coding, 7,000 of those invoices are handled straight through at $1.35 each, a genuine efficiency gain. Compare those two figures directly and the reduction is 70 percent, a number strong enough to justify cutting headcount.
Dr. Tebaa's argument is that this comparison, while arithmetically correct, describes only part of the workload.
The cost that did not make the report
The remaining 3,000 invoices in his example are not free successes waiting to be counted. They are flagged for human review because the AI's output was ambiguous or incomplete, and correcting a partial result, he notes, is often slower than coding the invoice from scratch. In his scenario that rework costs $6.20 per invoice, more than the original manual baseline.
That figure rarely appears in the report that reaches a budget committee. The 70 percent reduction was calculated only across the invoices that went through cleanly. The population that pushed into rework was excluded from the denominator, not out of deception, Dr. Tebaa argues, but because it is the more natural number to reach for: compare the new process to the old one only where the new process actually finished the job.
He observes the same shape wherever a system sorts work into an easy pile and a hard pile, from support ticket deflection to contract review to claims triage. The easy pile supplies the case study; the hard pile supplies the true cost of full deployment.
The gap between what a team believes an AI rollout delivered and what it measurably delivered has been tested directly elsewhere. In a randomised controlled trial of experienced open-source developers, participants estimated afterwards that the tools had sped them up by roughly 20 percent, while the trial itself found that when developers use AI tools, they take 19% longer than without, the same reversal Dr. Tebaa describes, arrived at from the other end.
Recalculating the true number
His corrective is to insist on a blended figure across the entire workload rather than the successful subset. Blending 7,000 invoices at $1.35 with 3,000 at $6.20, across all 10,000, produces a true cost of $2.805 per invoice. Against the $4.50 baseline, that is a 37.7 percent reduction, roughly half the size of the original claim, though still a genuinely strong result worth expanding.
The gap between the two figures, in Dr. Tebaa's view, is exactly where budget decisions go wrong. A headcount plan sized to a 70 percent efficiency gain will overcommit by close to double, leaving the invoices that still require a trained person to untangle them without anyone available to do the work.
Survey evidence suggests the blended figure is closer to the norm than the headline one. Stanford's AI Index reports that while roughly half of organisations using AI in service operations record any cost saving at all, most of them report cost savings of less than 10%, which makes a 70 percent figure less an outlier than a sign that something was left out of the denominator.
Four questions before the number becomes a decision
He recommends a short set of questions before any executive accepts an AI performance figure into a budget or staffing decision:
- What is the denominator: the full workload, or only the cases that went through cleanly?
- Does the number include the cost of rework, escalation, and human correction downstream?
- Is the baseline it is compared against measuring the same population, over the same time window?
- Has the number been checked again after a full quarter, or is it still the pilot-week figure?
None of these, Dr. Tebaa points out, require technical fluency. They require the discipline to ask what was excluded before approving what was included. In his experience, that discipline is what separates a defensible AI investment from an overstated one, and a team reporting its AI performance honestly will show the denominator without being asked.