What does Dr. Jonah Tebaa on the AI Rollout That Won on Average, Lost $276K mean in practice?
Dr. Jonah Tebaa explains that an AI rollout lost $276,000 despite improving average resolution time by 25.4 percent because performance was evaluated solely on mean metrics. While routine tickets saved $84,000 in labor, complex edge cases slowed down, tripling $150 SLA penalty breaches from 1,200 to 3,600 and generating $360,000 in penalties. To prevent such hidden deficits, Dr. Jonah Tebaa advocates implementing a variance audit to segment cases, track full distribution percentiles, and price tail risks.
Most AI return-on-investment reports lead with one number: average handle time down, average cost per case down, average resolution time down. Dr. Jonah Tebaa's recent analysis of a support operation's AI rollout argues that this single number is precisely the wrong basis for a renewal decision — and the case he uses to make the point comes with numbers exact enough to check yourself.
A 25 Percent Win That Was Actually a $276,000 Loss
The operation in question handles 40,000 tickets a quarter. Before its AI copilot went live, average resolution time was 14.2 minutes, and the slowest 5 percent of tickets — the 95th percentile — took 42 minutes. A $150 SLA penalty applied to any ticket over 60 minutes, and roughly 3 percent of tickets breached that threshold: 1,200 tickets, $180,000 a quarter in penalties.
After the AI copilot began drafting agent responses, the average dropped to 10.6 minutes — a 25.4 percent improvement, and the number that reached the quarterly business review. What did not reach it: the 95th-percentile case rose to 71 minutes, a 69 percent increase.
According to Tebaa, both numbers are true and both come from the same mechanism. The AI drafts near-instantly for the roughly 80 percent of tickets that are routine, which is where the entire average gain originates. For the harder 20 percent — ambiguous, multi-issue, or policy-exception tickets — the draft arrives shaped for the wrong case. Agents read it, correct it, and still complete the manual work the ticket required, which adds a step rather than removing one. That imbalance is easy to miss because both effects are real; only one of them is visible on a mean-based dashboard.
Why the Financial Story Reverses
Tebaa walks through the arithmetic in dollar terms rather than minutes, and that is where the reversal becomes visible. The labour saved — 3.6 minutes per ticket across 40,000 tickets, at a $35 loaded hourly rate — comes to $84,000 for the quarter. But the SLA breach rate tripled, from 3 percent to 9 percent, which is 3,600 breaching tickets instead of 1,200: 2,400 additional breaches at $150 each, or $360,000 in new penalties.
Net the two figures against each other, in the same quarter: $84,000 saved minus $360,000 in new cost is a $276,000 net loss — the same quarter the dashboard reported a 25 percent win.
Tebaa's Variance Audit
Tebaa does not conclude that the AI tool is at fault; in his view it did what it was built to do. His conclusion is that the measurement was incomplete, and he offers a five-step check he calls the variance audit:
- Report the full distribution, not the mean alone — p50, p90, p95, and p99 wherever a contractual penalty applies, before any "average improved by X percent" claim is presented.
- Segment before averaging — split routine cases from exceptions, since a blended mean will always flatter a tool that is strong on the easy majority and weak on the hard minority.
- Price the tail in dollars — attach the actual cost of the worst-case threshold (SLA penalty, churn, fine, chargeback) and multiply by the change in breach count, not the change in average time.
- Net both figures on the same period — mean-driven savings against tail-driven cost, same quarter, same unit; without that single net number, there is no value claim yet.
- Re-run the audit quarterly — the tail shifts independently of the mean as case mix and behaviour change, so a programme that nets positive in month one can net negative by month four with no change to the underlying code.
The Question Tebaa Says CFOs Should Be Asking
The larger point in Tebaa's work is aimed less at any single support desk and more at how AI value gets reported to the people who approve its budget. A mean-based dashboard, he argues, will show a win even when a programme is a net financial loss, because the losses accumulate in a part of the distribution that dashboards rarely display. The average is a summary of performance, not a verdict on value — and treating it as one is the actual error, not the AI system itself.
For the COOs, VPs of customer operations, and CFOs currently deciding whether to renew or expand an AI rollout, Tebaa's advice is specific: ask what happened to the worst-case tail, and price it in dollars before the renewal is signed. Whether that number turns out favourable or not, Tebaa notes, is beside the point — the point is that without it, the renewal decision is being made on half the ledger.