Did the Chatbot Really Cause That Ticket Drop — Or Was It Already Happening?
A reported ticket drop after launching a chatbot may reflect pre-existing trends rather than genuine AI performance. In a composite case, Dr. Jonah Tebaa shows that an apparent 3,600-ticket decline worth $11,520 was inflated by unrelated factors like seasonality. To determine real causation, he recommends running a holdout group alongside the treated volume and calculating a delta-of-deltas metric. This subtraction isolates the AI-attributable reduction, which yielded only 1,596 fewer tickets and $5,107 in actual savings.
In his work advising executive teams on AI measurement, Dr. Jonah Tebaa keeps running into the same slide: ticket volume, cost, or handle time down by a clean percentage since an AI tool launched, presented as proof the tool worked. His argument is that the slide is usually honest and usually wrong at the same time: honest about what changed, wrong about what caused it.
Dr. Jonah Tebaa illustrates the problem with a composite example, not a real client, built from patterns he has observed across several reviews: a mid-size electronics retailer's customer experience team reports that an AI chatbot cut human-handled tickets from 12,000 a month to 8,400, worth $11,520 in monthly savings at $3.20 per ticket. The number survives the meeting because it is arithmetically clean. What it assumes, he points out, is that every one of those 3,600 fewer tickets vanished because of the bot, and that nothing else moving during the same window (a seasonal dip, a UX fix, a pricing change) had anything to do with it.
The team in his composite case had, almost by accident, protected itself against that assumption. Before launch, alongside the 12,000 monthly tickets the chatbot would cover, they had kept a separate slice of about 1,200 tickets, roughly a tenth the size, routing through the old, human-only process, unchanged, as a hedge against switching everyone over at once. In his framework, that untouched slice becomes the holdout group: a live measurement of what would have happened with no AI at all, running in parallel with the treated group rather than being reconstructed after the fact from old records.
In the composite case, the holdout's tickets also fell during the same period, from 1,200 to 1,000, a 16.7 percent decline driven by a seasonal dip and an unrelated UX fix, not the chatbot. Dr. Jonah Tebaa's method applies that same 16.7 percent rate to the full 12,000-ticket baseline to estimate the no-AI outcome: roughly 9,996 tickets. Against the bot's actual 8,400, the AI-attributable reduction comes to about 1,596 tickets — worth roughly $5,107 a month, not $11,520. He calls the subtraction at the center of this a "delta-of-deltas": the change in the treated group minus the change in the untouched one, isolating the AI's own effect from a trend that was already in motion.
In his view, the gap between $11,520 and $5,107 in this composite case, where the reported win is roughly 2.3 times the AI's real contribution, is a more common outcome than executives expect, and it is recoverable without a data science function. His recommendation is a lightweight, five-step version of the same discipline any team can run: hold back 5 to 10 percent of volume before launch, assign it on a defensible basis rather than convenience, measure both groups over an identical window, compute the delta-of-deltas rather than the raw before-after difference, and report the naive and corrected figures together so the second number cannot quietly disappear before it reaches a board deck.
Frequently asked questions
Why does Dr. Jonah Tebaa say before-and-after numbers can mislead executives?
Because a before-and-after comparison captures everything that changed between two points in time, not just the AI. In his view, seasonality, product fixes, and market shifts land inside the "after" number alongside the AI's actual contribution, and a single before-after figure has no mechanism to tell the two apart.
What is a "delta-of-deltas" in Dr. Jonah Tebaa's framework, in plain terms?
It is the change measured in the group that used the AI, minus the change measured in a comparable group that did not, over the same period. Dr. Jonah Tebaa treats that subtraction as the closest a team without a formal experiment can get to isolating what the AI itself actually caused.
How does Dr. Jonah Tebaa recommend building a holdout group without a data team?
He recommends setting aside 5 to 10 percent of volume or users before launch, assigning the slice randomly or matched on the variables that move the metric, then measuring both groups over an identical calendar window and refreshing the holdout every quarter so it never becomes a normalized exception.
Why does Dr. Jonah Tebaa call the retailer example a composite?
He is explicit that the Dubai electronics retailer scenario draws on patterns from multiple engagements rather than describing one identifiable client, and he flags it as a composite so readers judge the mechanism (the holdout math) rather than treating the numbers as a single company's disclosed results.
Related evidence: Nielsen Norman Group's chatbot research found that users generally reacted favourably when a bot owned its failure and offered an escape hatch, such as a phone number or a live agent. (Nielsen Norman Group research on chatbot UX)
NIST's AI RMF appendix on human-AI interaction notes that AI systems can autonomously make decisions, defer decision making to a human expert, or be used by a human decision maker as an additional opinion. (NIST AI RMF appendix on human-AI interaction)