What does Dr. Jonah Tebaa on the AI Baseline Nobody Actually Recorded mean in practice?
Dr. Jonah Tebaa argues that organizations frequently distort AI ROI metrics because their pre-rollout baseline was merely recalled rather than archived. In a composite, illustrative case involving invoice triage, an unverified 38-minute baseline yielded an exaggerated 45 percent win, whereas the true contemporaneous baseline of 24.3 minutes revealed only a 13.6 percent gain. To resolve this provenance failure, Dr. Jonah Tebaa introduces the Baseline Capture Protocol, requiring teams to timestamp multi-week case data, assign an owner, and cross-check records before rollout.
In a recent piece on his site, Dr. Jonah Tebaa takes on a quieter failure mode in AI ROI reporting: not a miscalculation, but a missing piece of evidence. His argument is that a percentage improvement is only as trustworthy as the "before" number underneath it, and in most organizations he has observed, that before number was never actually captured. It was recalled.
Dr. Tebaa illustrates the point with a composite, illustrative case, explicitly built from a recurring pattern rather than a named client, involving a finance operations team running AI-assisted triage on invoice exceptions. The team's claimed baseline, 38 minutes per exception, had circulated in planning documents for months. Nobody could trace it to a dated file, a sampling exercise, or an export. It existed only as something people remembered being told.
A 45 Percent Win That Wasn't
After deploying the AI-assisted system, the team measured a genuinely solid result: 21 minutes per exception, calculated from ticketing timestamps across 1,400 exceptions over 60 days. Set against the recalled 38-minute baseline, that produced an improvement of 44.7 percent, reported upward as a 45 percent win.
The story shifts when someone in finance goes looking for where the original 38-minute figure actually came from. Dr. Tebaa describes the team locating a contemporaneous dashboard export from the month the baseline was supposedly measured, one nobody had thought to check against the planning slide. The real average that month was 24.3 minutes per exception. The 38-minute number, it turned out, had been anchored to a single atypical week, one in which two staff members were out and the exception queue backed up well beyond normal.
Recalculated against the true baseline, the improvement comes to 13.6 percent, closer to 14. In dollar terms, at a fully loaded labor cost of $42 an hour, or $0.70 a minute, across roughly 700 exceptions a month, the claimed savings of $8,330 a month collapse to a real figure of $1,617 a month. The team had, in good faith, been reporting savings 5.15 times larger than what was actually occurring.
A Provenance Problem, Not a Math Problem
What makes this case distinct, in Dr. Tebaa's framing, is that nobody involved did anything obviously wrong. The arithmetic on both sides was correct. The AI system did produce a real improvement. The failure sits one layer upstream, in how the "before" number was established in the first place.
He points to a specific and predictable mechanism: human memory anchors on whatever was most vivid, not whatever was most typical. A disruptive week, with a visible backlog and an escalation email, is memorable. Eleven ordinary weeks around it are not. When someone later reconstructs "what things used to look like" from memory rather than from archived data, the vivid week is what gets recalled and reported as if it were representative.
Dr. Tebaa is careful to distinguish this from two related but separate arguments he has made elsewhere: a scope error, where the wrong slice of workload gets used as the denominator, and a translation error, where a genuine time saving gets mapped onto the wrong financial line. Both of those failures assume the starting figure was at least real. In this case, he argues, the starting figure was never a measurement to begin with, only a memory that had calcified into one by repetition.
The mechanism matters because it is not a matter of anyone acting carelessly in the ordinary sense. Nobody set out to inflate a result. The 38-minute figure had, technically, a source once, a quick glance at a difficult week's worth of tickets by someone under deadline pressure, but that source was never dated, never archived, and never distinguished afterward from a properly sampled baseline. By the time it reached a planning slide, it was functionally indistinguishable from real data, and it stayed that way until someone happened to ask where it came from.
The Baseline Capture Protocol
To prevent the pattern from recurring, Dr. Tebaa lays out a six-point protocol for capturing a defensible AI baseline before a rollout begins:
- Freeze and timestamp a pre-change window of four to six weeks, never a single week.
- Archive the raw case-level data itself, not a summary average or a planning slide.
- Log anomalies during the window as they occur, including staffing gaps, backlog spikes, and policy changes.
- Apply an identical metric definition on both sides of the before-and-after comparison.
- Assign one named owner for the baseline file, with a fixed storage location and retention date.
- Cross-check the baseline against an independent, contemporaneous source, such as a finance report or ops dashboard, before any ROI figure reaches a slide.
For organizations that are already past the rollout stage with no archived baseline to fall back on, his advice is to reconstruct from whatever proxy data survives, timesheets, staffing records, backlog logs, and to report a range with honestly stated uncertainty rather than a falsely precise point estimate. A range, in his view, survives scrutiny. A confident single number that cannot be traced to a source does not.
What He Wants Boards to Ask First
Dr. Tebaa's closing recommendation is procedural rather than technical: before any AI ROI percentage reaches a board deck, someone should be required to answer a single question on the record: what was the before number, where is it stored, and who signed off on it. In his view, that question, asked early and answered honestly, catches more bad ROI claims than any amount of scrutiny applied to the after number, because everything downstream of an unrecorded baseline inherits its error.
Related evidence: Article 50 of the EU AI Act requires providers to design AI systems that interact directly with people so that those people are informed they are interacting with an AI system, unless that is already obvious in the circumstances and context of use. (EU AI Act Article 50 transparency obligations)
The UK government's introduction to AI assurance defines AI governance as a range of mechanisms, including laws, regulations, policies, institutions and norms, used to outline processes for making decisions about AI. (the UK government's introduction to AI assurance)