Why Does a 94%-Accurate AI Pilot Fail in Production, and What Fixes It?
A 94%-accurate AI pilot fails in production because live inputs diverge from curated pilot data, not from model defects. In Dr. Jonah Tebaa's composite logistics example, photo-checking accuracy dropped from 94% in testing to 61% under real dock conditions. Rather than retraining, the fix adds an intake-normalization step to screen blur, crop, orientation, and lighting before scoring. Dr. Jonah Tebaa evaluates this vulnerability before scoping builds by applying a three-question intake audit that interrogates input channels, failure rates, and rejection ownership.
Dr. Jonah Tebaa's clients keep bringing him the same kind of number, and it keeps disappointing them the same way. A pilot for an AI tool clears the room at 90-something percent accuracy, the board signs off, and within weeks of going live the number has fallen by a third. The usual reaction is to blame the model and retrain it, or to quietly widen the human-review queue. In his work, Dr. Tebaa argues that both reactions treat the wrong patient. The model is very often fine. What has changed is the input reaching it.
He illustrates the point with a composite example built from patterns he sees repeatedly across logistics and field-service clients in the MENA region: a proof-of-delivery photo checker built to confirm a package matched an order and auto-close the support ticket. The pilot corpus was 500 photos supplied by the operations team itself — well-lit, single package, correct framing — and against that corpus the tool scored 94% (470 of 500). That number justified the rollout. Once the system went live and drivers began forwarding real delivery photos through a messaging app at the end of their shifts, accuracy on a matched 500-photo production sample fell to 61% (305 of 500): 165 more misses than the pilot had predicted.
The Model Wasn't the Problem
In Dr. Tebaa's account, the drivers did nothing wrong. They photographed deliveries the way they always had, with the equipment they had, under the lighting the loading dock actually offered. What nobody had done, before the model was ever trained, was define what the production input channel would look like once it left the operations team's controlled desk and entered a driver's pocket. The pilot accuracy number was a true statement about a photo set the company had curated. It was never a forecast about the photo set the company would actually receive.
This is the distinction Dr. Tebaa returns to across his writing on AI deployment: a pilot result is a property of the channel it was measured against, not a portable fact about the model. Two different input distributions, run through the same weights, will produce two different accuracy numbers — and the gap between them tells you nothing about model quality. It tells you that nobody scoped the intake before scoping the build.
Closing the Gap Without Touching the Model
The fix Dr. Tebaa recommends does not involve retraining anything. In the composite example, the team spent roughly three weeks building an intake-normalization step: a lightweight pre-check for blur, crop, orientation, and lighting that runs before a photo ever reaches the scoring model. A photo that fails the check is never scored. Instead, an automated reject-and-resend prompt returns to the driver, asking for another shot before the ticket can close. On the same production sample, this single addition brought accuracy back up to 89% (445 of 500).
Dr. Tebaa is explicit that 89% is not a victory lap back to the pilot's 94%, and that it shouldn't be treated as one. A loading dock at the start or end of a shift will never match a controlled desk for lighting or composition, and a normalization gate can only screen out the worst inputs — it cannot manufacture pilot-grade conditions on a moving dock. What matters to him is that the remaining 11-point gap is now a known, monitored figure that someone in the organization owns, rather than a hidden 33-point miss that surfaces only when customers complain about tickets closed on bad data. In his framing, converting an unbounded risk into a bounded, visible one is the actual deliverable — not a number that flatters the original pilot deck.
The Audit He Runs Before Trusting Any Pilot Number
Out of this pattern, Dr. Tebaa has distilled a three-question intake audit that he now applies to any AI pilot involving a photo, scan, or document a customer or field employee has to capture themselves, before the build is ever scoped:
- Who actually captures the input in production, and with what device and behavior — not who supplied the pilot's test set?
- What share of real-world inputs would fail a basic quality gate for blur, crop, orientation, or file type?
- What is the reject-and-resend path once a bad input is caught, and who is named as its owner from day one?
None of the three questions require a data scientist to answer. They require someone in the room willing to interrogate the channel a pilot number came from before that number gets treated as a launch decision. In Dr. Tebaa's view, a 94% pilot accuracy figure is not dishonest. It is simply an answer to a narrower question than the one most rollout decisions actually need answered — and the businesses that ask the narrower question anyway are the ones that get surprised by the loading dock.