What does Dr. Jonah Tebaa on the AI Handoff Most Pilots Never Complete mean in practice?
Dr. Jonah Tebaa argues that the AI handoff most pilots never complete is transferring the unwritten, tacit operational judgment from the pilot creators to the permanent operations team. In a composite case, performance drops nineteen points because undocumented habits left with the creators. To resolve this, Dr. Jonah Tebaa introduces a six-item handoff package: an exception log, a written workaround list, named escalation contacts, baseline metrics, a freeze list, and shift-level correction authority.
Dr. Jonah Tebaa has spent much of the past year studying not whether AI pilots succeed, but what happens after they do. His latest analysis centers on a specific and often overlooked moment: the point at which the small team that built a working AI pilot hands daily operation to a separate, permanent team. In his view, this transition, not the launch itself, is where most AI deployments quietly lose ground.
To illustrate the mechanism, Dr. Tebaa uses a composite case, with identifying details changed, that he calls Meridian Retail. In the illustration, a mid-size retail chain assembles a three-person pilot team, a data scientist, a product manager, and an operations lead, to build and run an AI assistant that decides whether customer return requests should be automatically approved or routed to a human reviewer. Over eight weeks, the pilot reaches 91 percent approval accuracy against a trained human reviewer's judgment.
The Handoff Point
At the end of the pilot, Dr. Tebaa notes, the three-person team hands the system to a four-person operations team responsible for running it going forward. In the composite case, formal documentation covers only five of roughly thirty-five small judgment calls the pilot team had developed over the eight weeks. The other thirty exist only as tacit habits, absorbed by people who worked inside the system daily.
Two weeks after the handoff, accuracy in the illustration falls from 91 percent to 72 percent, a nineteen-point drop, without any change to the underlying model. Dr. Tebaa's point is precise: the system did not degrade. The people running it changed, and the habits that had kept it accurate left the building with the pilot team.
Among the undocumented behaviors in his composite example: re-triggering a weekly inventory sync that silently lagged after weekend processing, treating certain phrases in a customer's own refund explanation as an informal signal for human review even though no written rule said to, and manually adjusting return credit on product categories the model was known to consistently under-price. None of this, Dr. Tebaa argues, reflects carelessness on the pilot team's part. It reflects how quickly operational judgment becomes instinct rather than documentation when a small team lives inside a system every day.
His Central Argument
Dr. Tebaa's broader claim is that a pilot's performance is never purely a property of the AI system itself. It is partly a property of the specific people operating it during the pilot period. When an organization transfers only the technical system at handoff, and not the accumulated judgment behind it, he argues the transfer is incomplete regardless of how thorough the technical documentation appears.
He is careful to frame this as an operational and staffing problem rather than a governance question. His analysis explicitly avoids the language of AI autonomy levels or trust thresholds, focusing instead on a narrower, more mechanical claim: specific facts need to be written down, and specific people need to be named as responsible for specific failure types, before a pilot team disperses to its next project.
The Six-Item Handoff Package
To close this gap, Dr. Tebaa proposes a six-item handoff package that operations teams should receive before a pilot team moves on. It includes a detailed exception log that records the reason behind every human override rather than an aggregate override rate; a written workaround list turning informal habits into numbered operating steps; a named escalation contact, rather than a general ticket queue, for each class of failure in the first ninety days; a documented baseline for the informal metrics the pilot team was watching; an explicit freeze list distinguishing calibrated elements that should not be touched from adjustable elements that are safe to tune; and a named individual with real-time authority to correct a wrong output during any given shift.
Without the freeze list in particular, Dr. Tebaa notes, a new team often cannot tell which settings were carefully calibrated and which are arbitrary, and well-meaning adjustments can silently undo the fixes that made the pilot work in the first place.
A Test Before Calling It Done
Dr. Tebaa recommends a practical test for whether a handoff is genuinely complete: have the incoming operations team run the system alone for a full week while the original pilot team is truly unreachable. Any judgment call the operations team cannot resolve from the documentation alone becomes a gap to close immediately, rather than a surprise discovered weeks later in a falling accuracy chart.