What does The Exit Criteria Dr. Jonah Tebaa Writes Before Any AI Trial Begins mean in practice?
Before initiating any parallel artificial intelligence trial, Dr. Jonah Tebaa establishes five mandatory exit criteria to govern production launch decisions. These requirements demand a minimum match rate sustained across consecutive days, at least one full peak-volume or peak-complexity test day, a named individual review of every divergence case, a live rehearsal of the rollback path, and sign-off from a single accountable reviewer rather than group consensus.
Most conversations about AI failure focus on the model. Dr. Jonah Tebaa's attention goes somewhere less glamorous and, in his experience, far more decisive: the moment a working system stops being a trial and starts touching real customers. He argues that this handover is rarely governed by evidence. It is governed by a calendar, or by nothing at all.
His illustration is a shadow-mode deployment that reached day nineteen of a planned three-week parallel run. The system had been reading every case and generating a recommendation alongside the client's own team, changing nothing and visible to no one outside the project. On that day the operations lead asked to bring the launch forward by two days, because the coming Monday was short-staffed and going live would solve a real problem. Every qualitative signal in the room supported the request. Tebaa declined it.
The Missing Finish Line
The reason he could decline it, in his telling, is also the reason the pressure existed in the first place. The trial had a start date and an approximate duration. It did not have a written definition of "ready" agreed before any shadow data existed. Tebaa treats that absence as the structural fault, not the phone call.
Without a threshold committed to paper in advance, he observes, a team has only two available ways to end a parallel run: on the date printed in the kickoff deck, or never. He has watched both. Organisations sit in perpetual testing because nobody wrote down what passing looked like, so nobody can declare that the system passed. Others flip live on schedule because a date is concrete in a way an unwritten standard is not. Neither outcome, in his framing, is a decision. Both are defaults wearing the costume of one.
Why a Good Number Was Not Good Enough
The metric under observation was deliberately plain: on each day, how often the system's recommended action matched what the human team actually did. Tebaa's team logged it as a daily series rather than a closing average, on the reasoning that an average conceals both a bad stretch and a genuine drift.
By day nineteen the match rate was good but unstable. It had crossed the target on some days and fallen back on others. Tebaa makes a distinction here that he considers the practical heart of the matter: a figure checked once looks like success, while the same figure watched as a line looks unresolved. The standard set before the trial required the threshold to hold across consecutive days, and that run had not yet occurred.
He is candid that the pressure to override it was reasonable. The team enjoyed working alongside the system. No serious miss had been flagged in over a week. The staffing shortfall was immediate and real. That combination, he suggests, is precisely when a written number earns its cost — not when a system is visibly failing, but when everyone present, the consultant included, would prefer to say yes.
Waiting for a Condition, Not a Date
The two-day extension Tebaa requested was not about gathering more of the same evidence. It was about capturing a condition the scheduled window would have missed entirely: a genuine peak-volume day. Every day logged to that point had been an average day.
He regards this as the most commonly skipped step in the last mile. Systems that behave well under normal load are not reliably the same systems that behave well when volume spikes and the queue lengthens; the failure modes that emerge under pressure often differ in kind rather than degree. The peak day arrived on day twenty. The match rate held, and two edge cases surfaced that three weeks of ordinary traffic had never produced — neither disqualifying, both of which would have entered production unobserved had the original date been honoured.
The Criteria He Now Writes First
Tebaa's conclusion is that none of this called for a more sophisticated system. It called for a more complete definition of done, fixed before the trial rather than negotiated during it. The criteria he now commits to writing before any shadow-mode run begins, independent of client, sector, or tooling:
- A minimum sustained match threshold, held across a consecutive number of days rather than achieved on one strong day drawn from a noisy series.
- At least one full peak-volume or peak-complexity day inside the trial window, on the grounds that average days conceal the gaps that matter.
- Every divergence case reviewed by a named individual, with none deferred indefinitely.
- A rollback path rehearsed in practice at least once, not merely documented in a plan no one has executed.
- A single named reviewer accountable for signing off on the number, rather than a group consensus that quietly disperses ownership.
Individually these are modest. Together, in his argument, they convert "does this feel ready" into a question with a checkable answer — the only kind of question capable of surviving a call from someone with a legitimate and unrelated problem to solve.
Tebaa is careful not to claim novelty for shadow-mode testing itself. Running a system in parallel before exposing it to customers is established practice, and he notes that most teams he works with already do some version of it. What they omit is the harder, earlier act: deciding in advance what evidence would justify saying no to their own launch date, while that decision is still cheap and abstract rather than expensive and personal.
Writing the list on day one, as he puts it, costs considerably less than assembling it under pressure on day nineteen.