What does Dr. Jonah Tebaa's Five-Question Test for AI Implementations That Actually Hold mean in practice?
Dr. Jonah Tebaa's five-question test evaluates AI implementations before approval across five operational checks: identifying a named individual responsible for unresolved exception queues, verifying direct visibility into that workload, establishing real shutdown timelines and authority, determining whether staff adapt to or route around the system, and rehearsing the response to the first consequential mistake. Rather than abstract maturity models, this diagnostic requires staffing, instrumentation, and rollback plans to prevent post-launch operational failures.
Most write-ups of AI failure focus on the wrong moment. They dwell on the launch — the go-live date, the ribbon-cutting, the demo that either dazzled the room or fell flat. Dr. Jonah Tebaa's argument, developed across his recent operator-facing writing, is that launch day tells you almost nothing. The moment that actually determines whether an implementation survives arrives a few weeks later, quietly, and it rarely gets a name.
Tebaa's framework starts from a simple observation: every AI system operating on real volume will produce a stream of cases it cannot resolve on its own. That is not a defect to be engineered away — it is a structural feature of deploying automation into a messy world. What separates the implementations that hold from the ones that quietly decay, in his account, is not the sophistication of the model. It is whether anyone was ever assigned to handle what the model could not.
He frames this as a pre-approval problem rather than a post-launch one. Rather than treating the weeks after go-live as a period to monitor and react to, Tebaa argues the relevant decisions have to be made and staffed before the build even starts — because by the time the exception queue is visibly piling up, the organization is already reacting from a deficit, and the fixes available at that point are slower and more expensive than the ones available at the approval stage.
Five questions, not a framework
What distinguishes Tebaa's approach from more theoretical treatments of AI risk is its refusal to stay abstract. He does not propose a maturity model or a scoring rubric. He proposes five plain-language questions that an operator can ask in an approval meeting, each pointing at a specific, checkable fact rather than a sentiment.
The first two questions concern capacity and visibility: is there a named individual — not a team, not the vendor — responsible for clearing the cases the system cannot handle, and can that person see the volume of that work directly, without having to request it from someone else. The same instinct is already codified in public-sector rules: Canada's Directive on Automated Decision-Making obliges departments to document "human overrides of the decision or assessment made by the system and other system failures", and to use those findings to take corrective action — a requirement that only functions when a specific person owns that queue and can see it. Tebaa treats these as staffing and instrumentation decisions, not commitments to be taken on faith. A responsibility that lives only in a kickoff slide, he suggests, tends not to survive contact with a busy week.
The second pair addresses the system's relationship to the organization around it: how long a genuine shutdown would actually take once dependencies have formed, and whether the team using the system day to day is adapting its workflow to it or quietly routing around it. This second point is, in Tebaa's telling, the harder one to see from above — workarounds rarely announce themselves, and the only reliable way to catch them, he argues, is to ask the team directly rather than infer their behavior from adoption metrics.
Why the fifth question is the load-bearing one
The final question in Tebaa's test concerns the system's first visible, consequential mistake — not whether one will happen, which he treats as close to inevitable given sufficient real-world volume, but whether the organization's response to it has been rehearsed in advance or will be improvised under pressure. That expectation is not idiosyncratic: the UK National Cyber Security Centre's guidelines for secure AI system development state that "The inevitability of security incidents affecting your AI systems is reflected in your incident response, escalation and remediation plans", and expect responders to have been trained for those incidents before one occurs. This is where, in his analysis, most implementations that fail actually fail: not through a design flaw, but through a chaotic, ad hoc response to an entirely foreseeable event.
What makes this diagnostic distinct from typical AI adoption advice is its refusal to moralize. Tebaa is not arguing for caution over ambition, or for slowing down adoption in the name of risk management. His five questions are operational, not philosophical — they concern staffing lines, dashboards, rollback timelines, and a rehearsed communication plan, all of which can be answered in a single approval conversation rather than deferred to a future audit.
The practical form of the test is deliberately compact. As Tebaa presents it, the five questions can be copied directly into an approval checklist: who owns the unresolved cases, can that ownership see its own workload, how fast can the system actually be reversed and by whose authority, is the team adapting or working around the system, and has its first real mistake been rehearsed. Asked at the point of approval rather than after launch, they cost a few minutes of a meeting. Left unasked, they tend to resurface a few weeks later as a queue nobody is clearing and a room with no plan for the moment that follows.