What does The Audition Protocol: Five Tests a Practitioner Runs Before Adopting an AI Tool mean in practice?
Before adopting an AI tool, a practitioner evaluates it using Dr. Jonah Tebaa's five-part framework called the audition protocol. This process consists of the Boring-Task Test to time repeated five-minute workflows and correction overhead, the Failure-Signature Test to distinguish loud flags from confident silent errors, the Edit-Distance Test to measure unedited output surviving into final deliverables, the Version-Drift Test to monitor silent software updates against baselines, and the Handback Test to avoid dependency.
Most professionals evaluate a new AI tool the same way: they hand it the hardest problem on their desk and see if it survives. Dr. Jonah Tebaa argues this is exactly backwards. In his view, a hard task is too forgiving an environment — almost anything looks impressive next to a blank page, and a single dramatic success tells an operator nothing about the hundreds of routine tasks the tool will actually be asked to do. The demo, whether staged by a vendor or self-administered on a showcase problem, is built on ground chosen for the tool's benefit. A real workflow is not that ground.
His alternative is a five-part evaluation he calls the audition — a deliberate attempt to see a tool fail before trusting it to succeed. Where a demo shows what a system can do under ideal conditions, Dr. Jonah Tebaa's framework is designed to surface what a tool will actually cost an operator once it is inside a messy, exception-filled week of real work.
Why the Demo Fails as an Evaluation Method
In his work advising operators on tool adoption, Dr. Jonah Tebaa draws a sharp distinction between a demo environment and a live workflow. A demo is staged inside the tool's designed use case: clean inputs, the precise task it was built to handle, run by someone who already knows its blind spots. A live workflow, by contrast, is shaped by years of accumulated exceptions, inconsistent formatting, half-finished context, and constraints that never made it into any product specification. A tool that performs flawlessly in a vendor's controlled environment, he notes, has told an operator nothing about how it behaves once those controls are removed.
This is why Dr. Jonah Tebaa discourages testing a new tool on the hardest task available, a common instinct he considers intuitive but misleading. A hard task carries so much natural variance that nearly any output looks strong by comparison. It also happens rarely — once a quarter, perhaps — while the tool's real cost or value shows up in the routine work it touches daily. His audition framework reorders the test entirely, starting with the boring and working toward the existential.
The Five Tests
Dr. Jonah Tebaa's audition protocol consists of five sequential tests, each designed to expose a different failure mode a demo is built to hide.
- The Boring-Task Test. He recommends giving a new tool the single most repeated five-minute task in a person's workflow, ten times consecutively, while timing the entire loop — including the human review and correction step that follows the tool's output. In his framing, this check-and-correct time is where a tool's real cost lives, and it is precisely what vendor demos never show.
- The Failure-Signature Test. Here the operator hands the tool a task just outside its stated scope and observes how it fails. Dr. Jonah Tebaa treats the distinction between loud and silent failure as more consequential than raw accuracy. A tool that flags uncertainty when it is out of its depth is workable; a tool that produces a fluent, confident, wrong answer is dangerous regardless of how rarely that happens, because the moment it occurs is precisely when a user is least likely to catch it.
- The Edit-Distance Test. Rather than relying on subjective satisfaction, Dr. Jonah Tebaa measures how much of a tool's first output survives, unedited, into the version a person actually ships. A tool that consistently requires heavy rewriting has not saved craft time, in his view, even when it produces something that "feels" helpful — the labor has simply moved from writing to editing.
- The Version-Drift Test. Because many tools update silently and on schedules users don't control, Dr. Jonah Tebaa argues operators need a built-in way to detect when a tool's behavior has quietly changed. Without a periodic rerun of the boring-task and edit-distance checks against a fixed baseline, drift goes unnoticed until it has already degraded a workflow.
- The Handback Test. The test Dr. Jonah Tebaa says most people skip is also the one he considers most revealing: can the operator still perform the task themselves, cold, if the tool disappeared tomorrow? If the honest answer is no, he argues, the tool hasn't created leverage — it has created a dependency the operator never explicitly agreed to.
Leverage, Not Dependency
Underlying all five tests is a distinction Dr. Jonah Tebaa returns to often in his broader writing on AI adoption: the difference between a tool that extends a person's capability and one that quietly replaces it. A tool an operator can step around when necessary is leverage. A tool that has become structurally load-bearing, one whose absence would leave a skill unrecoverable, is something else entirely — and in his framing, most adoption decisions never test for that difference because the vendor demo was never designed to reveal it.
This isn't presented as an argument against adopting AI tools quickly. Dr. Jonah Tebaa is explicit that some tools earn a permanent place in a workflow within days of a proper audition. His point is narrower and, he suggests, more useful: evaluation should happen on the operator's terms rather than the vendor's — on routine work rather than showcase tasks, with failure deliberately induced rather than avoided, and with an honest accounting of what survives untouched rather than a general impression of helpfulness. The demo answers what a tool can do. The audition, in his framing, answers the only question that actually determines whether adoption was worth it: what will this tool cost, and what will it quietly replace, once it's living inside the real work.