Why Dr. Jonah Tebaa Puts the Human Checkpoint After the AI Reply?
Dr. Jonah Tebaa puts the human checkpoint after the AI reply for low-risk messages to prevent costly response delays that destroy buyer intent. Testing revealed that pre-send human review added a fourteen-minute delay, dropping booked meeting conversions from thirty-one percent to twelve percent. Through his category-based framework evaluating whether an imprecise answer costs more than a delayed one, instant AI delivery is reserved for scheduling and qualification, while post-send human audits monitor drift.
Most debates about AI in customer-facing roles focus on whether a human should review a reply before it goes out. Dr. Jonah Tebaa's recent work on AI sales chat argues that this is the wrong question. The more consequential decision, in his analysis, is not whether a checkpoint exists at all, but where in the sequence it sits — and for which category of message. Move that single checkpoint from before a reply to after it, he found, and the effect on results is not marginal.
A Controlled Comparison of Two Identical Assistants
The pattern comes out of testing Dr. Tebaa ran on AI-driven lead qualification for B2B service and SaaS pipelines. He deployed two configurations of the same assistant — identical script, identical qualifying questions, identical tone and calendar-offer language — against comparable inbound volume, roughly one hundred leads a week over four-week windows. He is careful to frame this as an illustration of a pattern he has observed recur across several deployments, not a single definitive case study, and that hedge matters: the numbers below describe a repeatable dynamic, not a universal constant.
In the first configuration, the AI answered leads on its own. A question came in, the AI replied in roughly eleven seconds, worked through two or three qualifying questions, and offered an open calendar slot — no person in the loop before the message reached the lead. Review happened afterward, on a sample of transcripts, to catch drift or tone problems rather than to approve individual replies.
In the second configuration, the AI drafted the identical reply, but it sat in a queue until a person read and approved it. That single design choice — a pre-send human gate — added an average of just over fourteen minutes to every reply. Nothing about the substance of the message changed between the two setups. Only its position relative to a human's attention did.
What a Fourteen-Minute Delay Actually Costs
The results, according to Dr. Tebaa's data, were not close. The instant-reply configuration held a seventy-four percent contact rate and converted thirty-one percent of leads to a booked meeting. The reviewed configuration held a thirty-nine percent contact rate and converted twelve percent. Same script, same logic, same offer — and a gap of roughly two and a half times in booked meetings, produced entirely by where a person's judgment was inserted into the sequence.
That gap tracks a broader pattern he documented in response speed generally. Leads answered within the first minute book meetings at roughly thirty-four percent; once the wait stretches past five minutes, that figure falls to about nine percent. Buyer intent, in other words, behaves like it is on a short clock, and it keeps ticking whether the delay comes from a slow system or a careful reviewer. Leads have no way to know that a compliance-minded human is the reason an answer took fourteen minutes instead of eleven seconds; they experience only the wait, and by the time the approved reply arrives, several have already opened a competitor's tab.
The Real Decision Is Made Per Category, Not Per System
What makes Dr. Tebaa's argument more than a speed-versus-caution tradeoff is the framework he proposes for resolving it. He rejects the instinct to ask, for an AI reply, whether it is "good enough to send unsupervised." That question, he suggests, invites a blanket policy — review everything, or review nothing — when the honest answer changes depending on what is actually at stake in a given message. His reframing is narrower and more useful: for this category of message, does an imprecise answer cost more than a delayed one?
A pricing exception or a discount request has real downside if the AI gets it wrong — a wrong number can create a commitment the business has to honor. A standard qualifying question about company size or timeline has almost none; if the AI phrases it slightly off, the cost is trivial next to the cost of a lead going cold while waiting for approval. Sorting message categories by that question, rather than applying one rule to the whole funnel, is the core of the framework — and Dr. Tebaa lays out the fuller argument, including the full data set behind these numbers, in his own account of the testing.
Regulators sorting high-risk AI systems by the same logic have landed in roughly the same place. The EU AI Act's human oversight article requires that oversight measures be commensurate with the risks, level of autonomy and context of use of the high-risk AI system — a pricing exception and a scheduling question are not the same risk, and the rule that governs them should not be either.
Applied to a typical inbound sales chat funnel, the split looks roughly like this:
- Route through pre-send human review: any message that could create a binding commitment before a signed agreement exists, custom quotes and pricing exceptions, requests for a discount or an altered contract term, and complaints or messages carrying sensitive or emotional language.
- Let the AI send instantly, review after the fact: standard qualifying questions such as company size, timeline, or budget range, availability and scheduling offers, FAQ replies covered by an approved script, and initial acknowledgments confirming a message was received.
Separating a Bottleneck From a Design Choice
One distinction Dr. Tebaa draws that is easy to collapse into the rest of the argument: a slow-moving review queue is not the same problem as a badly placed checkpoint. If a company has enough reviewers and a fast approval process, it can still be routing categories through a human gate that never needed one — the queue moves quickly, but the delay it introduces on low-risk messages is still costing contact rate for no corresponding reduction in risk. Fixing the throughput of a review process and fixing where that process sits in the message flow are two separate projects, and solving the first does not solve the second.
The broader implication of his testing is that "responsible AI deployment" and "AI deployment with a checkpoint on everything" are not synonyms, even though they are often treated as interchangeable in rooms where stakeholders are nervous about a mistake reaching a prospect. A pre-send review on a pricing exception is a defensible safeguard. The same review applied to a scheduling offer is a tax on speed with no offsetting reduction in risk — and in a channel where buyer intent decays by the minute, that tax is measured in lost meetings, not just lost time.