What does Dr. Jonah Tebaa on the AI Quality Drop That Was a Queue Problem mean in practice?
Dr. Jonah Tebaa explains that an AI quality drop is often a queue problem caused by resource contention rather than model drift. When an RFQ qualification override rate spiked from 8 percent to 34 percent, the issue stemmed from unmanaged shared queues handling follow-up emails, RFQs, and pipeline reports simultaneously. The fix established weekly ceilings on each job overseen by named roles, successfully restoring the override rate to roughly 9 percent without modifying the underlying model.
An override rate of 8 percent climbed to 34 percent over six weeks, and nobody on the team had changed the workflow it measured, the rules behind it, or the AI system running it. That gap — a quality number moving while everything local to it stayed still — is the starting point Dr. Jonah Tebaa uses to argue that most quality complaints about an AI system get diagnosed backwards.
His argument is that when the instinct is to interrogate the exact spot where a number moved, it is usually the wrong instinct. If nothing local to a failing job has changed, in his framing, the cause is rarely local either — and the case he uses to demonstrate it involves a sales team, a shared queue, and a symptom that pointed everyone toward the wrong department.
The Override Rate Nobody Could Explain
The team in question is an 11-person sales operation at a regional distributor, running one AI system across a single open mandate on its sales side. Over time, three separate jobs had come to run through it: drafting roughly 40 follow-up emails a week to quoted customers, qualifying about 25 inbound RFQs a week against pricing rules, and assembling the VP's pipeline report once a week.
The number the sales director watched was his own override rate on RFQ qualification — how often he corrected the system's call before it went out. It had held near 8 percent for months, then climbed to 34 percent over six weeks. By Dr. Tebaa's account, everyone closest to the RFQ side had already checked themselves out as the cause: the pricing rules were unchanged, the RFQ format was unchanged, and the director was applying the same standard he always had. Only the number had moved.
The Explanations Dr. Tebaa Ruled Out
Dr. Tebaa walks through three explanations before naming the real one, and treats all three as worth stating precisely because ruling them out is what clears the path forward. A drifted model was the first to go — no update had shipped to the system in that window. Harder RFQs went next — deal size, complexity, and client mix were flat against the prior quarter. A stricter reviewer was the explanation that held longest, until the overrides were plotted against the calendar: a reviewer tightening his own standard produces overrides spread evenly across his working days, and these were not spread evenly. They were bunched on specific days, and that clustering is what redirected the diagnosis entirely.
The Queue Behind the Number
The days RFQ overrides spiked, in Dr. Tebaa's account, were the same days follow-up email volume spiked. Both jobs ran through the same AI system with no weekly limit on either, meaning they shared one queue with no rule for which job took priority when both arrived heavy. Whichever job landed second on a busy day absorbed whatever attention was left. The 2.1-day delay that had crept into follow-up emails that same quarter was not a separate issue — in his telling, it was the identical fault surfacing in the other direction, on the days RFQ work happened to land first instead.
Once the cause was identified as contention rather than competence, the fix he applied touched scheduling, not the model: a weekly ceiling on each job, and one named person watching each ceiling — a Follow-Up Drafter capped at 40 a week under the sales ops manager, an RFQ Qualifier capped at 25 a week under the sales director, and a Pipeline Reporter producing one report a week under the VP. Within three weeks, the override rate was back to roughly 9 percent and follow-up emails were again going out same-day, with no change made to the model, the rules, or the reviewers in between.
Dr. Tebaa's broader point is that a quality drop inside an AI-augmented team is frequently a resourcing collision dressed up as a model failure, and the signature is consistent: the number that falls belongs to the job that was never touched. His recommendation, before anyone retrains or replaces a system, is to check what else was quietly added to its queue first.