Why Does Dr. Jonah Tebaa Say Not Every AI Task Deserves the Best Model?
Dr. Jonah Tebaa argues that model choice must be a per-task decision rather than a habit, because default tiering over-serves routine work while leaving critical tasks under-protected. To match capability to risk, he introduces a three-question framework evaluating the cost of an error, whether output is checkable in under a minute, and monthly run volume. In an illustrative composite firm, routing routine summaries to fast models and contract comparisons to top models dropped monthly spend from roughly 1,900 dollars to 700 while improving accuracy.
A monthly AI bill can look reasonable and still be wrong in a way that matters. Dr. Jonah Tebaa, an AI strategist based in Lebanon, describes a composite case he uses when advising operations and finance leads: most of the spend goes to summarising meeting notes on the most expensive model available, while the single task that needed careful reasoning, a comparison of a supplier contract against standard terms, was run once on a quick, cheap setting. The money and the risk sit in different places, and nobody decided that they should.
His argument is that model choice is a per-task decision, and that most teams make it by habit. The full write-up is on his own site: Not Every AI Task Deserves Your Best Model.
Why the safe default is not safe
Dr. Tebaa meets two opposite habits in firms of twenty to five hundred people. One group defaults every task to the top model on the theory that the most capable option is the safest. The other puts everything on the cheapest option and is puzzled by uneven quality. In his view both skip the same step: looking at what each task actually demands.
For routine work, he points out, the ceiling of a fast model is rarely the limit. Summaries, ticket tagging and standard replies are cleared comfortably by a lighter tier, so the extra capability is paid for and never used. The real exposure is concentrated in a few jobs, such as long documents, multi-step reasoning and figures that feed decisions. Spreading budget evenly leaves those few under-protected while the many are over-served.
Three questions instead of a habit
Dr. Tebaa asks three questions about each recurring task, in a fixed order.
- What does a wrong answer cost? He grades it as low, real or serious.
- Can a competent person verify the output in under a minute? This is the checkability test.
- How often does the task run in a month? This measures how much a mistake in tiering will cost.
The answers produce a tier. A fast tier suits tasks with low cost of error that can be checked quickly. A standard tier suits real but checkable stakes, or cases where the team is unsure. The top tier is reserved for serious stakes or output that cannot be verified quickly, which in practice means multi-step reasoning, long documents and numbers that flow into pricing, legal or financial decisions.
Why checkability carries so much weight
The second question does the most work in Dr. Tebaa's rule. A cheaper model is an acceptable gamble only when the person receiving its output can catch a mistake at a glance. A summary of a meeting someone attended passes that test. A comparison of two forty-page contracts does not, because finding an error costs almost as much as doing the comparison. When errors are hard to spot, the quality of the first answer is the only protection available, and that is when the top tier earns its price.
Volume, by contrast, never promotes a task in his framework. A task that runs thousands of times at low stakes remains a fast-tier task. High volume simply makes a wrong choice more expensive, which is a reason to get the tiering right rather than a reason to buy more capability.
The worked example, in numbers
Dr. Tebaa illustrates the rule with a composite 60-person distribution firm, labelled as illustrative rather than a real client. Eight recurring tasks were sorted in one 45-minute session by the operations lead, the finance lead and a customer-service manager. Meeting summaries, ticket categorisation and routine order-status replies went to the fast tier. Catalogue rewrites, Arabic and English customer notices and weekly sales commentary went to standard. Contract clause comparison and supplier rebate reconciliation, about nine runs a month between them, went to the top tier.
In the illustration, monthly spend fell from roughly 1,900 dollars to roughly 700, a drop of about 63 percent. The more interesting result was on the quality side. The contract task had been the one running on the quick setting. Moving it up, and adding a step where every flagged clause is read against the source, caught a payment-terms difference that the earlier pass had missed. Cost fell because the many routine tasks came down, and quality rose because the few consequential ones moved up.
Running the sort
The exercise he recommends is short. A team lists its recurring AI tasks with rough monthly counts, notes the cost of a wrong answer for each, marks which outputs can be verified in under a minute, and assigns a tier. Where the team hesitates between two tiers, it picks the higher one for now. Each task gets a one-line upgrade trigger, for example moving up when a reviewer finds two material errors in a month. One person owns the list, and a review is booked for a month later.
Where teams go wrong
Dr. Tebaa names three recurring mistakes. The first is tiering by department rather than by task. Marketing does not live entirely on the fast tier, since it also writes pricing pages that inform decisions, and finance does not belong entirely on the top tier, since it also handles hundreds of routine invoices. The second is saving small amounts on a task whose result feeds a pricing or legal decision, where one wrong figure costs far more than the saving. The third is treating the sort as final. Models change, and a tier that fit last year may not fit now, so the list needs an owner and a date.
The management parallel
Dr. Tebaa closes on an analogy from ordinary management. A capable manager does not give every task the same hour from their most experienced person. Routine work gets a glance and a contract gets a careful read. His claim is that choosing an AI model tier is the same skill: matching effort to stakes, and keeping the best instrument for the few jobs where a mistake would hurt.