What does Dr. Jonah Tebaa: The AI Budget That Gets More Expensive When It Works mean in practice?
In a composite case, Dr. Jonah Tebaa shows that usage-priced AI budgets escalate when rollouts succeed because real customer queries trigger longer reasoning chains and context windows, driving costs up from $0.16 to $0.2275 per query. To prevent a projected $64,000 bill from reaching $91,000, Dr. Jonah Tebaa applies a four-question board tool that tracks true per-unit costs across volume scenarios, establishes distinct per-unit cost ceilings, and classifies commitments properly.
Dr. Jonah Tebaa spends a lot of time in rooms where a board has just approved an AI system, and a lot less time in the rooms where that same board finds out what it actually costs to run. In his latest analysis, he walks through a composite scenario built from a pattern he says recurs across board approvals for usage-priced AI: a pilot that ran at 40,000 queries a month for roughly $6,400, a board that reasonably projected a 10x rollout to land near $64,000, and an actual invoice at 400,000 queries a month that came in near $91,000. He is careful to label the figures composite and illustrative rather than a single client's raw invoice, but he argues the arithmetic is exactly what he sees recur across real approvals.
The point of the piece is not that the project overspent. Dr. Jonah Tebaa's argument is closer to the opposite: nothing went wrong. Adoption succeeded, usage grew exactly as intended, and the bill grew faster than the board's own math predicted anyway. He frames this as the central governance blind spot in how boards currently approve usage-priced AI systems — a blind spot that has nothing to do with vendor behavior or model quality, and everything to do with which variable gets modeled at approval time.
Two Pricing Models, Two Different Ceilings
Dr. Jonah Tebaa's mechanism starts with a comparison most finance committees already understand intuitively but rarely apply to AI line items. Seat-priced software has a built-in ceiling: cost is a fixed number multiplied by headcount, and headcount grows in increments a board can see coming a quarter in advance. Usage-priced AI has no equivalent ceiling, because its cost depends on two variables rather than one — how much volume moves through the system, and how expensive each individual interaction is to serve. In his account, nearly every approval memo he reviews models the first variable in detail and treats the second as a constant, carried forward from the pilot without ever being written down as an assumption that could move.
Why Success Is the Trigger, Not the Failure
The second half of the mechanism explains why that constant moves exactly when a rollout works. Dr. Jonah Tebaa points to how pilot data gets built in the first place: a smaller, more engaged group of early users, staff still reviewing outputs before they reach customers, and requests that skew simpler because early adopters are still learning what the system can do. Production usage, once a rollout succeeds, looks nothing like that sample — real customers ask multi-part, ambiguous questions that trigger longer reasoning chains, more retries, and larger context windows than the pilot ever exercised. He walks through the arithmetic directly: the pilot's $6,400 on 40,000 queries works out to $0.16 per query; a truly linear 10x projection holds that same $0.16 per query at $64,000 on 400,000 queries; the actual $91,000 on 400,000 queries works out to roughly $0.2275 per query, about 42 percent higher than the pilot rate. In his framing, that 42 percent is the part no approval memo priced in, because no approval memo asked the question that would have surfaced it.
The Four-Question Test He Uses Before Sign-Off
The practical center of the piece is a four-question board tool Dr. Jonah Tebaa says he now runs before any usage-priced AI system gets approved:
- What is the true cost per unit of work at pilot volume — per resolved ticket, per generated report, per completed transaction — rather than the total monthly bill?
- Modeled across 2x, 10x, and 50x volume scenarios, is per-unit cost assumed flat, falling, or rising in each case, and who is accountable if that assumption is wrong?
- Is there a per-unit cost ceiling set as its own governance trigger, distinct from and in addition to the total monthly budget ceiling — one that forces a review even while the total budget line still has room?
- Is the commitment correctly classified as capacity-like (bounded, seat-priced) or consumption-like (unbounded, usage-priced), with budget cycles and board reporting matched to whichever it actually is?
Dr. Jonah Tebaa's point in laying the questions out this way is speed as much as rigor — he argues the exercise takes roughly fifteen minutes inside an approval meeting, against the composite scenario's three billing cycles of unnoticed cost creep before anyone caught it.
Rewriting What the Approval Memo Is Actually For
Dr. Jonah Tebaa is explicit that his argument is not a case against usage-priced AI. It is a case for splitting two questions that most memos currently collapse into one: whether the system will work, which most approval processes already test reasonably well, and what happens financially once it works better than the pilot suggested, which he says almost no memo addresses at all. He connects the idea to two adjacent pieces of his own thinking — his earlier argument about how boards misallocate AI budgets they have already approved, and a separate framework he uses for ranking competing AI investments by what it would cost to unwind each one rather than only by projected return. Taken together, his position is that a usage-priced system without a per-unit ceiling functions as an investment with an unpriced, unbounded downside case.
His closing argument reframes the composite scenario's $91,000 bill: not a cost-control failure, in his telling, but a pricing failure at the design stage, months before a single production query ran. Most approval memos are built to catch overruns caused by failure. Dr. Jonah Tebaa's contention is that the overruns boards should be planning for are the ones caused by success — because a board that has priced what happens if the system works even better than expected is the board that does not get blindsided by its own win.
Related evidence: The EU AI Act obliges providers of high-risk AI systems to report a serious incident to the market surveillance authorities immediately after establishing a causal link to the system, and in any event not later than 15 days after becoming aware of it — a disclosure deadline fixed in law rather than decided during the incident. (the EU AI Act's 15-day serious-incident reporting deadline)
The UK government's introduction to AI assurance defines AI governance as a range of mechanisms, including laws, regulations, policies, institutions and norms, used to outline processes for making decisions about AI. (the UK government's introduction to AI assurance)