Cost overruns on AI pilots rarely trace back to a technology failure. In Dr. Jonah Tebaa's assessment, they trace back to a spreadsheet that was incomplete from the start. His argument, built around a specific worked example, is that vendor pricing pages describe only one part of what it actually costs to run an AI tool from Lebanon or much of the wider MENA region, and the missing parts are precisely what surface at renewal.
Tebaa is careful to frame the numbers behind his framework as illustrative. The scenario he uses, a 10-person Beirut SaaS company running a customer-support copilot, is a constructed example built to make a pricing structure concrete, not a case study drawn from a real client engagement or a market survey. The figures are meant to show a pattern, not report a measurement.
A Quote That Rarely Survives Contact With the Invoice
The pattern Tebaa describes starts with a familiar sequence: a company runs a small AI pilot, gets it approved against the vendor's quoted price, and then finds the first production invoice running well above that number, often by 50 to 70 percent. His point is that this gap is not a sign the vendor overcharged or the tool underperformed. It's a sign that three cost categories specific to this market were never included in the original budget line.
The Four-Line Cost Stack
Tebaa names this structure the four-line cost stack, and applies it to a specific scenario: a Beirut support desk handling roughly 2,000 tickets a month across WhatsApp and web chat, most conversations mixing Arabic, English and Franco-Arabic.
The first line is the List Price itself, the number on the vendor's pricing page, set in his example at $380 a month for an estimated 19 million tokens. The second is what he calls the Payment Access line: because a Lebanese business account often cannot settle a recurring USD subscription directly, payment typically routes through a UAE-registered card or an intermediary billing service, adding forex spread and service fees, roughly 11 percent, or $42, in his scenario. The third is the Usage Pattern line: vendor benchmarks are usually built on English-only demo traffic, while real bilingual and code-switched conversations consume meaningfully more tokens, adding roughly 38 percent, or $144. The fourth is the Continuity Reserve, not a charge but a budgeting discipline, set aside because a subscription tied to a single card or a single person's account represents a single point of failure. Tebaa sizes that reserve at roughly 17 percent, or $66.
The Usage Pattern line is the one founders push back on hardest, and it is also the line with the clearest outside support. A NeurIPS 2023 study of language-model tokenizers found that the same text translated into different languages can have drastically different tokenization lengths, with differences up to 15 times in some cases, a disparity its authors tie directly to the cost of accessing commercial language services. Measured against a spread like that, a 38 percent premium for Arabic-English code-switching reads as a conservative assumption rather than a padded one.
Stacked together, the four lines bring the realistic monthly run-rate in his example to roughly $632, against a quoted $380, a gap of about 66 percent. His framing is direct: nothing about the tool changed. Three cost lines were simply priced at zero until the invoice priced them instead.
What Tebaa Recommends Before the Next Renewal
Tebaa's guidance for founders and finance teams evaluating a production commitment centers on pricing these lines in advance rather than discovering them after the fact:
- Confirm Payment Access with finance before a pilot is approved, not after the first invoice.
- Test pilots against real bilingual or mixed-language traffic samples rather than a vendor's English-only demo.
- Request MENA or emerging-market usage benchmarks directly from vendors, rather than relying on their global averages.
- Build a 15 to 20 percent continuity reserve into any AI budget line tied to a single card or account.
- Re-price the tool against 60 to 90 days of real usage data before signing an annual contract.
His broader conclusion is that the risk in most failed AI pilots was never the model's performance. It was a budget line that assumed a vendor's global pricing page was the whole story, when in this market it was only the first of four.