Why Dr. Jonah Tebaa Tells MENA Businesses to Stop Waiting for "Enough" Data?
Dr. Jonah Tebaa advises MENA businesses to stop waiting for enough data because unmet minimum thresholds reflect vendor technical defaults rather than true organizational readiness. Instead of requiring thousands of clean, labeled records for fine-tuning, existing messy assets like eight hundred multilingual WhatsApp threads or cash ledgers can immediately power retrieval-augmented generation and few-shot prompting. Dr. Tebaa introduces a three-question evaluation framework asking whether approaches require volume or retrieval, identifying the smallest real dataset to start narrow, and determining vendor default biases.
In his work advising businesses across Lebanon and the wider MENA region, Dr. Jonah Tebaa keeps running into the same document: an AI proposal that reads as though it were written for a company several sizes larger than the one holding it. The number that gives it away, he says, is usually buried in a single line — a minimum data threshold, often in the thousands, handed to a business whose entire operating history lives in a WhatsApp inbox.
He describes a composite scenario he uses to make the pattern concrete, not a single client file: a fourteen-month-old business with roughly 800 WhatsApp threads — unlabeled, mixing Arabic, Arabizi, and English, with no CRM behind any of it — is told it needs 5,000 labeled conversations before a support AI can be fine-tuned for it. Taken at face value, that gap looks like a verdict: the business isn't ready. Dr. Tebaa's argument is that it isn't a verdict at all. It's a mismatch between the business and the one technical approach the vendor happened to reach for.
The data was never missing
Dr. Tebaa's central point is that MENA SMBs almost always have real operating data — it simply doesn't look like a dataset, because nobody built it to. A WhatsApp Business line carries hundreds of genuine customer interactions. A cash ledger, however informal, tracks who buys what and how often. An owner's memory of which customers pay late or which products move before a holiday is, in his framing, a form of institutional knowledge that most enterprise-grade AI proposals never account for.
What matters, he argues, isn't whether that data is clean or structured. It's whether it's representative of how the business actually operates. Eight hundred messy, real threads beat five thousand clean examples pulled from a company that looks nothing like the one being pitched.
Two techniques, two very different appetites
The technical distinction Dr. Tebaa insists on is between fine-tuning and retrieval-based approaches. Fine-tuning adjusts a model's underlying weights, which by design requires substantial volume to work reliably — this is where vendor minimums in the thousands originate. Retrieval-augmented generation and few-shot prompting solve a different problem: they surface the most relevant real examples at the moment a model needs them, a task that hundreds of genuine examples can support well.
The split is a fact about what each technique consumes, not a matter of vendor taste. IBM's technical explainer puts it plainly: fine-tuning is a supervised learning method, which means the data used in training is organized and labeled. Eight hundred unlabelled, code-switching WhatsApp threads are precisely what that is not — which is how a five-thousand-conversation minimum ends up in a proposal written for a business that has never labelled a conversation in its life.
In his view, treating these as interchangeable — and applying a fine-tuning-scale threshold to a business that would be far better served by a retrieval-based approach — is where most "you're not AI-ready" verdicts actually originate.
A framework instead of a wait
Rather than telling business owners to accumulate more data before revisiting AI, Dr. Tebaa proposes three questions to ask before accepting, or shelving, a proposal built on an unmet data threshold:
- Does the proposed approach need volume, or does it need retrieval and examples — and which problem is the vendor actually solving?
- What is the smallest real dataset — WhatsApp threads, ledger entries, call notes — that would let the business start narrow and prove value before scaling?
- Is "not enough data" describing the business, or describing the one approach the vendor defaulted to?
For Dr. Tebaa, the last question does the most work. A data threshold is a design decision made by a vendor, not an objective measurement of a business's readiness. When the answer changes with the approach, he argues, the business was never the obstacle.
A different readiness gate for the region
The broader implication, in his framing, is for how Lebanese and MENA business owners should evaluate AI vendors going forward: not by asking how much data a company has, but by asking whether a proposed approach actually fits the data that already exists. He suggests owners ask vendors to show what they would build with the data on hand before asking what a larger threshold would make possible. A vendor who can only answer the second question, he argues, has told you more about their own default than about your business.