Why Does Dr. Jonah Tebaa Say 'A Human Reviewed It' Needs Arithmetic?
Dr. Jonah Tebaa argues that claiming a human reviewed an AI output requires arithmetic because it is a testable statement of fact regarding time, visibility, and authority, not an unmeasured quality claim. Introducing his Review Claim Ladder framework, he uses an illustrative composite insurance broker to show how two reviewers handling 240 daily drafts spent a median of just 38 seconds per item without source documents. Through four audit tests—SAW, TIME, POWER, and TRACE—organisations must calculate whether review hours mathematically support the published claim.
A customer receives a reply with the wrong excess amount. The firm's website says every reply is reviewed by a human, so the customer asks the obvious question: did a person check this? According to Dr. Jonah Tebaa, an applied-AI strategist working across the MENA region, that question is where a large number of AI-assisted operations discover that their most reassuring sentence was never measured. He has published a framework, the Review Claim Ladder, to help organisations describe human review in terms they can defend.
A claim that can be tested
Dr. Tebaa's starting point is that "a human reviewed it" is not a statement of quality. It is a statement of fact about what a person saw, how long they had, and what they were allowed to do. Customers, auditors and regulators can test those facts, and in his work he finds that many firms have never tested them first.
His full argument, with a worked example, is in the original article on jonahtebaa.com. This piece summarises the reasoning for readers who want the structure quickly.
The arithmetic in his composite example
Dr. Tebaa illustrates the problem with a composite: a Gulf-based insurance broker that uses AI to draft replies to policy queries. He labels it as composite, with illustrative numbers, rather than a real client.
- About 1,200 replies a week, or 240 a day.
- Two reviewers, each with roughly six hours a day available for review. That is 720 minutes, so 3 minutes per reply.
- Timestamps show a median of 38 seconds from opening a reply to sending it.
- Eleven of the 1,200 replies were edited, under 1 percent.
- Reviewers see the draft, but not the policy document it cites.
When the customer's challenge arrives, the truthful answer is that a reviewer had the message open for roughly 38 seconds, with no policy document in view. Dr. Tebaa is careful to say this is not a story about negligent staff. In his account it is a story about volume and screen layout, and the remedy belongs in staffing and process, not in blame.
Four tests, applied to every item
Dr. Tebaa asks four questions, each answerable from records.
- SAW: did the reviewer see the whole output and the source it depends on?
- TIME: does the available review time, divided across the daily volume, leave enough minutes to read each item?
- POWER: could the reviewer reject or change the item without penalty or a second sign-off?
- TRACE: is there a record of who reviewed it, when, and what was changed?
Four levels of wording
The results map to four phrases, from the most demanding to the most limited.
- Approved by a named role. All four tests pass on every item, and that role signs each one.
- Checked. Each item is assessed against a written checklist, and the time spent is recorded.
- Skimmed. Every item is opened, but without a checklist. In his framing, this is a real process step, but it should never be described as review.
- Sampled. A stated share of items gets a full review, and the remainder is machine-checked on named fields.
Dr. Tebaa's rule is that an organisation should publish the level it actually passes. If it wants a stronger claim, it has to change headcount, volume or tooling. Rewording is the one lever he rules out.
What the broker should say instead
Applying the tests, the composite broker fails SAW and TIME, which places it at Skimmed at best. Dr. Tebaa proposes two changes. First, show the relevant policy extract next to each draft and verify amounts and names against the policy record before a person sees the reply. Second, replace the blanket sentence with the narrower wording he proposes: "One in five replies is fully checked by a named reviewer, and every reply's figures are verified against the policy record."
He then tests that claim against capacity. One in five of 240 replies means 48 full reviews a day. At six minutes each, that is 288 minutes of the 720 available, which leaves room for escalations. The new wording promises less, but each part of it can be demonstrated.
A short self-test for operations leaders
Dr. Tebaa suggests a short exercise for COOs, customer operations heads and compliance leads.
- Collect every place the review sentence appears: website, contracts, email footers, tenders.
- Extract last week's reviewer timestamps and find the median time per item.
- Divide review hours by volume and compare the result with the measured time.
- Check whether reviewers could see the source material, and whether rejecting an item carries any cost.
- Confirm you can show who reviewed one specific item and what changed.
- Restate the published claim at the level the evidence supports.
His closing test is practical. Could the firm give its reviewers' timestamps to the customer without hesitation? If the answer is no, the wording should be revised before a complaint forces the test.
Why it matters in regulated markets
Dr. Tebaa addresses this work to leaders at insurers, banks, logistics firms, telecoms and healthcare administrators, where a customer or supervisor can ask for evidence. His position is that a smaller claim that holds up is worth more than a larger one that depends on nobody asking. He also notes that the framework describes a process accurately and is not legal advice, and that any wording a regulator or contract prescribes takes precedence.
Related evidence: Article 14 of the EU AI Act requires high-risk AI systems to be designed and developed, including with appropriate human-machine interface tools, so that natural persons can effectively oversee them throughout the period they are in use. (EU AI Act Article 14 on human oversight)
The Model Cards paper proposes short documents that accompany trained machine learning models and report benchmarked evaluation across a variety of conditions. (the Model Cards for Model Reporting paper)