What does Dr. Jonah Tebaa on Why Most AI-Visibility Audits Are Wrong Before They Start mean in practice?
Most AI-visibility audits are wrong before they start because their prompt sets rely on generic keyword guesses rather than actual customer speech. Dr. Jonah Tebaa argues that a score is only as credible as this measurement instrument, demonstrated when an ungrounded sixty-one percent visibility score collapsed to twenty-two percent once audited against sales calls and support tickets. To establish a defensible baseline, Dr. Tebaa introduces a five-step method for building a defensible question set sourced directly from buyer language.
The Report Said 61 Percent. Nobody Could Say 61 Percent of What.
A CMO recently forwarded Dr. Jonah Tebaa a vendor report with a headline number: 61 percent "AI visibility." Her board wanted to know if that was good news. Dr. Tebaa asked a different question first: 61 percent of what, exactly, and answered by whom.
The vendor's audit had run twenty prompts against major answer engines and counted how often the brand appeared. Those twenty prompts were built from generic category keywords — the kind of terms an SEO tool suggests by default — not from anything a real buyer had actually said.
Working with the client's sales and support teams, Dr. Tebaa rebuilt the question set from ninety days of real call transcripts and support-ticket language. Same brand, same answer engines, same measurement window. The defensible baseline came back at 22 percent — a 39-point gap from the number the board had almost acted on.
His conclusion: the vendor hadn't made an arithmetic error. It had measured a market that didn't exist.
The Prompt Set Is the Instrument, Not the Score
Dr. Tebaa treats an AI-visibility score the way a researcher treats any survey result: it's only as credible as its sampling method. A poll of twenty people from one zip code isn't wrong because the math is bad — it's wrong because the sample was never the population it claims to represent.
In his view, the industry has quietly normalized exactly this error. A percentage looks rigorous regardless of what produced it, and the prompt list behind the number is rarely published or checked. Most vendor audits, he argues, never disclose their prompt set at all — which means clients have no way to know whether the sample matched the market it claimed to describe.
His position is that the prompt set is the actual instrument under test. The score is just a readout of that instrument. Get the instrument wrong, and every number downstream inherits the error.
A Five-Step Method for Building a Defensible Question Set
The method he uses with clients has five parts, and none of it requires specialized tooling — only discipline about where the questions originate.
- Source real buyer language, not guessed keywords — pulled from support transcripts, sales-call notes, site search, and search console data, not assumed by a content team.
- Segment by purchase stage, because awareness, comparison, objection, and post-purchase questions behave differently inside answer engines.
- Include competitor-comparison and "vs" phrasing explicitly, since category terms alone miss how buyers actually evaluate alternatives.
- Include reputational and failure-mode questions — "is X reliable," "is X worth it," "did X get acquired" — categories brands routinely skip despite their outsized weight in how answer engines characterize a company.
- Set a refresh cadence, quarterly at minimum, since buyer language shifts faster than SEO keyword lists ever did.
What the Real Baseline Changes
The honest number tends to come in lower, and Dr. Tebaa argues that's the entire point. A defensible baseline doesn't just correct a metric — it redirects budget.
In the case above, once the question set was rebuilt, close to 40 percent of real buyer queries turned out to be comparison phrasing ("X vs Y for a fifty-person team") and reputational phrasing ("is X still around") — categories the original audit had never asked about.
Two case-study pages already dominated the real question set, so further investment there would have been largely wasted. The actual gap was an entire comparison-page category that didn't exist yet. Once built, it moved coverage on that segment within a single content cycle. Dr. Tebaa is careful not to attach an overly precise after-number to that shift — he treats a suspiciously clean before-and-after as a sign the measurement is still broken somewhere — but the direction, he says, was unmistakable.
The Trap Is Treating the Baseline as Fixed
Buyer language keeps moving, which means a question set starts drifting the moment it's finished. Dr. Tebaa has seen sets older than two quarters where 15 to 20 percent of the questions no longer matched how a sales team said prospects were actually asking.
His argument is that a static baseline fails quietly — it keeps producing a stable-looking number while the market underneath it has already changed. He treats the question set the way any researcher treats an instrument measuring a moving target: revisited on a fixed schedule, not whenever someone happens to remember.
For Dr. Tebaa, the score was never the deliverable. The question set is — and getting that instrument right is what finally makes the score worth trusting enough to act on.