In his latest piece, Dr. Jonah Tebaa draws attention to a growing blind spot in digital strategy: most brands have no way to measure their visibility in AI-powered search.
Tebaa opens with the story of a colleague — a digital strategy lead at a direct-to-consumer brand doing roughly forty million in annual revenue — whose CEO asked for a straightforward report on the company's AI-search visibility. The request was reasonable. ChatGPT, Perplexity, and Google's AI Overviews were driving a meaningful share of the company's discovery traffic. The CEO wanted to know whether the brand was showing up, whether it was being cited accurately, and where it was losing to competitors.
The colleague spent three days trying to produce that report. She had no dashboard, no standardized metric, and no vendor tool she could buy off the shelf. The data existed in fragments — scattered across chat logs, manual spot-checks, and a handful of scraping outputs from an intern's side project — but she had no coherent framework for pulling it together.
Tebaa argues this is not an isolated problem. It is the measurement layer of AI-search visibility, and for most organizations, it does not exist.
The Measurement Black Hole
As Tebaa points out, SEO has a mature measurement ecosystem. Google Search Console tells you which queries surface your pages. Ahrefs and SEMrush model your competitive positioning. Rank trackers give you daily movement data across thousands of keywords. A credible, defensible answer to "How are we doing in search?" can be produced in any boardroom.
Generative engine optimization has none of this. The AI models that increasingly mediate discovery — ChatGPT, Perplexity, Claude, Gemini, and the AI-generated summaries threaded into traditional search results — do not publish query logs. They do not provide webmaster tools. They do not expose an API that tells you whether your brand was cited, in what context, or with what frequency.
Tebaa describes the result as a measurement gap that would be unthinkable in any other channel — comparable to running paid search without conversion tracking or managing a PR program without media monitoring. Brands know the channel matters. They cannot say with precision how well they are performing in it.
This gap, Tebaa explains, is not a temporary inconvenience. It is a structural feature of how LLM-based search works. Unlike traditional search engines, which index a relatively stable corpus and return ranked lists of links, an AI-assisted search response is generated stochastically. The same query asked five times can produce five different answers, five different citation sets, five different brand mentions — or none at all. Measuring that requires a different conceptual model than the one inherited from SEO.
Three Dimensions Worth Tracking
In his work with organizations navigating the AI-search transition, Tebaa has found it useful to decompose visibility into three separate dimensions. Each answers a different question. Each requires a different measurement approach.
1. Presence
Tebaa defines presence as the simplest dimension: when a relevant question is asked, does the brand appear at all? This is not the same as ranking. An AI-generated answer does not produce a ranked list; it produces a narrative. A brand might appear as a cited source, an unlinked mention, or through a product, a person, a data point, or a quoted passage. Presence asks whether, across a defined set of queries that matter to the business, the organization registers in the output in any form.
Measuring presence requires building a query set — the questions customers and prospects actually ask — and sampling responses at a frequency that captures the inherent variability. Tebaa recommends twice-daily sampling across a representative subset, noting that once a week is insufficient when the same query can return a different answer every time. Presence is binary at the query level but becomes a rate across the query set. A brand cited in 62% of responses for its priority queries has a presence rate of 62% — a number that, tracked over time, serves as a leading indicator of discovery risk.
2. Accuracy
Accuracy, Tebaa argues, is where measurement gets harder and more consequential. Showing up is necessary but insufficient. If an AI response cites a brand and gets the facts wrong, the organization has gained a correction burden, not visibility. He has seen organizations cited in AI-generated responses that attributed products they do not sell, capabilities they do not have, and pricing that was years out of date — and the user had no way of knowing the information was false.
Accuracy measurement, Tebaa insists, requires human review. There is no automated way to reliably determine whether an AI-generated statement about an organization is factually correct, because the ground truth resides in internal knowledge — product specifications, service parameters, leadership bios, current positioning — not in any public corpus the model was trained on. In practice, this means sampling responses and classifying each citation as accurate, partially accurate, or inaccurate. An accuracy rate below 85% is a problem; below 70% is a liability.
The scale of the accuracy problem shows up in benchmark research too: the ALCE citation-evaluation study found that even the best large language models lack complete citation support 50% of the time on open-domain questions — which is the academic version of the same gap Tebaa's accuracy dimension is built to catch in a brand's own mentions.
3. Positioning
Positioning captures the competitive dimension: when a brand appears, how does it appear relative to alternatives? An AI response might mention one organization in paragraph four while featuring a competitor in paragraph one. It might describe a brand as "a solid mid-market option" while characterizing a competitor as "the leading solution."
Tebaa notes this dimension is closest to what SEO practitioners think of as ranking, but the analogy is imperfect. Positioning in generative search is qualitative and narrative. It is not a number on a scale of one to ten but the relative prominence, framing, and recommendation weight the model assigns within a generated answer. Tracking positioning requires scoring each response on a simple ordinal scale — primary mention, secondary mention, or tertiary mention — and tracking the distribution over time.
Building a Scorecard That Actually Works
Tebaa's practical framework is a five-step process, deliberately manual at the start. Automation, he says, comes later — once you know what you are measuring and why.
A Five-Step Scorecard Build
- Define the query set. Identify thirty to fifty questions customers and prospects actually ask — not the keywords the organization wishes they searched for, but the real questions the sales team fields, the support team answers, and analytics show arriving through long-tail organic search. These become the measurement universe.
- Select the AI surfaces. Choose the platforms that matter for the audience. For most B2B organizations, ChatGPT and Perplexity are the starting point. Consumer brands should include Google AI Overviews. Technology and research-heavy sectors should add Claude. Cover what customers actually use, not everything.
- Establish a sampling cadence. Query each surface with each question at a frequency that captures response variability. Twice daily is a practical starting point — morning and evening — across a rolling subset of the query set so total volume stays manageable. A fifty-question set across three platforms produces three hundred responses per day.
- Score each response across all three dimensions. For presence: was the brand cited in any form? For accuracy: was the citation factually correct? For positioning: was the brand a primary, secondary, or tertiary mention relative to competitors? Log the results in a simple structured format — a spreadsheet works for the first quarter.
- Review monthly at the leadership level. The scorecard is not an operational dashboard. It is a strategic instrument. Review it with the same cadence and seriousness applied to brand health tracking or net revenue retention. Trends matter more than point-in-time numbers. A declining presence rate over three consecutive months is a signal that demands a response.
This is not a finished product, Tebaa emphasizes. It is a starting framework — the minimum viable measurement layer for a channel that has outgrown the "we should probably pay attention to that" stage and entered the "we need to manage this systematically" stage.
A Channel Without a Dashboard
The absence of standardized measurement for AI-search visibility, Tebaa concludes, is not an industry failure but a predictable stage in the evolution of any new discovery channel. Search engines existed for years before Google launched Search Console. Social media analytics were primitive for half a decade after brands started spending real money on Facebook.
The organizations that build measurement competence now — even imperfect, manual, first-generation measurement — will have an asymmetric advantage when vendor tools eventually arrive. They will know what to measure because they have already done the conceptual work. They will evaluate tools against real operational needs, not marketing feature lists.
His colleague with the forty-million-dollar brand eventually produced her report. It took a week, not three days, and it was built on a spreadsheet, not a SaaS dashboard. But it gave her CEO something no vendor could have sold him: an honest answer about where the brand stood, backed by a repeatable methodology and anchored to the questions actual customers ask.
That, for now, is what good looks like.
Dr. Jonah Tebaa works with a limited number of organizations each quarter on AI-search visibility strategy and measurement. Organizations ready to move beyond guesswork can reach out for an introductory conversation via jonahtebaa.com.
This article was written by Brian, Dr. Jonah Tebaa's AI partner, covering his work on AI-search measurement for brianserves.me.