AI visibility: how to measure and improve it
Ranking well no longer guarantees being cited. This page covers how AI visibility is measured, what the measurement cannot tell you, and what the evidence says actually moves it.
AI visibility is how often and how prominently a brand, domain or page is mentioned and cited by AI answer engines such as ChatGPT, Perplexity, Google AI Overviews, Gemini and Claude when people ask questions in your category. It is measured by running a fixed set of prompts against each engine on a schedule and recording who gets named, cited and recommended.
Optimization for AI
Why is AI visibility a different metric from search rankings?
Because ranking well no longer guarantees being cited. Ahrefs analysed 863,000 SERPs and 4 million AI Overview URLs in a March 2026 study and found that only 37.9% of URLs cited in AI Overviews also ranked in the organic top 10 for the same query. Around six in ten cited sources came from somewhere other than the visible top ten.
The mechanism behind that gap is query fan-out. Google states that AI Mode "uses our query fan-out technique, breaking down your question into subtopics and issuing a multitude of queries simultaneously on your behalf" (Google, May 2025). You are no longer competing for one keyword. You are competing across a spread of sub-questions the user never typed and never sees.
The second difference is what a win looks like. A ranking win is a click. An AI win is often a mention with no click behind it. Pew Research Center tracked 68,879 Google searches from 900 US adults and found that when an AI summary appeared, people clicked a traditional search result in 8% of visits, against 15% of visits without a summary. Clicks on a link inside the summary happened in 1% of visits.
| Search ranking | AI visibility | |
|---|---|---|
| Unit of competition | One query, one URL | A cluster of fan-out sub-questions |
| What you win | A position and a click | A mention, a citation, a recommendation |
| Where it is visible | Rank trackers, Search Console | Only by querying the engines directly |
| Failure mode | Page 2 | Named competitor, no mention of you |
Traffic is not the only payoff either. Adobe data reported in June 2026 found AI referrals to US travel sites grew 194% year over year in May 2026, and those visitors spent 70% longer per visit with bounce rates 41% lower than non-AI traffic, although they still converted 28% less. That figure is travel-specific, so treat it as a signal about visitor behaviour rather than a number for your own vertical.
For the underlying mechanics, see how AI visibility differs from classic SEO.
Why can’t Search Console or analytics tell you this?
Google added a generative AI performance report to Search Console in 2026. It covers impressions from AI Overviews and AI Mode, grouped by page, country, date and device. It is genuinely useful, and it is also bounded: impressions only, and only Google’s own surfaces.
Nothing in that report sees ChatGPT, Perplexity or Claude. Analytics sees only the people who arrive, and most AI answers end without a visit. If you want to know whether an engine recommends you, you have to ask the engine.
How is AI visibility actually measured?
By sampling, not by counting. A tracker defines a set of prompts a real buyer might type, sends them to each engine on a schedule, then parses every answer for brand mentions, cited URLs and tone. That is the entire method, and its credibility depends on being upfront about what it is.
What is a prompt set, and why does it decide your numbers?
A prompt set is a synthetic list of questions you choose to stand in for real demand. Fifty prompts about "best project management software for agencies" will produce a very different score than fifty prompts about your brand name. Neither is wrong. They answer different questions.
Build the set from three layers: unbranded category questions, comparison and shortlist questions, and branded questions. Then freeze it. A prompt set you keep editing produces a trend line that measures your edits, not the market.
Which numbers are worth watching?
| Metric | Question it answers | What it does not tell you |
|---|---|---|
| Visibility Score | How present am I overall? | Whether anyone actually asked those prompts |
| Share of Voice | How present am I next to the field? | Anything about brands outside your tracked set |
| Citation share | Is my own site the source, or someone else? | Whether the citation was flattering |
| Sentiment | How am I described when named? | Why the model formed that view |
| Prompt coverage | Where are the holes? | How hard each hole is to close |
Citation share is the one most teams underuse. Being described accurately by a model that cites a review site is a different problem from being described accurately by a model that cites you, and the fixes are different. Visibility Score and Share of Voice are covered in more detail on the features page.
What can prompt-based tracking not tell you?
Four honest limits, worth saying out loud before you present a dashboard to a CMO.
- It does not measure demand. A prompt set is a sample you invented, so a Visibility Score is not a market share figure and should never be presented as one.
- It does not attribute revenue. Mentions inside an answer leave no referrer, and the Pew data above shows how few of them turn into clicks at all.
- It is noisy at the single-run level. Language models are non-deterministic, so the same prompt run twice can cite different sources. One check is an anecdote. A fixed prompt set run repeatedly over weeks is a signal.
- It is not your personal result. Answers vary by country, language, account history and index freshness, so a tracked answer is one plausible answer, not the only one.
What is a good AI visibility score?
There is no published industry benchmark, and any vendor quoting a universal "good" number is guessing. Scores depend on how broad your prompt set is, how crowded your category is and how old your brand is, so two companies with identical scores can be in completely different positions.
Use three internal reference points instead. First, your own trend: is the score moving up across a stable prompt set. Second, relative position: your Share of Voice against the set of names that actually appear alongside you. Third, coverage: the count of tracked prompts where you never appear at all, which is usually the most actionable number on the page.
For a live reference point that needs no account, the public visibility leaderboard shows how brands currently rank in AI answers.
Why is your brand invisible in AI answers?
Usually one of five reasons, and they are diagnosable.
- Your pages answer a keyword rather than a question, so no self-contained passage matches the sub-queries an engine actually issues.
- The answer exists but is buried below 800 words of preamble, past the point where a retrieval step would extract it.
- The content is unciteable rather than invisible: no dates, no numbers, no named sources, nothing a model can safely repeat.
- The problem is off your site entirely, because the engine leans on third-party pages that discuss your category without you.
- It is technical, and the substance lives in JavaScript, PDFs or images that never resolve into crawlable text.
What actually moves AI visibility?
The strongest evidence is a peer-reviewed paper. Aggarwal and colleagues, in "GEO: Generative Engine Optimization" (KDD 2024), tested content edits across a 10,000-query benchmark and found that adding direct quotations from credible sources raised a source’s visibility by roughly 40% on their headline metric, with adding statistics and explicit citations both worth around 30%. Keyword stuffing performed worse than the untouched baseline.
Read that with the caveat the authors themselves supply. Most of those experiments ran against a generative engine the researchers built, not a live commercial one. When they validated on Perplexity, the quotation gain measured 22%. The direction is trustworthy. The exact percentages are not promises, and they were measured on 2023-era retrieval.
Off-site signals appear to matter too, though the evidence is weaker. Ahrefs studied 75,000 brands in December 2025 and found brand mentions on YouTube correlated with AI citation at about 0.74 and branded web mentions at roughly 0.66 to 0.71, while Domain Rating sat between 0.27 and 0.33. The authors are explicit that correlation is not causation, and big brands accumulate both mentions and citations for the same underlying reason. Read it as a lean toward earning real brand presence, not as a lever with a guaranteed return.
The practical priority order that follows: put a liftable answer near the top of every page, attribute your facts to named sources with dates, cover the sub-questions rather than the head keyword, and build genuine mentions on the platforms your buyers already read. That is the core of generative engine optimization.
How does AI visibility differ from engine to engine?
Substantially, which is why a single blended score hides more than it shows. An analysis of more than 680 million citations collected between August 2024 and June 2025 found each engine leans on a different source mix.
| Engine | Top cited domain | Share of that engine’s citations | Runner-up |
|---|---|---|---|
| ChatGPT | Wikipedia | 7.8% | Reddit, 1.8% |
| Google AI Overviews | 2.2% | YouTube, 1.9% | |
| Perplexity | 6.6% | YouTube, 2.0% |
Two practical consequences. A brand can be well established in one engine and absent from another, so per-engine breakdowns matter more than the average. And these mixes shift with platform deals, licensing agreements and model updates, so any engine-specific tactic needs re-checking rather than being set once.
How often should you measure AI visibility?
Weekly for the tracked prompt set, monthly for the analysis. Weekly runs give you enough samples to average out the non-determinism described above. Monthly is the honest cadence for drawing conclusions, because content changes need time to be crawled, indexed and picked up by retrieval.
Add an off-cycle run after two events: a significant content push on your side, and a visible model or engine update. Comparing the same frozen prompt set before and after is the closest thing to a controlled test available in this channel.
How do you turn AI visibility data into action?
A score on its own changes nothing. The loop that works is measure, find the gap, publish against it, then measure the same prompts again.
Optimization for AI runs that loop end to end. It tracks your domain across ChatGPT, Perplexity, Google AI Overviews, Gemini and Claude, alongside your classic Google organic position, and computes Visibility Score and Share of Voice from the prompt set you define. It then surfaces the prompts where you are missing, generates and publishes articles built to answer those specific questions, and re-measures the same set. You can see how the measurement works in detail.
The point of the closed loop is falsifiability. If a published page does not move the prompts it was written for, you learn that within a measurement cycle instead of assuming it worked.
AI visibility FAQ
Sources
- https://ahrefs.com/blog/ai-overview-citations-top-10/863,000 SERPs and 4M AI Overview URLs, March 2026: 37.9% of cited URLs also rank organic top 10
- https://blog.google/products-and-platforms/products/search/google-search-ai-mode-update/Google’s own definition of the query fan-out technique in AI Mode
- https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/Pew Research Center, 68,879 searches from 900 US adults: 8% vs 15% click rate, 1% in-summary clicks
- https://searchengineland.com/ai-referrals-engagement-travel-sites-adobe-480445Adobe data, June 2026: AI referrals to US travel sites +194% YoY, 70% longer visits, 28% lower conversion
- https://support.google.com/webmasters/answer/16984139?hl=enSearch Console generative AI report covers AI Overviews and AI Mode impressions only
- https://arxiv.org/abs/2311.09735Aggarwal et al., GEO (KDD 2024): quotations ~40%, statistics and citations ~30%, keyword stuffing below baseline; simulated-engine caveat
- https://ahrefs.com/blog/ai-brand-visibility-correlations/75,000 brands, December 2025: YouTube mentions ~0.74, branded web mentions 0.66-0.71, Domain Rating 0.27-0.33
- https://www.tryprofound.com/blog/ai-platform-citation-patterns680M+ citations, Aug 2024 to Jun 2025: per-engine source mix for ChatGPT, AI Overviews and Perplexity
Last reviewed: 2026-07-31
See where you stand today
Point it at your domain and let one prompt set run. The first useful output is usually not the score, it is the list of category questions where a model recommends someone else.
Check your AI visibility