Guide

Why do AI visibility numbers change, and why do tools disagree?

Jan NovakGEO Expert, FullReach AI · Updated
Share this

An AI visibility number depends on five choices: what counts as a hit, the platforms, the prompts, the rivals you monitor, and the day. In our measurement, five common ways of counting gave one brand between 5% and 38% on the same 1,008 answers. Counted on one platform only, it scored 30% on ChatGPT and 8% on Perplexity. Counted over every platform, its daily figure stayed between 18% and 21%.

What did we measure?

We used our own project, which tracks the category we sell in: AI visibility tools. It asks 48 prompts every day in Ireland, in English. It asks them on 7 AI platforms: ChatGPT, Gemini, Perplexity, Copilot, Google AI Mode, Google AI Overviews and Grok. Over three days, from 25 to 27 September 2026, it collected 1,008 answers, 144 on each platform.

The project monitors five rivals, Brand A to Brand E. None of them is FullReach AI. Brand A appears in the most answers. A brand counts when its name, or a spelling that FullReach AI knows for it, appears in the text of the answer.

Then we calculated Brand A's figure again and again, from the same 1,008 answers. Each time, we changed one choice that a tool or a report makes.

What we changedBrand A's figure
What counts as a hit5% to 38%
Which platforms count8% to 30%
Which prompts count14% to 55%
Which day18% to 21%

Some figures below carry a margin of error, because a different set of prompts would give a different figure. To find it, we drew 48 prompts at random from our 48, with repeats allowed, 4,000 times. The margin holds 9 of every 10 draws. This is one category in one market over three days, so read the numbers as an example, not a law.

How much does the figure move from day to day?

Counted over every platform, very little. Brand A's visibility, the share of answers that named it, was 19%, 18% and 21% on the three days. Both daily changes lay within the margin of error.

One platform can really move while the total stays calm. On Gemini, Brand A appeared in 8%, 15% and 19% of the answers. The rise from 8% to 19% had a margin of error from 4 to 19 points. Even its low end is above zero, so the rise was more than noise. On ChatGPT and Google AI Mode, Brand A fell over the same days, so the total hardly changed.

Single answers move the most. A pair here is one prompt on one platform, asked on two days in a row. For Brand A, 8% of the 672 pairs changed: the brand appeared where it had not, or the other way round.

The list of names changes even more. In 211 pairs, at least one of the two answers named a monitored rival. Only half of those, 106 pairs, named exactly the same rivals on both days.

These figures rest on two day-to-day changes only. Compare weeks, not days, and read one platform's daily figure as noise until a whole week shows the same change.

Why do two tools show different numbers for the same brand?

Each tool makes choices before it shows a number. When two tools make different choices, they can both be right and still disagree by a factor of three or more.

What counts as a hit

We counted Brand A and Brand B in five common ways on the same answers:

What countsBrand ABrand B
Answers that name the brand (visibility)19%15%
Its share of all rival mentions (share of voice)35%27%
Share of voice, weighted by position38%28%
Answers that name the brand first11%8%
Answers that recommend the brand5%6%

Some tools give an early mention more weight. For the weighted row, the first brand in an answer counts 1, the second counts ½ and the third counts ⅓.

The gap between the two brands also depends on the count. Brand A led Brand B by 4 points on visibility and by 8 on share of voice. On recommendations, the two were level within the margin of error.

"Recommended" also depends on a model that reads each answer and labels it. Two tools with two such models can label the same answer differently. We did not check these labels by hand.

So first ask what a number counts. A share of voice of 35% and a visibility of 19% can describe the same brand in the same answers. Our guide on share of voice explains the difference.

Which platforms count

Brand A's visibility was 30% on ChatGPT (margin of error 20% to 40%) and 8% on Perplexity (3% to 15%). On Perplexity, Brands B and C appeared in about as many answers as Brand A.

Platforms countedBrand ABrand B
ChatGPT only30%20%
ChatGPT, Perplexity and Google AI Overviews18%13%
All 7 platforms19%15%
Perplexity only8%10%

We tried every choice of three platforms, 35 in all. They gave Brand A between 12% and 28%. So a tool that covers three platforms and a tool that covers seven can report the same answers correctly and still disagree. Our comparison of AI visibility tools shows which platforms each tool's plans include.

When you compare totals over several platforms, check how many answers each platform gave. A platform with fewer answers counts less in the total. This happens when a platform fails often, or when Google shows no AI Overview for many questions. In our data, every platform gave 144 answers.

Which prompts count

The prompts are the biggest choice. From the same answers, Brand A's visibility was 14% over the 41 prompts that name no rival. Over the 8 prompts that ask for tools, prices or trade-offs in our category, it was 55%. Their margins of error, 8% to 21% and 39% to 68%, do not overlap.

Even two random sets of 25 prompts disagree. We drew 3,000 pairs of such sets from our 48 prompts. In a typical pair, the two sets differed by 4 percentage points on Brand A. In 10% of the pairs, they differed by 10 points or more. Our guide on which prompts to track shows which kinds of prompt name brands.

Which rivals count

Share of voice divides by the rivals that you monitor. When we took Brand A off the list, Brand B's share of voice rose from 27% to 42%. This is arithmetic, not a change in the market: not one answer changed. Two tools that monitor different rivals show different shares for the same brand.

Answers that are missing

Some prompts get no answer on a day. A platform can fail, or Google can show no AI Overview for a question. One tool leaves these out. Another counts them as answers that name nobody. In our data, 16 attempts failed, and FullReach AI asked each of them again the same day. So every prompt got an answer on every platform on every day. For questions where Google often shows no AI Overview, the two tools can report different figures. Our data had no such question, so we could not measure how different.

Your figure dropped. What do you check?

A drop can be real. Check these causes in this order, and stop when one explains the drop:

  1. A setting changed. Somebody added or paused prompts, changed a wording, added or removed a rival, or added a spelling of a brand's name. A filter now shows another platform, market or period.
  2. One platform moved. Break the figure down by platform. When one platform explains the drop, read its answers from before and after, and check its model version.
  3. The window is small. In our data, the daily figure over all platforms moved by at most 2 or 3 points. A week holds seven times as many answers as a day, so its figure moves less by chance. We did not measure whole weeks.
  4. The answers changed. When no setting changed and the drop shows on several platforms for a whole week, the answers really name you less. Read them to see who took your place.

How do you compare two numbers fairly?

  • Compare a brand with itself, in one tool, over time. A trend inside one tool keeps every choice fixed. A level from one tool next to a level from another compares their choices, not the brands.
  • Name the unit. Visibility, share of voice and recommendation answer three different questions. Write the unit beside every number that you report.
  • Name the platforms, the prompts and the rivals. Nobody can check a figure without them. When one of them changes, start a new trend line or recalculate the old weeks.
  • Compare weeks, not days. One platform on one day holds too few answers.

One line in a report can carry all of this. For example: "Visibility 19%, over 1,008 answers from 25 to 27 September. 48 prompts, 7 AI platforms, Ireland in English, 5 rivals monitored."

FullReach AI keeps these choices fixed and visible. Every number opens into the answers behind it, with the date, the platform, and the model version where the platform reports one. After you change the list of rivals, it calculates past weeks again with the new list, so the trend compares like with like.

For how FullReach AI collects the answers, read the methodology.

More guides

Keep reading.

  • GuideUpdated

    Best AI visibility tools in 2026: nine compared

    Compare nine AI visibility tools on the platforms each plan includes, the price per prompt, the evidence behind each number and the limits. Each price carries the date we read it.

  • GuideUpdated

    How many prompts do you need to track AI visibility?

    How many prompts give a stable AI visibility number? We measured it on real answers from 7 AI platforms. The guide also covers markets, languages and which questions to track.

  • GuideUpdated

    How to measure share of voice in AI answers

    What does share of voice in AI answers measure, and how does it differ from visibility? We show why it moves when the monitored rivals change, on real answers from 7 AI platforms.

  • GuideUpdated

    Which prompts should you track for AI visibility?

    Which kinds of prompt name brands in AI answers? We compare requests for tools, how-to questions, neighbouring topics and prompts that name a rival, on 1,008 answers from 7 AI platforms. A checklist ends the guide.

All guides