Resources · Strategy
Reading a visibility score: trend over number
4 min read · updated July 7, 2026
The first thing everyone does with an AI visibility tool is fixate on the number. "We're at 34%!" The second thing they do is panic when it's 28% the next day. Neither reaction is warranted, and understanding why makes the metric genuinely useful.
AI answers are samples, not facts
Ask the same assistant the same buying question five times and you will not get five identical brand lists. Models sample; retrieval varies; the round-up article cited at 9am may not be the one cited at 4pm. So a visibility score is an estimate from repeated sampling: of the times your buyers' questions were asked, how often were you named?
That has an immediate consequence: any single run (including a manual "let me just check ChatGPT") is one draw from a distribution. It can't tell you your visibility any more than one coin flip can tell you a coin's bias.
What the confidence band is telling you
An honest visibility score comes with a band ("34%, ±8"). The band narrows with more runs and widens with fewer: it's the tool admitting how much evidence sits behind the number. Two rules of thumb:
- Within-band movement is noise. 34% → 30% with overlapping bands is the same number twice. Don't report it, don't react to it.
- A move is real when the bands separate, or when the same direction persists across several consecutive readings. Direction sustained over weeks is signal; any single day is not.
Trend beats level; share beats trend
The absolute level of your score is shaped by things you don't control: how listy your category's answers are, how many brands fit in one. Comparing your 30% to another industry's 60% is meaningless. Two comparisons that do mean something:
- You vs. you, over time. The whole point of a stable prompt set is that the trend line is causally interpretable: you changed sources, the line moved.
- You vs. your competitors, same questions, same days. Share of voice controls for everything about the category. If answers name four brands and you're moving from the fifth to the third, revenue follows, whatever the absolute percentage says.
Judging whether your fix worked
The discipline that makes visibility tracking worth the money: when you ship a change (a comparison page, a review push, a round-up placement), snapshot the baseline and evaluate only the runs after the change. Then wait for enough runs: judging on two days of data is how teams convince themselves nothing works. Give a change a few weeks and a dozen runs before you call it, and expect platforms to move at different speeds: live-browsing surfaces react in days, model-memory surfaces take months.
Geofound bakes this in: baselines on every completed action, before/after measured automatically, and verdicts withheld until there's enough data to be honest. If you don't have a baseline yet, the free scan is a one-minute start: today's answer, today's competitors, in writing.
See who AI recommends in your category
The free scan asks ChatGPT three of your buyers' questions and shows whether you're in the answer, in about a minute.
Run the free scan →