GEO · 7 MIN
How to measure AI search visibility, and why one ChatGPT answer proves nothing
AI answers are non-deterministic, so a single screenshot is not a measurement. The query set, the run count, the metrics worth logging, and a weekly loop you can run yourself.
Ask ChatGPT the same commercial question five times and you can get five different sets of brand names back. That single fact decides how AI search visibility has to be measured, and this guide covers the metric that actually works, the query set to build first, what to log every week, and the loop you can run without buying a tool.
- AI answers are non-deterministic, so the useful metric is mention frequency across repeated runs, not presence in one answer.
- A fixed query set built before any optimization begins is what makes later numbers comparable.
- Logging the sources an engine cited matters more than logging the answer text, because those sources are your target list.
- Search crawlers and AI crawlers are different bots, so verify access before you measure anything.
- Nactore runs this loop as part of one program where GEO, SEO, and the engineering behind them sit with a single partner.
What AI search visibility means
AI search visibility is how often your brand appears inside the answers generative engines compose for the questions your buyers ask. It is not a ranking, because there is no results page. It is closer to share of voice inside a recommendation, and it is the outcome that Generative Engine Optimization exists to move.
The practical unit is the query. For any given buyer question, either your name shows up in the composed answer or it does not, and the only honest way to describe your position is a rate. Named in seven of ten runs is a measurement. Named once in a screenshot is an anecdote.
Why a single answer is not a measurement
Generative engines sample. They also retrieve fresh pages at query time, weigh sources differently between runs, and change behavior when a model is updated. Two people asking the identical question in the same hour can get different lists.
This means three common practices are worthless. Taking one screenshot as a baseline. Declaring victory from one good answer. Panicking about one bad one. All three treat a probabilistic system as if it were a deterministic one.
If your measurement cannot survive being run again tomorrow, it was not a measurement.
Build the query set before you build anything else
The query set is the whole apparatus. Get it right once and every number you produce afterward is comparable. Change it midway and you have lost your history.
Build it in three buckets, and keep the buckets separate in reporting because they behave differently.
- Proof queries. The ones a prospect types about your category, like best provider for a given service in a given city. Low volume, high stakes, because this is the query your buyer runs before the call.
- Buyer queries. The problem-shaped questions people ask before they know a vendor exists. These carry the actual pipeline and are usually easier to win.
- Education queries. Definitional and how-to questions in your domain. These are where citable content earns mentions that later support the other two buckets.
Thirty queries split across the three buckets is enough to start. Fifty is comfortable. Two hundred is a tool vendor's problem, not yours.
The five things worth logging
Per query, per engine, per run, record these and nothing more. The discipline of a small schema is what keeps a weekly loop alive past month two.
| What you log | Why it matters |
|---|---|
| Mentioned, yes or no | The raw signal that becomes your mention rate |
| Position among named brands | First mention behaves very differently from fifth |
| Sources cited in the answer | Your target list for placement work, the most actionable field here |
| Competitors named | Tells you who the engines currently trust in your category |
| Engine and date | Answers drift after model updates, so undated data is noise |
The third row is the one most people skip and the one that pays. Every source an engine cites for your category is a page you now know matters. That list, not your own blog, is where the next quarter of work comes from.
Cover the engines your buyers actually use
Four surfaces are worth the effort, and they retrieve differently enough that results do not transfer between them.
- ChatGPT, because of raw usage, and it is the one your buyer most likely opens first.
- Perplexity, because it cites heavily and visibly, which makes it the best early diagnostic for whether your sources are working.
- Google AI Overviews, because it sits on top of the search results your SEO already fights for.
- Claude or Gemini, to catch the case where one engine's retrieval is simply not seeing you.
Before your first measurement run, check that AI crawlers can reach your site at all. They use different user agents from Googlebot, and a site can serve Google perfectly while returning errors to the crawlers that fetch pages for AI answers. Measuring a site that cannot be read produces a real number for the wrong reason. Our free AI crawler check tests this in about ten seconds, no signup.
A weekly loop you can run yourself
This is deliberately small. A loop that takes twenty minutes survives. A dashboard project does not.
- Run the set. Every query, every engine, five runs each. Script it if you can, do it by hand if you cannot.
- Log the five fields. A spreadsheet is genuinely fine for the first quarter.
- Compute mention rate per bucket. One number for proof queries, one for buyer queries, one for education queries.
- Update the source list. Add every newly cited domain. Mark the ones you could plausibly appear on.
- Pick one action. One placement to pursue or one page to fix. Not ten.
Review the trend monthly, never weekly. Week-to-week movement in a probabilistic system is mostly noise, and reacting to it will send you in circles.
Set expectations against the buckets while you do it. Education queries move first, often within weeks, because a genuinely useful page gets retrieved on its merits, which is largely why AI tools cite some brands and not others. Buyer queries follow. Proof queries move last and slowest, because they depend on third-party pages you do not control. A reasonable first target is a measurable improvement in education-query mention rate inside six weeks, and any movement at all on proof queries by week ten. Anyone promising faster than that on comparison queries is describing something they cannot control.
Frequently asked questions
How many runs per query are enough?
Five is the practical floor for spotting a real change, and it keeps a thirty-query set inside a manageable weekly session. Go to ten runs only for the handful of proof queries where you need confidence in a small shift.
Do I need a paid AI visibility tool?
Not to start. A spreadsheet and a fixed query set will tell you more in the first quarter than a tool you have not calibrated, because you will actually understand where the numbers come from. Buy tooling once the loop is a habit and the manual runs are the bottleneck.
Why log cited sources instead of the answer text?
Because sources are actionable and answer text is not. Knowing an engine trusts a particular directory or comparison page for your category tells you exactly where to work next. The prose it generated tells you nothing you can act on.
Does this replace SEO reporting?
No. It sits beside it. The same content and technical foundations feed both, and the honest picture of your visibility now needs both numbers on the same page.
The takeaway
Fix the query set, run it repeatedly, log the sources, and read the trend monthly. That is the entire method, and it is more rigor than most of the market currently applies. Nactore runs GEO and SEO as one program on a site we engineer for both, so measurement, content, and the code underneath it stay with one partner. Build, ship, rank, grow.
Want to apply this to your business?
Tell us the goal. A founder will reply with an honest next step within one business day.