GEO · 6 MIN
How do you track AI citations and mentions for a SaaS product?
A practical method to track whether AI tools mention and cite your SaaS: prompt panel design, what to record, how to avoid false confidence, and how to report it.
You track AI citations for a SaaS product by running a fixed set of real buyer questions against the AI tools on a schedule, recording whether you are mentioned, whether a page of yours is cited, and whether the description is accurate, then comparing the trend over months. No single score is reliable, because answers vary by wording, location, and day. The value is in a consistent method and in the errors it exposes.
- A prompt panel is a fixed list of buyer questions run on a schedule under controlled conditions.
- Record four things per answer. Mention, citation with URL, position among alternatives named, and factual accuracy.
- One answer is a sample. Only repeated runs across many prompts show a trend.
- Combine the panel with Search Console, referral analytics, and qualified enquiries so the report ties to revenue.
- Be explicit about what you did not test, and never claim absence from an engine you did not query.
- Nactore builds the measurement harness as software with evals, then uses the findings to guide content and engineering work.
What should you actually measure?
Four views together give a fair picture. None of them is a ranking.
| View | Question it answers | Where the data comes from |
|---|---|---|
| Prompt visibility | Are we named and cited for buyer questions | Your prompt panel |
| Answer accuracy | Is what AI says about us correct | The same panel, reviewed by a human |
| Referral traffic | Do people arrive from AI tools | Analytics referral sources |
| Business outcomes | Do those visits become pipeline | CRM and signup attribution |
Google adds a fifth data point for its own features. Its documentation says traffic from AI Overviews and AI Mode is included in the Search Console Performance report under the Web search type. There is no separate AI filter, so you cannot isolate it there, and you should say so in reports.
How do you design a prompt panel?
The panel is the core of the method. Build it from real language, not from your own product vocabulary.
- Collect the source questions. Use sales call notes, support tickets, community threads, and Search Console queries. Real phrasing matters.
- Group by intent. Category discovery ("best tools for X"), task questions ("how do I do X with Y"), comparisons ("A vs B"), and evaluation ("is A secure for regulated data").
- Aim for 20 to 40 prompts. Fewer is too noisy, and many more becomes expensive to review honestly.
- Include named competitors. Comparison prompts show how you are described next to alternatives.
- Freeze the wording. Changing a prompt breaks the trend. Add new prompts as a new batch.
- Set conditions. Same market, same language, signed out or in a clean profile, same cadence.
We use the intent groups to report separately. A product can be well described on task questions and invisible on category questions, and those need different fixes.
What do you record for each answer?
Capture enough to audit later. Store the raw text, because summaries lose detail.
- Date, tool, and mode. Note whether a search or browsing feature was active.
- Mention. Was your product named at all.
- Citation. Was a link given, and to which URL, yours or a third party.
- Alternatives named. Which competitors appeared, and in what order, labeled as observed and not as a ranking.
- Accuracy. Mark each factual claim about you as correct, outdated, or wrong.
- Sources. The domains the answer relied on.
The sources column is the most useful. If assistants keep citing a directory, a forum thread, or a competitor comparison, that tells you where evidence about your category lives. For how the evidence layer works, see why AI tools cite some brands.
How do you run it without fooling yourself?
Measurement of AI answers is easy to get wrong. These are the traps we watch for.
- Personalization. A logged-in account with history can favor you. Test cold, from a clean profile.
- Averaging away errors. A 60 percent mention rate hides a recurring wrong pricing claim. Review inaccuracies separately.
- Single runs. Repeat each prompt a few times and note the spread, since answers can differ between runs.
- Claiming absence. If you did not query an engine, say "not tested," not "not present."
- Confusing search and model knowledge. A tool with live search may answer differently from the same tool offline.
- Changing the prompts mid-stream. It destroys the baseline.
Automating the runs saves time and reduces inconsistency. We describe a browser-based harness in measuring AI answers with Playwright, and the same discipline of fixed inputs and logged outputs appears in LLM observability.
Keep a "not tested" list in every report. Naming the engines and modes you skipped protects the credibility of everything you did measure.
How do you connect citations to revenue?
Citation counts alone will not satisfy a CFO. Link them to outcomes in three steps.
- Referral grouping. Create a channel grouping in analytics for visits from AI tools. Our GA4 walkthrough is in measuring AI search visibility in GA4. Referrers are often missing or stripped, so treat the number as a floor.
- Ask at signup. Add a "How did you hear about us" field with an option for AI assistants. Self-reported data is imperfect, and it catches visits that analytics miss.
- Review the pipeline. Check whether deals that mention an assistant in notes close differently from others, and say plainly when the sample is too small to conclude anything.
Report these as directional. Overclaiming is the quickest way to lose trust in the whole program.
What does a monthly report look like?
Keep it to one page that a head of growth can use.
- Coverage. Which engines and modes were tested, and which were not.
- Trend. Mentions and citations by intent group compared to last month.
- Errors. Wrong or outdated claims about the product, each with the source it came from.
- Source map. The domains most often cited in your category.
- Actions. The three fixes planned, such as updating a page, publishing a comparison, or correcting a directory listing.
- Outcomes. Search Console trends, AI referral visits, and enquiries.
Frequently asked questions
Can we track AI visibility with a single tool or score?
Third-party tools exist, and each uses its own sampling method, so read how a tool collects data before trusting a score. We prefer a documented panel you control, with a tool as a cross-check.
How often should we run the panel?
Monthly is enough for most SaaS teams. Run it weekly during a major content push or after a big site change, so you can see the effect.
Do citations equal traffic?
No. Many answers cite a page and few readers click. Track mentions, citations, and visits as separate numbers.
What should we do when an AI tool states something wrong about us?
Find the source it relied on, correct or update that page, and align your own pages on a single description. You cannot edit an answer directly, so the fix is the underlying evidence.
Measure, fix, repeat
A steady panel, honest reporting, and quick fixes to bad sources beat any one-time audit. Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.