Field Notes · 6 MIN
How do you measure what AI assistants say about your brand?
Query the real assistants from a cold browser, run each question several times, and save the full answer. A method learned from shipping one wrong audit.
You measure it by asking the real assistants your real buyer questions from a clean, logged-out browser, running each question several times, and saving the full text and a screenshot of every answer. Anything weaker measures something else. We learned this by shipping an audit that reported the wrong thing, and this post is the method we rebuilt afterward.
- A search tool is not an AI assistant. If a report names an engine, that exact engine must have been queried and the answer saved.
- A signed-in session that has already discussed your brand measures your chat history, not the market. Measure cold.
- Answers vary between runs, so run each query several times and report a rate, never a yes or no.
- Record whether the brand was named, whether its own site was cited, and whether a different company was returned, as three separate findings.
- Nactore is an AI-native software engineering partner, and we build this kind of measurement harness into the products we ship.
What did we get wrong the first time?
We produced an audit that said a brand was invisible to AI assistants. We had never queried one. Subagents had run a web search tool, and the report then described the results as "Google and AI assistants." A search tool is not Google's results page, and it is not ChatGPT.
The founder opened ChatGPT, asked the pricing question, and got a correct answer that cited the brand's own pricing page. The error was caught before it reached the client. A second claim in the same report, that the brand had no presence in an app store, was inferred from the fact that competitors held the ranking URLs. A single direct check showed the listing was live. Two mistakes, one root cause. We reported on a system we had not measured, and we asserted an absence we had not tested.
Why must the session be cold?
A signed-in assistant that has already talked about the brand will find it. You are then measuring your own conversation history. We ran the same pricing question twice on the same day.
| Session | What the assistant returned |
|---|---|
| Signed in, primed thread | Correct pricing, cited the brand's own site |
| Cold, anonymous browser profile | Answered about a different company with a similar name and asked if we meant another service |
The cold session is what a stranger sees, so the cold session is the measurement. Use the signed-in session only to demonstrate the contrast.
Results are also query by query. On the same cold profile, a safety question found the brand and cited its app listing, while the pricing question returned the wrong company. Never generalize from one query.
What does the harness look like?
We use Playwright for Python to drive a real browser, with no API keys, because the consumer product is the thing buyers actually use. The harness runs each query, saves answers/<engine>/qNN.txt and a matching screenshot, and does the classification later from the saved text.
- Load a clean profile. No cookies, no login, no history.
- Submit the query exactly as a buyer would phrase it.
- Wait for the answer to finish. Wait until the answer text stops growing for a set number of seconds, instead of chasing per-engine "done" selectors that change every few weeks.
- Save the full text and a screenshot. The screenshot is the evidence a client cannot argue with.
- Repeat three times for anything client-facing.
- Classify offline from the stored text.
Engines behave differently in practice. In our runs, ChatGPT worked headless and logged out. Gemini also worked anonymously, though its page always renders a "Sign in" button, so that string is not a signed-out signal, and it returns the whole page text so you must split on the answer marker. Google blocked headless access almost immediately, so AI Overviews needed a real headed window. Those are observations from our runs, and they can change, so re-test before relying on them.
Why store the full answer instead of a score?
Classification logic is the part you will fix most often. If you save only booleans, every matcher fix means re-running the whole query set. If you save the full text, a fix costs a second.
That paid for itself. A pattern meant to catch a competitor with a brand-like name flagged the verb "befriend" in an ordinary sentence. Namesake patterns must require product context, never a bare substring.
One more rule matters here. An answer you could not read is measured: false, never a miss. A failed extraction recorded as "not found" is how the first wrong report happened, twice.
What should you record for each answer?
Collapsing everything into one visible or not visible number is what produced the wrong report. Record three distinct things.
- Named. The brand appears in the answer at all.
- Cites own site. The answer links to the brand's own domain.
- Wrong brand returned. The answer served a competitor or namesake instead.
A brand that is named in two of three runs and replaced by a namesake in the third has an unstable identity. That inconsistency is a finding worth reporting on its own. For the wider measurement picture, see how to measure AI search visibility.
Why run each query several times?
Answer engines are not deterministic. The same cold query on the same day returned the wrong company on one run and the correct one on the next. This is not a footnote. A single run is an anecdote.
Run each query at least three times, report the rate, and keep the wording, market, and cadence fixed so trends mean something. Never write "the assistant does not recommend you" from one run. Write "named in one of three runs."
Frame engines as coverage, not a contest. Show what each engine returns side by side, then the combined gap. An engine that already names the brand is proof the fix works, not a hole in the argument.
What are the limits of this method?
It is a sample, not a ranking. Different locations, accounts, and product versions will give different answers. It also depends on third-party interfaces that change, so the harness needs maintenance. And it measures what the engines say, not whether anyone clicked.
Pair it with referral data and Search Console, as covered in how to track AI citations. For measuring model output inside your own product, the equivalent discipline is AI evals before production.
Frequently asked questions
Can I use a web search API instead of querying the assistant?
Not if your report names the assistant. A search API returns search results, which is a different system from the assistant's answer. Query the engine you claim to measure.
Why not use an official API for each assistant?
An API answer is not always what the consumer product shows a buyer, and keys cost money at volume. We drive the consumer product directly and save screenshots as evidence.
How many runs are enough?
Three is our minimum for client-facing work. More runs tighten the rate, but the key habit is to report a rate instead of a binary.
What if an engine blocks automated access?
Record the query as not measured, use a real browser window if the engine requires one, and say so in the report. Never log a blocked query as a miss.
Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.