Measuring AI Search Visibility: A Practical Tracking Method
Ranking third on Google says nothing about what ChatGPT tells a prospect. Here is the method and the sheet to track your AI visibility, month after month.

Why AI visibility needs its own measurement
A strong Google ranking says nothing about what ChatGPT tells a prospect who asks for a recommendation. Answer engines do not rank pages, they compose a synthesis from several sources, often without citing a single link. A company can dominate its market on Google and stay invisible in Perplexity, or the reverse. Position tracking on its own no longer covers what matters, which is why answer engines need a separate measurement protocol.
That protocol is not complicated. It rests on three building blocks: a prompt panel representative of your industry, a repeatable questioning method, and a tracking sheet that captures change over time rather than a single snapshot. The hard part is not technical, it is methodological: without discipline in repeating the same protocol month after month, the numbers you collect stop being comparable and the dashboard turns decorative.
The three metrics that actually matter
Presence rate is the percentage of industry-representative prompts for which your brand shows up in the generated answer, across every engine tested. It is the reference metric: it reads independently of context and stays comparable month to month.
Two other metrics complete the picture. Citation share measures how often your domain appears as an explicit source, not merely as a text mention: a mention without a link generates no measurable traffic. The third, sentiment, qualifies the tone of the answer (neutral, favourable, or comparative against a named competitor). Tracked monthly on the same prompt panel, these three numbers form the basis of a usable GEO dashboard.
A fourth signal is worth noting without tracking it monthly: where the brand name lands inside the answer. A brand named in the opening paragraph does not carry the same weight as one buried in a list of ten names at the end of the response, even though both count as a "presence".
Building a representative prompt panel
The most common mistake is testing three or four obvious queries and drawing a general conclusion from them. A useful panel covers at least three question families:
- Generic queries about the service ("which agency should I pick for an e-commerce site in Geneva").
- Comparison queries that explicitly name competitors.
- Reassurance queries, asked by someone who already has a shortlist and wants to validate it.
Every question should be run in a logged-out session, with no conversation history and no personalised account: a connected account adjusts the answer based on your own search history, which distorts the comparison over time. Track at least twenty questions before a monthly presence rate starts to carry statistical meaning; below that, a single different answer swings the percentage by several points.
Weight the panel by stage in the buying journey too: a question asked by someone discovering the topic carries less commercial value than one asked by someone ready to sign. A small business gains more from weighting its panel toward bottom-of-funnel queries, where a citation is most likely to turn into an actual contact.
Keep the panel itself stable once you have built it. Swapping questions in and out every month feels tempting when a new competitor or a new service line appears, but it breaks the one thing that makes the exercise useful: a like-for-like comparison. Add new prompts as a separate, clearly labelled batch instead of replacing the original panel, so the historical trend line keeps its meaning.
Measuring engine by engine
Each answer engine behaves differently when it comes to citations, which rules out treating "AI visibility" as a single score.
| Engine | Citation source | Trackable via Search Console | Trackable via GA4 |
|---|---|---|---|
| Google AI Overviews | Standard Google index | Yes, via the search appearance filter | Partially, depending on outbound clicks |
| ChatGPT (browsing on) | Real-time web search | No | Yes, via a utm parameter or the chatgpt.com referrer |
| Perplexity | Proprietary index plus web search | No | Yes, via the perplexity.ai referrer |
| Copilot (Bing) | Bing index | No, but Bing Webmaster Tools exists | Yes, via the bing.com/chat referrer |
Two rows of that table overlap with tools you already have: Search Console for AI Overviews, GA4 for any engine that sends an identifiable referrer. Everything else, ChatGPT without browsing and most Perplexity answers, sits entirely outside standard analytics and can only be measured through manual querying or a dedicated tool.
One technical detail changes everything for ChatGPT: the answer differs depending on whether web browsing is turned on in the user's settings. A brand absent from a non-browsing answer can perfectly well appear once web search is enabled, which means a test protocol must record, at every measurement, which mode was used.
Manual tracking or a paid tool: the real cost comparison
Manual tracking costs nothing in licence fees but takes two to three hours a month to run twenty questions across four engines and log the results. That is enough for a company testing the exercise for the first time. A dedicated tool automates daily or weekly querying on a fixed number of contracted prompts (the Otterly.ai Lite entry plan, priced at 29 dollars a month in 2026, covers fifteen tracked prompts across six engines) and adds sentiment detection, something manual tracking can only approximate.
The right choice depends on volume: below thirty prompts to track, the gain from a tool remains marginal compared with its subscription cost. Above that threshold, the human time required to stay rigorous becomes the dominant cost itself, and automation becomes cheaper than the time it replaces.
Building a monthly tracking sheet
A tracking sheet that works fits on a single page and always follows the same structure: one row per prompt, one column per engine, each cell holding three pieces of information (presence yes/no, citation yes/no, position of the brand name in the answer). Add a date column and keep every month as a separate snapshot rather than overwriting previous data: it is the comparison between snapshots, not the absolute value of a single month, that reveals whether a technical fix actually paid off.
A serious AI visibility audit never stops at reading the engines once: it links every change in the sheet to a published action (a new page, a Schema fix, publishing an llms.txt file) to establish, month after month, what actually moved the needle. Without that "action of the month" column next to the numbers, the sheet quickly turns into a passive log nobody reopens.
Reading results without misleading yourself
A few points of change from one month to the next mean nothing while the panel stays small: answer engines are not deterministic, the same question asked twice on the same day can produce two different answers. Wait for at least two or three measurement cycles before concluding that a trend is real rather than statistical noise. The most reliable comparison remains a three-month rolling average, not the latest single reading.
Be equally careful comparing across engines: a 40% presence rate on Perplexity against 15% on ChatGPT does not mean Perplexity "prefers" your content, it may simply reflect very different query volumes for your industry on each platform.
When this method is not the right approach
This protocol assumes a minimum of differentiated content already published: without a page that directly answers your panel's questions, measuring your own absence every month adds no new information, you already know you will be absent. For a one-person structure without regular editorial output, a manual quarterly check of five to ten questions is enough; setting up a full monthly dashboard before there is content to measure amounts to instrumenting an empty room.
Similarly, if your business depends almost entirely on direct referrals or a local professional network, the time invested in this tracking returns less than the same hour spent on direct outreach. GEO tracking earns its keep when a measurable share of your prospects starts from a generic search, not when it starts from word of mouth you already have locked in.
There is also a data-quality floor to respect: if your industry generates fewer than a handful of relevant prompts worth asking in the first place, no amount of tooling turns that into a meaningful sample. In that case, a lighter qualitative check, reading a dozen real answers once a quarter, tells you more than a dashboard built on a panel too thin to be representative.
From measurement to action
A tracking sheet only has value if it triggers a decision. A presence rate that stalls despite recent content often signals a structural problem (see which Schema.org types actually matter for answer engines) rather than a content problem. Conversely, strong presence without an associated citation shows that content is being reused but not sourced, a signal that points to the clarity of your standalone passages more than to publishing volume.
For a full view of the levers that drive these three metrics, our overview of GEO and visibility in AI answer engines covers the six levers that come before measurement. Some teams prefer to hand this tracking to a tool that measures this score automatically rather than keeping the sheet by hand every month.
Browse all articles on the blog, or check our packages and pricing if you would like this protocol set up for your business.
Frequently asked questions
How many prompts do I need for a reliable panel?
Track at least twenty questions split across generic, comparison and reassurance queries. Below that threshold, a single different answer swings the presence rate by several points and the number loses its tracking value.
Does a logged-in ChatGPT account distort the measurement?
Yes. A personalised account draws on your own conversation and browsing history, which changes the answer. Any tracking measurement should run in a logged-out session or private browsing.
Can Google Analytics track traffic coming from ChatGPT?
Partially. GA4 can identify inbound traffic via the chatgpt.com or perplexity.ai referrer when a user clicks a cited link, but it captures nothing for answers that mention your brand without a clickable link.
Do I need a paid tool from month one?
No. Below thirty prompts to track, a manually kept sheet each month is enough and only costs time. A tool becomes worthwhile once prompt volume or tracking frequency exceeds what two to three hours a month can cover.
How often should the measurement be repeated?
A monthly cadence gives a stable reading for a company publishing regularly. A structure with no content output can limit itself to a quarterly check while it builds up content worth measuring.


