AI Visibility Tracking: A Monthly Measurement Routine
AI Visibility Tracking: A Monthly Measurement Routine
AI visibility tracking means checking, on a fixed schedule and with a fixed list of questions, whether assistants like ChatGPT, Perplexity, Gemini and Copilot name your brand when someone asks what your buyers ask. You need a prompt list, an identical method every month, and a spreadsheet, because one lucky answer proves nothing and one missing answer proves nothing either.
That last point is the whole reason this article exists. Assistants do not return the same answer twice, so the instinct of typing your category into ChatGPT once and drawing a conclusion is worse than useless: it produces confidence without information. What follows is a routine you can run in under an hour a month, what to write down, what your analytics will and will not show you, and where the measurement stops being able to tell you anything.
TL;DR
- Ask the same fixed list of 10 to 20 buyer questions every month, in the same assistants, and log the result. Same list, same method, or the numbers are not comparable.
- Run each prompt more than once. Answers vary between runs, accounts, regions and model versions, so a single check is closer to a coin flip than a measurement.
- Four signals are worth recording: prompt citations, how often you appear across the list, referral visits from assistants, and branded search demand.
- Your analytics can show visits arriving from assistants. It cannot show the answers where you were named and nobody clicked, which is most of them.
- Look at direction over three months, never at a one point move. And no tool, ours included, can promise you a citation.
Table of contents
- What AI visibility tracking actually measures
- Why a single check tells you nothing
- The four signals worth recording
- Building the prompt list your buyers would use
- The monthly routine, step by step
- What your analytics can and cannot show
- Reading the numbers without fooling yourself
- What to do with a month of bad results
- What tracking will never tell you
- How much time this takes
- FAQ
- Conclusion
What AI visibility tracking actually measures
Search rankings have a fixed object. A query, a page, a position, and a report that agrees with itself from one day to the next. AI visibility tracking has none of that, and pretending otherwise is where most of the confusion starts.
What you are measuring is a rate, not a rank. Out of the questions your buyers actually ask, in what share of the answers does your brand get named, and in what share does it get named with a link. There is no position 3. There is only present or absent, over and over, until a pattern shows.
That reframing has a practical consequence. You cannot track a keyword here, you track a question set. And because the object is a rate, the sample size and the stability of your method matter more than any individual result you will ever see on screen.
Why a single check tells you nothing
Assistants are not deterministic. Ask the same question twice in two fresh conversations and you can get two different sets of sources, two different orderings, sometimes two different framings of the answer. Nothing about your site changed in those thirty seconds.
Four other things move the answer independently of your content:
- Personalization and memory. A logged in account that has been discussing your category for weeks is not a neutral judge. It has context you did not intend to give it.
- Region and language. The same question in French and in English often surfaces a different source list, because the underlying material is different. If you sell in two markets, you have two measurements, not one.
- Model version. Assistants ship changes without telling you. A drop in your numbers can be their release note, not your fault.
- Whether the assistant searched the web. Some answers come from browsing live results, some from the model alone. Those two paths do not favour the same sources.
None of this makes measurement pointless. It makes single observations pointless. The fix is boring and it works: fixed prompt list, several runs per prompt, fresh sessions, and a written note of the conditions.
The four signals worth recording
One number cannot carry this. Four can, and each one answers a different question.
| Signal | Where you see it | What it tells you | How often |
|---|---|---|---|
| Prompt citations | Your own runs in each assistant | Whether you get named at all, and who gets named beside you | Monthly |
| Coverage rate | The same runs, counted across the list | Whether you are a one question fluke or a recognised option | Monthly |
| Assistant referrals | Your analytics, referral sources | How many people actually clicked through to you from an answer | Weekly glance |
| Branded search | Search Console, queries containing your name | Whether upstream exposure is creating demand for you by name | Monthly |
The first two you have to produce yourself, by hand or with a tool. The last two already exist in tools you have. Most teams watch only the third, which is the smallest of the four, and conclude that AI answers do nothing for them.
Branded search deserves a special mention because it usually moves first. Someone reads an answer that names four options, clicks none of them, and searches your name two days later. That journey leaves no referral, no source, no attribution. It shows up as a slow rise in queries containing your brand, and it is the closest thing to an early indicator you will get.
Building the prompt list your buyers would use
The list is the instrument. If it is badly built, everything downstream is decoration, so spend the hour it takes to get it right and then leave it alone.
Write 10 to 20 prompts across four families:
- Category questions. How someone with the problem describes it before they know any brand exists. "How do I get my shop found in AI answers." These are the ones that matter, and the ones you are least likely to appear in early.
- Solution questions. The shortlist requests. "Best tools to publish content on several channels automatically." Expect to be absent here for a long time on a young site, and record it anyway.
- Comparative questions. "What are the options for a small team with no marketing department." Answers here often list four or five names, and the list is the whole result.
- Brand questions. "What is distrify." Sounds vain, and it is the most useful diagnostic you own: if the assistant describes you wrongly, that is a content problem you can fix this week.
Two rules. Phrase every prompt the way a buyer would, not the way your marketing deck would, which usually means shorter and less flattering. And freeze the list once written, because the moment you swap prompts you have lost your baseline and your first three months of data with it.
The monthly routine, step by step
Same day every month, same order, one sitting.
- Open a fresh session in each assistant you track. Logged out where possible, no memory, no prior conversation. Two or three assistants is plenty to start; the ones your buyers actually open.
- Run every prompt twice, three times for the handful you care most about. Do not rephrase between runs. You are sampling, not negotiating.
- Score each run on a scale you can keep: 0 for not mentioned, 1 for mentioned, 2 for mentioned with a link to your site. Write the score, not your feeling about the answer.
- Note who else appears. The other names in the answer are your real competitive set for this channel, and they are often not the ones you assumed.
- Copy one full answer per family into your sheet. In three months, those excerpts will tell you more about how you are being described than any score.
- Record the conditions. Date, assistant, model version if visible, language, region, logged in or out. This is what makes next month comparable.
- Add the two easy signals. Referral sessions from assistant domains, and impressions on queries containing your brand name.
Your sheet ends with one line per month per assistant: total score, coverage rate, referrals, branded impressions. Ten minutes to read, and it will not lie to you.
What your analytics can and cannot show
Two things are genuinely visible in tools you already have.
Assistants that link out send referral traffic, and it identifies itself. In your analytics, look for referral sources like chatgpt.com, perplexity.ai, gemini.google.com and copilot.microsoft.com. That traffic is usually small and unusually well qualified, because the person arrived after reading a considered answer rather than a headline.
Search Console tells you about Google. As things stand, Google does not give you a separate performance report for its AI experiences: what happens in those surfaces is folded into your normal Search performance data, and Google documents its behaviour for site owners in its guidance on AI features and your site. So you can watch total impressions and branded queries move, but you cannot isolate an AI Overviews line, because there is not one.
Three things are invisible, and it is worth being blunt about them:
- Answers where you were named and nobody clicked. That is the majority of your exposure, and no analytics tool will ever report it. It is exactly why you run prompts by hand.
- Conversations inside the assistant. You cannot see the follow up questions, the objection, the moment your name was dropped from a shortlist.
- Why you were chosen. You will see that you appeared. The reason stays a hypothesis, which is fine as long as you label it as one.
If your CTA depends on a click, distrify is built for exactly the messy part described above: one campaign, written, published and repurposed across search, social and AI answers so the material exists to be found in the first place. Join the waitlist for early access.
Reading the numbers without fooling yourself
With 15 prompts, two runs each, your monthly score sits somewhere between 0 and 60. A move of two or three points is noise. That is not caution, it is arithmetic: with that many samples, run to run variance alone will move your total by a couple of points in either direction.
So read it like this:
- Direction over three months, not month over month. Two rises in a row is a signal. One rise is weather.
- Coverage before intensity. Being named in six prompts out of fifteen is worth more than being named twice in the same prompt, because it means the association generalised beyond one question.
- Read the wording, not just the score. Being described as a content generator when you publish is a positioning failure that a rising score will happily hide from you.
- Compare yourself to your own last month. Comparing your score to another brand's is meaningless unless they ran your list, your way, on your date.
One discipline saves the whole exercise: if you change the list, the assistants or the scoring, start a new sheet and say so. A baseline you quietly edited is worse than no baseline at all.
What to do with a month of bad results
Zeros in the first months are the normal outcome for a young site, not a verdict. The work that changes them is the same work that earns organic rankings, with one addition.
Three levers, in order of how fast they move:
- Be describable. If nothing on your site states plainly what you do, for whom, and what makes you different, an assistant has nothing to reuse. This is the cheapest fix on the list and it is usually the missing one. The mechanics are in our guide to getting mentioned in ChatGPT.
- Cover the questions. Answers get built from material that exists. Pages that answer real buyer questions in a way that can be quoted are what generative engine optimization is about, and there is no shortcut around having them.
- Get mentioned elsewhere. Public analyses of AI answers keep pointing the same way: brands that get talked about across the web get named more often in assistant answers, and that is not something you can write on your own site. Consistent presence off site is what a real distribution routine is for, and turning one article into a week of posts is the cheapest version of it, which we broke down in nine social formats.
What not to do is stuff your pages with claims about being the best option, hoping an assistant repeats them. That is the behaviour that gets content classified as made for ranking rather than for people, a line we covered in detail on the supposed AI content penalty.
What tracking will never tell you
Measurement earns its keep by having limits you can state out loud.
It will not tell you causation. You published, you were mentioned, and both of those can be true while the actual cause was a forum thread you never saw. Log your hypothesis, keep the log, and let three months of it be the argument.
It will not give you control. Nobody can guarantee that an assistant will cite you, and any vendor who does is selling you a feeling. You can make yourself easier to cite. That is where the influence ends.
It will not stay stable. A model update can move every number on your sheet in a week, in either direction, for reasons that have nothing to do with you. This is why the conditions column exists.
And it will not substitute for revenue. A rising visibility score with flat sales is a hypothesis that needs testing, not a result to celebrate. Keep the score next to something that pays, or drop the score.
How much time this takes
Forty five minutes to build the list once. Thirty to fifty minutes a month to run it, depending on how many assistants you keep. Ten minutes to read the sheet.
Tools shorten the running part, and that is genuinely what they sell: frequency and scale, not truth. They run more prompts more often than you would by hand. They do not remove the variance, they average over it, and they cannot tell you why you appeared either.
For a small team, the honest sequence is to run it by hand for a quarter first. You learn what your buyers are actually being told, you find the positioning gaps nobody would have flagged, and only then do you know whether paying to automate the sampling is worth it.
FAQ
How often should I run AI visibility tracking?
Monthly for almost everyone. Weekly only if you publish something every day, because otherwise you are sampling the assistants' variance rather than your own progress. Quarterly is too slow to catch a positioning problem while it is cheap to fix.
Do I need a tool for this?
No. A spreadsheet, a fixed prompt list and one recurring calendar slot cover the whole method described here. A tool buys you frequency and scale, which matter once you have several markets or languages, and it still cannot tell you why you were named.
Why do I get a different answer every time I ask the same question?
Because assistants are not deterministic, and because personalization, region, language, model version and whether the assistant browsed the web all change the source list. That is why the routine samples several runs and records the conditions instead of trusting one screen.
Can I see AI Overviews performance separately in Search Console?
Not as a dedicated report. Activity in Google's AI experiences is included in your overall Search performance data rather than broken out, so watch total impressions and branded queries for direction, and keep your own prompt log for what happens inside the answers.
Does being cited in an answer actually bring visitors?
Sometimes, and less than you would hope per citation. Many answers satisfy the question without a click. Treat citations as brand presence measured by your prompt log, and treat clicks as a separate line coming from assistant referrals in your analytics.
Conclusion
AI visibility tracking is not a dashboard you buy, it is a habit you keep: one frozen list of buyer questions, several runs, four signals, and a note of the conditions each time. Done that way it answers the only question worth asking, which is whether you are becoming a name that assistants reach for, in a direction you can see over a quarter.
Done any other way it produces confident nonsense. One check, one screenshot, one conclusion is how teams end up celebrating variance and steering by it.
The part that actually decides your numbers is upstream of the measurement: publishing material worth reusing, every week, on the channels where you can be found. That is the load distrify is designed to carry, and you can browse the rest of the blog meanwhile. Join the waitlist to be told when it opens.