GEO fundamentals2026
How to measure AI citations without analytics
AI citations cannot be read off your analytics, because most of them produce no click at all. Measuring them means asking the engines a fixed set of questions on a schedule and recording who they name.
Every team that takes AI visibility seriously hits the same wall in the second week. Someone asks what the number is, and there is no number.
There is a workable method. It looks less like analytics and more like a standing experiment, which is why teams built around dashboards tend not to find it on their own.
01Why your analytics cannot see this
In short
A citation that answered the question outright produces no click, so it leaves no trace in any tool that counts sessions.
Your analytics measures visits. A citation is a mention inside a generated answer, and the whole design of that answer is to resolve the question without a visit. When it works best for the reader, it is most invisible to you.
Filtering referral traffic for the assistant domains is still worth doing. It captures the subset who followed a link, and those visitors tend to arrive with unusually clear intent because their question has already been answered and the click is a verification step.
But that subset is a floor, not a total, and the gap between the two is unknown and varies by query type. Reporting referral sessions as though they were citations is the most common measurement error in this space, and it consistently understates the effect.
02What a citation actually is, for measurement purposes
In short
A named mention of your brand or a link to your domain inside a generated answer, which is deliberately broader than a clickable link.
Decide this before you start counting, because the definition determines the number and a definition that shifts mid-quarter makes the trend meaningless.
A workable definition has three cases. A linked citation, where your URL appears in the sources. A named mention, where the answer says your brand name without linking. And an unattributed lift, where the answer clearly uses your framing without naming you at all.
Count the first two. The third is real and genuinely valuable, but it cannot be detected reliably, and a metric you cannot compute consistently is worse than one you leave out and mention in the notes.
03How do you build a prompt set?
In short
Write the questions a buyer would actually type at each stage, keep them unbranded, and then freeze the list.
Start from real buyer language rather than keyword tools. The phrasing people use with an assistant is longer and more conversational than what they type into a search box, and it usually carries their constraints: the sector, the team size, the region.
Cover three intents. Discovery questions, where the buyer is asking what exists. Comparison questions, where they are asking which option fits their situation. And problem questions, where they describe a symptom without knowing the category name.
Keep branded queries in a separate list. An engine naming you when asked about you proves almost nothing, and mixing those results into the same figure will flatter it until nobody trusts it.
Then freeze the set. This is the discipline that makes the whole thing work, and the one most teams break first. Every prompt you add or reword changes the baseline, so improvements and regressions become indistinguishable from edits. When you want new coverage, open a second list and track it separately.
04What to record on each run
In short
The prompt, the engine, the date, whether you were named, and every competitor named alongside you.
That last column is the one teams leave out and later wish they had. Your own citation rate tells you how you are doing. The competitor set tells you who is winning the slots you are not, and which of them to go and read.
Record the raw answer text too. It costs nothing to store and it is the only way to answer the question that always follows a bad quarter, which is what the answers actually said about the category.
Because generated answers vary between runs, repeat each prompt several times per cycle. Report the share of runs in which you were named. A single check is an anecdote; a rate across repeated runs is stable enough to trend even though no individual run is.
This is the method behind our own tracker: a fixed per-project prompt set replayed against each engine on a schedule, with every response parsed for brand and domain mentions to produce a rate over time (Osoro Solutions, 2026).
05The metrics worth reporting
In short
Citation rate per engine, share of voice against the competitor set, and the trend on both, with the raw counts kept underneath.
Citation rate is the share of runs in which you were named, computed per engine and never averaged across engines that behave differently. A blended figure hides the case that matters, which is being strong in one engine and absent from another.
Share of voice is your citations as a proportion of all brand mentions across the set. It is the number that survives contact with a leadership team, because it answers the comparative question directly.
Report both as a trend. The absolute level depends on how hard your prompt set is, which makes it close to meaningless in isolation and genuinely useful over time.
06What this method cannot tell you
In short
It cannot measure shortlist influence, which is probably the larger commercial effect and has no trackable touchpoint at all.
Say this out loud in the reporting, because it is the honest limit of the method and someone will find it eventually.
When a buyer asks an assistant to shortlist options and your brand is named, that shapes a decision that may not surface for months and will arrive attributed to something else entirely, or to nothing. No sampling method reaches it.
The method also cannot distinguish a citation shown to thousands of people from one shown to nobody, because the engines do not publish query volumes. Treat every prompt as equally weighted and be aware that they are not.
None of this makes the measurement worthless. It makes it directional, which is enough to decide what to work on next, and considerably better than the alternative of asserting that the work is going well.
07How often should you run it?
In short
Weekly for a tracked set, and hold the reporting cadence to monthly so you are reading a trend rather than reacting to noise.
Weekly sampling gives enough data points that a real change separates from run-to-run variance within a month or so. Daily adds cost and noise without adding signal, given how slowly the underlying signals move.
Report monthly. Technical and content fixes surface on the engines' crawl cadence, so a change you shipped on Tuesday will not show up by Friday and reading the chart that way will have you undoing work that was fine (Google Search Central, 2025).
Set the baseline before you change anything. The most common regret here is starting the measurement after the remediation, which leaves you with a number and no way to say whether it improved.
Frequently asked
Can you see AI citations in Google Analytics?
Only the fraction that produced a click, which understates the total by an unknown margin. Filtering referrals for the assistant domains captures visitors who followed a link, but a citation that answered the question outright generates no session at all. Treat referral data as a floor on the effect, never as the measurement.
How many prompts should a tracking set contain?
Twenty to fifty is enough for most sites, and stability matters far more than size. The purpose is comparison across time, so a set you keep unchanged for six months tells you more than a larger set you revise every month. Add prompts to a second, clearly separated list rather than editing the baseline.
Why do the same prompts give different answers each time?
Generated answers are probabilistic, so the same question can surface different sources on different runs. This is why a single check is an anecdote rather than a measurement. Run each prompt several times per cycle and report the share of runs in which you were named, which is stable enough to trend even though any individual run is not.
Should you track branded queries?
Track them separately, and do not let them into your headline number. An engine naming you in response to a question about your own product proves very little, and mixing those results into the same rate as unbranded discovery questions will flatter the figure until it is useless. The unbranded set is where the commercial answer lives.
Which engines are worth tracking?
Start with the ones your buyers actually use, and prefer engines that show their sources, because visible citations let you diagnose as well as measure. Whichever you choose, keep the list fixed alongside the prompt set. Adding an engine mid-quarter changes your aggregate number for reasons that have nothing to do with your site.
Sources
Every figure on this page traces to one of the following. Methodology and sample are stated so you can judge the evidence rather than take it on trust.
- 01
Google Search Central, 2025Institutional analysis
Publisher documentation describing how Google surfaces indexed web content inside AI Overviews and AI Mode, including the indexing requirement and the preview controls available to site owners. Reviewed July 2026.
- 02
Osoro GEO citation sampling method
Osoro Solutions, 2026Institutional analysis
First-party measurement approach used by the Osoro GEO rank tracker. A fixed per-project prompt set is replayed against each tracked answer engine on a schedule, and every response is parsed for brand and domain mentions to produce a citation rate per engine over time. Presented here as practitioner judgement, not third-party research.
See where your site actually stands
Reading about citation mechanics is one thing. Osoro GEO scores your site against the four signals in about 30 seconds, with no account and no card.
Keep reading
GEO fundamentals
The four signals that decide whether AI cites you
Citation likelihood breaks into four dependent signals: technical, content, entity and trust. They only work in that order, and here is why.
GEO fundamentals
How AI engines choose which sources to cite
AI answer engines cite sources they can parse, extract from, identify and corroborate. Here is how that selection works and what it means for your site.
GEO fundamentals
What structured data actually does for AI citation
Schema markup is not a ranking lever. It removes ambiguity about what a page is and who published it, which is what makes a claim safe to attribute.