GEO fundamentals2026

How AI engines choose which sources to cite

AI answer engines do not cite the highest-ranking page. They cite the page they can retrieve, parse into a clean claim, attribute to a known entity, and corroborate somewhere else.

Gideon Twum6 min readLondon

Most advice about getting mentioned by ChatGPT or Perplexity treats the problem as a harder version of search ranking. It is not. Ranking and citation are two different decisions, made by two different mechanisms, and a page can pass one while failing the other.

Understanding which decision you are failing is the whole job. This article walks through how the selection actually works, then what follows from it.

01What "being cited" actually means

In short

A citation is two decisions, retrieval then extraction, and a page has to pass both before it is quoted.

When an answer engine responds to a question, it does two separable things. First it retrieves a set of candidate documents, usually with a search index behind it. Then it reads those documents and composes an answer, attaching source links to the specific claims it lifted.

That second step is the one people skip. Retrieval decides whether your page is in the room. Extraction decides whether anything you wrote survives into the answer with your name on it. A page can be retrieved on every relevant query and still never be quoted, because nothing in it could be pulled out as a clean, standalone claim.

This is why site owners sometimes see their domain in an engine's list of consulted sources but never in the answer body itself. The retrieval half is working. The extraction half is not.

02How does an AI engine pick its sources?

In short

The engine has to fetch the page, isolate a self-contained claim, attach it to a known entity, and find that claim corroborated elsewhere.

Four things have to hold, roughly in order.

The engine has to be able to fetch and parse the page. If the substance only exists after client-side rendering, or the crawler is blocked, nothing downstream matters. Google publishes explicit guidance that content must be indexable to appear in its AI features, and documents the separate controls that govern preview and training use (Google Search Central, 2025).

The engine has to be able to isolate an answer. Generated answers are assembled from spans of text, not from whole documents. A span that depends on the previous three paragraphs for its meaning cannot be lifted, so it will not be.

The engine has to know who you are. Attribution needs a stable entity, not just a hostname. When your organisation is described consistently across your own site, your structured data and the third-party places that mention you, an engine can confidently say "according to X". When those descriptions conflict, the safe move is to quote someone else.

The engine has to find the claim corroborated. Models are tuned to avoid asserting things on the word of a single unfamiliar source. A claim that appears only on your site, in your words, is a claim an engine will hedge on or drop.

03The four signals that decide whether you get quoted

In short

Technical parseability, content extractability, entity clarity and third-party corroboration, applied in that order.

We group these into four pillars, which we call the CITE framework (Osoro Solutions, 2026). The grouping is ours, but the underlying mechanics are not controversial.

Technical: can the engine read you at all

Crawlability, server-rendered content, clean HTML, correct canonicals, and structured data that matches what is visible on the page. Schema.org types such as Article, Organization and FAQPage let you state authorship, publisher and publication date without the engine having to infer them from your layout (Schema.org, 2025).

This pillar is binary in effect. Fail it and the other three are unreachable.

Content: can it lift a clean answer

Write the answer before the build-up. Put the direct claim in the first sentence under each heading, then qualify it. Phrase some headings as the questions people actually ask, because those headings are what a retrieval system matches against.

Concretely: a section that opens with "There are several factors to consider when evaluating this" gives an engine nothing. A section that opens with "Schema markup does not improve ranking directly; it removes ambiguity about what the page is" gives it a sentence it can quote and attribute.

Entity: does it know who you are

The same organisation name, the same description, the same founder, the same location, everywhere you appear. Structured data on your own site, then consistency across the directories, profiles and listings that describe you. This pillar is slow to move and hard to fake, which is exactly why it carries weight.

Trust: does anyone it already trusts mention you

Co-citation from sources the engine already treats as reliable. Being referenced in an industry roundup, quoted in a publication, or discussed in a forum thread does more for citation likelihood than another page on your own domain saying the same thing louder.

04Why ranking first does not guarantee a citation

In short

Ranking rewards relevance for a query while citation rewards extractability for a claim, and an answer is assembled from several sources rather than awarded to a winner.

The two systems reward overlapping but distinct things. Ranking rewards relevance and authority for a query. Citation rewards extractability and attributability for a claim.

A comprehensive guide that ranks well can be a poor citation candidate if its useful assertions are buried mid-paragraph, hedged, or split across sections. A narrower page that states one thing clearly and attributes it properly can be quoted repeatedly while ranking nowhere in particular.

There is a second reason. Answer engines assemble a response from several sources at once, so the question is not "who is best" but "who supplies this particular sentence". A page that owns one specific claim can win that slot against a page that covers the whole topic vaguely.

05What to change first

In short

Fix retrieval, then extraction, then identity, then trust, because a later fix cannot compensate for an earlier failure.

Work in the order the mechanism implies, because later fixes cannot compensate for earlier failures.

Start with retrieval. Confirm your pages are server-rendered, indexable, and not blocked from the crawlers whose answers you want to appear in. This is usually a short list of concrete defects rather than a strategy.

Then fix extraction. Take your highest-intent pages and rewrite the opening sentence of every section so it answers the section heading on its own. Add a short direct answer near the top of the page. Convert the questions you get from customers into headings, and answer them in the first sentence underneath.

Then work on identity. Publish Organization structured data, make sure your description matches everywhere it appears, and correct the places where an old name or an old positioning is still live.

Only then invest in trust building, which is the slowest lever and the one least under your direct control.

06How long does it take to see a change?

In short

Technical and content fixes surface within a crawl cycle, entity consistency takes as long as third parties take to update, and trust accumulates over months.

Technical and content fixes surface on the engine's crawl cadence, so days to weeks. Entity consistency takes as long as third-party sites take to update. Trust signals accumulate over months.

The practical consequence is that the first month of GEO work should be almost entirely defect removal, not content production. Publishing more pages into a site an engine cannot parse cleanly produces more pages that do not get cited.

Frequently asked

What is the difference between SEO and GEO?

SEO optimises for a ranked list of links, GEO optimises for being quoted inside a generated answer. The mechanics diverge because a ranking system returns your page and lets the reader judge it, while an answer engine reads your page, extracts a claim from it, and puts its own credibility behind that claim. That second step rewards clean structure and verifiable attribution far more heavily than it rewards keyword coverage or link volume.

Do AI engines only cite pages that rank on the first page of Google?

No. Retrieval for a generated answer is a separate process from the classic ranked results, and engines routinely quote pages that would not appear near the top of a conventional search. Ranking well correlates with being retrieved, because both reward crawlable, relevant content, but the correlation is loose enough that plenty of first-page results are never quoted and plenty of quoted pages sit well down the results.

Does structured data help you get cited by AI?

Structured data helps mainly by removing ambiguity about what a page is and who published it. Schema.org markup states the article type, the author, the publisher and the publication date in a form that does not depend on parsing your layout correctly. It is not a ranking lever on its own, and marking up content that is not visible on the page is a violation of Google's guidelines rather than a shortcut.

How long does it take for changes to affect AI citations?

Parsing and extraction fixes can show up within a crawl cycle, typically days to a few weeks. Identity and trust signals move on a much slower clock, because they depend on other sites and directories updating their descriptions of you. That is why the sensible order of work is technical first, content structure second, and authority building as a continuous background effort rather than a launch task.

Should you block AI crawlers from your site?

Blocking AI crawlers removes you from the answers they generate, which is a real cost if those answers are where your buyers now start. The decision is genuinely a trade-off between protecting content from training use and remaining visible in the surfaces that increasingly sit in front of search. Google documents separate controls for training use and for appearing in AI features, so the choice is not all-or-nothing.

Sources

Every figure on this page traces to one of the following. Methodology and sample are stated so you can judge the evidence rather than take it on trust.

  1. 01

    AI features and your website

    Google Search Central, 2025Institutional analysis

    Publisher documentation describing how Google surfaces web content inside AI Overviews and AI Mode, including indexing requirements and the preview controls available to site owners. Reviewed July 2026.

  2. 02

    Schema.org vocabulary

    Schema.org, 2025Institutional analysis

    Open structured-data vocabulary maintained collaboratively by Google, Microsoft, Yahoo and Yandex, published with versioned releases and type definitions. Reviewed July 2026.

  3. 03

    The CITE framework

    Osoro Solutions, 2026Institutional analysis

    First-party audit framework applied to every site scanned on Osoro GEO. Four pillars, Technical, Content, Entity and Trust, scored per scan from crawl output and answer-engine response sampling. Presented here as practitioner judgement, not third-party research.

See where your site actually stands

Reading about citation mechanics is one thing. Osoro GEO scores your site against the four signals in about 30 seconds, with no account and no card.

Keep reading