Playbook2026

How Perplexity picks the sources it cites

Perplexity picks its sources by running a live search, retrieving a set of pages, and assigning numbered citations to the ones it quotes as it writes the answer, favouring pages that are relevant, recent, and from sources it treats as authoritative.

6 min readLondon

Perplexity is the answer engine that shows its work. Every answer comes with numbered citations you can click, which makes it the easiest of the major engines to reason about: you can see exactly which sources it chose. The useful question is why it chose those and not yours.

The mechanics are more knowable here than with most engines, partly because Perplexity documents its crawlers and partly because the citations themselves are visible evidence. This is how the selection works, and what to do about it.

01What is different about how Perplexity works?

In short

Perplexity retrieves live pages and quotes them with numbered citations, so the goal is to be a retrieved and quoted source rather than a ranked result.

A classic search engine returns links and lets you read. Perplexity runs a search, reads a small set of pages itself, and writes an answer that cites the specific pages it drew claims from. Those numbered citations are not decoration: they mark the sources whose sentences the answer is built on.

That changes what you are optimising for. Position on a results page is not the target, because Perplexity is not showing a results page. Being one of the handful of sources it retrieves and then finding your sentence clear enough to quote is the target. The visible citations make this unusually measurable, since you can read any answer and see whether you were named.

02How does Perplexity find and read your pages?

In short

It uses PerplexityBot to build its search index and Perplexity-User to fetch pages live during a query, and both are governed by robots.txt.

Perplexity documents two user-agents. PerplexityBot builds and maintains the search index over time, and Perplexity-User fetches pages live on behalf of a person during an active question (Perplexity, 2026). The distinction matters for the same reason it does with other engines: one keeps a standing index of you, the other reads you in the moment a user asks.

Access is governed by robots.txt, which the Robots Exclusion Protocol defines and compliant crawlers are expected to honour (IETF, 2022). Perplexity states that if you disallow PerplexityBot, it will not index the text of your pages, though it may still show your domain, your headline, and a brief factual summary (Perplexity, 2026). So a block is not a clean disappearance: it removes your quotable text while leaving a thin reference behind.

The practical first step is to confirm you are actually retrievable. Your important pages need to be reachable, allowed in robots.txt for the Perplexity agents you want, and served as real content in the HTML rather than assembled only after JavaScript runs. A page a crawler cannot read is a page that cannot be cited.

03What does Perplexity look for in a source?

In short

Relevance to the exact question, recency, and the authority of the source, which practitioners see consistently even where it is not spelled out in documentation.

Three signals show up again and again in which pages get cited. The first is relevance to the specific sub-question, not the broad topic: Perplexity is filling a particular slot in its answer, and it wants the page that addresses that slot directly. The second is recency, which matters most for questions whose answers change over time, where a current page beats a stale one making the same claim. The third is source authority, where a claim carried by a source Perplexity treats as reputable is a safer thing for it to repeat.

It is worth being honest about which of these is documented and which is observed. Perplexity documents its crawlers and its robots.txt behaviour. The weighting of relevance, recency, and authority is what practitioners infer from watching which pages get cited across many answers, and it is consistent enough to plan around, but treat it as informed observation rather than a published ranking formula (Osoro Solutions, 2026).

04How do you become the source it cites?

In short

Be retrievable, answer the exact question in a liftable sentence, keep the page current, and earn standing on sources Perplexity already trusts.

Start with retrieval, because nothing else matters if you are not in the set. Confirm indexing and crawler access as above.

Then make the page answerable. Because citations are assigned while the answer is written, Perplexity needs a clean, self-contained sentence it can lift for the question at hand. Open each section with the direct answer to its heading, phrase the heading the way a person asks the question, and keep the claim specific enough to be worth quoting. A page that buries its answer in a long build-up gives Perplexity nothing to cite even when it retrieved the page.

Then keep it current. For any page answering a question whose answer moves, a genuine update with a visible date is a real signal, not a cosmetic one. Do not fake it, but do maintain the pages that matter.

Finally, build corroboration. Authority here is partly about being referenced and discussed on the sources Perplexity's retrieval respects, which is the slow off-page work of earning mentions rather than buying links. It is the same discipline that helps you across every answer engine, and it compounds.

05Should you block PerplexityBot?

In short

Only if you value restricting access more than being cited, because a block removes your quotable text while still leaving a domain-level reference.

This is a genuine decision, not an obvious one. If your priority is keeping your content out of Perplexity's index, blocking PerplexityBot does that for the text. But remember the documented behaviour: Perplexity may still surface your domain, headline, and a short summary even when blocked, so you do not vanish, you just stop being quotable (Perplexity, 2026).

For most brands that want to be recommended inside answers, blocking is a cost rather than a protection, because it removes exactly the quotable material that earns a citation. Decide it deliberately, set robots.txt to match, and do not let a blanket AI-crawler block make this choice for you by accident.

06How do you tell if it is working?

In short

Sample your target questions in Perplexity directly and read the citations, because the visible sources are the measurement.

Perplexity makes this easier than any other engine. Ask your real buyer questions, read the answer, and note which domains are cited and for which part of the answer. Do it on a schedule so you can see movement, and record when a competitor holds a slot you want. There is no analytics event for a citation, so this direct sampling is the measurement, and with Perplexity the evidence is right there in the numbered links.

07Where should you start?

In short

Confirm PerplexityBot access and indexing, rewrite your highest-intent pages to answer their own sub-questions, then keep them current and corroborated.

Work in dependency order. First the retrieval checks, because they gate everything. Then the editorial pass on the pages that matter most to your pipeline, opening each section with its answer and phrasing headings as questions. Then the ongoing work of freshness and corroboration, which is slower but is what separates a page that gets cited once from a source Perplexity returns to. Measure by reading the citations, and let what you see there tell you which question to win next.

Frequently asked

How is Perplexity different from a normal search engine?

Perplexity answers the question directly instead of returning a list of links, and it shows numbered citations to the specific pages it drew from. A normal search engine ranks pages and leaves the reading to you; Perplexity reads a handful of pages itself, composes an answer, and attributes the claims it lifted. The outcome you are optimising for is being one of those cited sources, not holding a rank position.

Which crawler does Perplexity use, and does it respect robots.txt?

Perplexity runs two user-agents: PerplexityBot, which builds and maintains its search index, and Perplexity-User, which fetches pages live during a user's query. Perplexity states that PerplexityBot will not index the text of a site that disallows it in robots.txt, though it may still list the domain, headline, and a short factual summary. Set your robots.txt to match whether you want to be indexed and quotable.

Why does Perplexity keep citing my competitor instead of me?

Usually because their page is a cleaner answer to the specific question than yours, is more recent, or sits on a source Perplexity treats as more authoritative. Because citations are assigned per sub-question while the answer is written, the win goes to whoever states that particular point most clearly, not to whoever is bigger overall. Find the exact questions where they are cited and make your page the better answer to those.

Does content freshness matter for Perplexity citations?

Yes, recency is one of the signals practitioners see Perplexity weight, especially for questions where the answer changes over time. A page with a visible, genuine update date and current information is a stronger candidate than a stale one making the same claim. Refresh the pages that answer time-sensitive questions rather than letting them drift out of date.

Should I block Perplexity to protect my content?

That depends on whether you value being cited in its answers more than restricting its access, and the two are a genuine trade-off. Blocking PerplexityBot removes your text from the index it cites from, but it may still show your domain and a brief summary. If being named as a source matters for your reach, blocking is a direct cost rather than a free protection.

Sources

Every figure on this page traces to one of the following. Methodology and sample are stated so you can judge the evidence rather than take it on trust.

  1. 01

    How does Perplexity follow robots.txt?

    Perplexity, 2026Institutional analysis

    Publisher help documentation describing Perplexity's two user-agents, PerplexityBot for building the search index and Perplexity-User for live fetches, and stating that a page disallowed in robots.txt is not indexed for its text though the domain, headline, and a brief summary may still appear. Reviewed August 2026.

  2. 02

    RFC 9309: Robots Exclusion Protocol

    IETF, 2022Institutional analysis

    The standards-track specification defining how robots.txt user-agent and disallow rules are expressed and how compliant crawlers are expected to interpret them. Reviewed August 2026.

  3. 03

    The CITE framework

    Osoro Solutions, 2026Institutional analysis

    First-party audit framework applied to every site scanned on Osoro GEO. Four pillars, Technical, Content, Entity and Trust, scored per scan from crawl output and answer-engine response sampling. Presented here as practitioner judgement, not third-party research.

This article was drafted with AI assistance, then fact-checked, edited and approved by Gideon Twum before publication. Every statistic traces to a named source listed above.

See where your site actually stands

Reading about citation mechanics is one thing. Osoro GEO scores your site against the four signals in about 30 seconds, with no account and no card.

Keep reading