Article contents
Perplexity is an answer engine: every sentence points to a numbered source
Unlike ChatGPT, which chooses whether or not to search, Perplexity searches on every question. It relies on its own index, built by the PerplexityBot crawler, reads the pages it has selected, then writes an answer in which every claim is followed by a numbered reference to its source. The number of sources per answer is higher than on the other engines, often between five and ten, and they are displayed before the text itself.
This architecture has two consequences for a brand. The first: there are more places per answer, so more opportunities to be cited, but also more competitors visible side by side. The second: the citation can be verified by the user in one click, which pushes the engine towards sources that say precisely what it makes them say. Vague pages, sales pitches and unconditional promises are retained even less here than elsewhere.
In our observations, Perplexity is used by profiles who are documenting a decision: B2B buyers, consultants, lawyers, engineers, journalists. Questions there are longer and more technical than on Google, and the answer is often read in full, sources included. Being cited there has a value that traffic volume underestimates.
Anatomy of an answer: five zones, two that count
A Perplexity answer reads like a composed page. Understanding its zones lets you know where a brand actually appears, and count citations in a way that is comparable from one week to the next.
Two zones count: the strip, because it is seen, and the numbered references, because they attest that a passage was used. A source can appear in the strip without being referenced in the text, when it was read but little used. In our records, we count a citation when the domain is referenced at least once in the text; presence in the strip alone is noted separately.
How Perplexity builds its answer: two crawlers and a step-by-step search
Perplexity documents two agents. PerplexityBot crawls the web to build and refresh the search index; according to the company, it is not used for model training, and it respects the robots.txt file. Perplexity-User opens a page at a user's request, when the answer requires reading its full content; the documentation states that, since the visit is triggered by a person, this agent may not respect robots.txt rules. Blocking PerplexityBot therefore removes you from the index; blocking Perplexity-User has no guaranteed effect.
- 01ReformulationThe question is translated into several searches, displayed on screen as they run.
- 02SearchEach search queries the index built by
PerplexityBot; results appear as cards. - 03ReadingThe pages retained are opened and read, by
Perplexity-Userif necessary. A slow or empty page at this stage is abandoned. - 04WritingThe text is composed with one numbered reference per claim.
- 05Follow-upsRelated questions are suggested, feeding the next session.
The deep search mode makes these steps visible, which turns it into a free diagnostic tool: you can see which searches were launched, which cards came out, and which were retained in the text. If your page appears in the cards but never in the references, the problem lies in the content; if it never appears in the cards, the problem lies in access or in the index. The internal re-ranking between cards and text is not documented; we describe it by its effects.
The weight of third-party sources: why Perplexity cites forums and media so often
On our panels, Perplexity is the engine where the share of third-party sources is highest on choice questions. Three reasons explain this. The engine searches systematically, so it always encounters the comparisons, discussion threads and articles that cover the subject. It cites a lot, so it has room to include these sources alongside brand sites. And it favours passages that answer in one sentence, a form that forums and comparison sites adopt naturally, while brand sites avoid it.
Perplexity also offers targeted search modes, including one oriented towards community discussions, and it has set up a programme with press publishers. Without knowing their exact effects on selection, we do see that partner media and communities are cited abundantly. For a brand, this means the question "how do I get cited by Perplexity?" splits in two: how do I get my pages cited, and how do I become present, with the right words, on the third-party pages Perplexity already cites. We address this second question in Reddit, forums and communities: why AI engines cite them.
The formats that win
The brand pages that earn numbered references on our panels have recognisable forms. We group them into four families, which often combine on a single page.
- 1Short dated answerA question as the heading, a two-sentence answer, a visible date.Otherwise: the forum that answers in one sentence takes the place.
- 2Comparison tableOptions in rows, criteria in columns, precise values, including those that do not favour you.Otherwise: the third-party comparison is cited, with its errors about you.
- 3Original dataA figure that only you publish: a measured lead time, a price, a test result, a named method.Otherwise: nothing distinguishes your passage from ten others.
- 4Definition pageA term from your trade, defined in three sentences, with an example and a nuance.Otherwise: the encyclopaedic reference takes the place.
The comparison table deserves a clarification. Many companies refuse to compare their offers with those of competitors on their own site. On Perplexity, this reluctance has a direct cost: the engine then cites a third-party comparison, often incomplete or out of date about you, and the user reads it with the same credence. An honest, dated "X or Y" page, with verifiable criteria, is one of the most profitable formats we have measured on this engine.
B2B use cases: what buyers ask and who gets cited
The B2B questions put to Perplexity fall into a few stable families, each with dominant sources and a main lever. The table summarises what we observe on our panels; the proportions vary by sector, the structure does not.
| Question type | Example | Dominant sources | Main lever |
|---|---|---|---|
| Choosing a provider | "Which agency for an ISO 27001 security audit in France?" | Directories, rankings, trade press | Presence on third-party pages, with the right vocabulary |
| Comparing two solutions | "X or Y for fleet management?" | Comparison sites, forums, vendors' "vs" pages | Honest, dated "vs" page on your site |
| Regulatory question | "E-invoicing obligations for an SME in 2026?" | Public sites, firms, software vendors | Short dated answer, updated at every change |
| Prices and terms | "How much does payroll software cost for 50 employees?" | Vendor sites, comparison sites | Published pricing, with explicit terms |
| Opinion on a named brand | "Is [Brand] reliable?" | Reviews, forums, press, the brand's site | Replies to reviews, proof page, consistent entity |
Table: scroll horizontally.
The regulatory-question case is the most favourable to B2B brand sites. Public sources are accurate but rarely written as short answers; a firm or a vendor that publishes a dated, precise answer, updated at every rule change, earns numbered references durably. This is often where we start a programme on this engine.
Illustrative data. What to read into it: the questions where the site can answer for itself progress quickly; choosing a provider depends on third-party pages and moves more slowly.
The levers for appearing there, in the order we pull them
The order matters, because each lever presupposes the previous one. It follows the four conditions of a citation, accessible, citable, identified, corroborated, detailed in How to get cited by ChatGPT, Claude and Perplexity, with the particularities of this engine.
- Check
PerplexityBot's access inrobots.txt, in the CDN rules and in the server logs. A site absent from the cards on its own subjects almost always has an access problem, not a content problem. - Serve complete, fast HTML. The reading step abandons slow or empty pages; the content must be in the server response, not built by the browser.
- Rewrite priority pages as short dated answers, starting with regulatory questions, prices and terms, where the site is the legitimate source.
- Publish one piece of original data per important page: a figure, a lead time, a measured result, with its method and its date.
- Create the "vs" pages that buyers ask for, honest and verifiable.
- Exist on the third-party pages Perplexity already cites on your choice questions: answers from identified experts in communities, presence in directories and comparison sites, press coverage using the same words as your site.
- Measure every week on a panel of questions, distinguishing presence in the strip from references in the text.
Perplexity and Claude share one trait: citations are more precise there than on ChatGPT, and the passages retained are shorter. The comparison between the two is the subject of the next article, Claude and web search: how Anthropic selects and cites its sources. For the mechanics specific to ChatGPT, see ChatGPT Search: what we know about source selection.
What to remember
- Perplexity searches on every question and numbers its sources; the citation that counts is the reference in the text, not presence in the strip alone.
- PerplexityBot builds the index and respects robots.txt; Perplexity-User reads on request and, according to the company, may not respect it.
- Third-party sources weigh more here than elsewhere on choice questions; the programme splits between your pages and those that talk about you.
- Four formats win: short dated answer, comparison table, original data, definition page.
- In B2B, start with regulatory questions, prices and terms, where the site is the legitimate source and progresses quickly.
Frequently asked questions
Does blocking PerplexityBot remove my site from Perplexity?
Yes for the index: PerplexityBot builds the search index and, according to Perplexity, respects robots.txt; a block removes your pages from the cards and from the numbered references. On the other hand, Perplexity-User, which opens a page at a user's request, may not respect robots.txt according to the company's documentation. A block therefore does not prevent a user from having your page read, but it deprives you of spontaneous citations.
How do you count a citation in Perplexity?
Count the numbered references in the text of the answer, not only the cards in the source strip. A page can appear in the strip because it was read, without any of its claims being used. In our records, a citation is counted when the domain is referenced at least once in the text; presence in the strip is noted separately, as an access indicator.
Why does Perplexity cite forums rather than my site?
Because the engine searches systematically, cites a lot and favours passages that answer in one sentence, a form that forums and comparison sites adopt naturally. On choice questions, brand sites cannot recommend themselves. The lever consists in existing, with the right words, on those third-party pages, through answers from identified experts and a presence in directories and the press, while writing your own pages as short answers.
Should I publish an "X or Y" page on my own site?
Yes, if it is honest. On Perplexity, a comparison missing from your site is replaced by a third-party comparison, often incomplete or out of date about you, and read with the same credence. A dated "vs" page, with verifiable criteria and precise values, including those that do not favour you, is one of the most profitable formats we have measured on this engine.
Which B2B questions should I work on first on Perplexity?
Regulatory questions, prices and terms. These are the ones where your site is the legitimate source and where public sources, accurate but rarely written as short answers, leave room for a dated, precise page. On our programmes, these families progress within a few weeks; choosing a provider, which depends on third-party pages, takes longer.

