How do I get cited by Perplexity?
- How do I get cited by Perplexity?
- Why does Perplexity never cite my site?
- Should I block PerplexityBot in robots.txt?
- Does Perplexity respect robots.txt?
- How do I see which sites Perplexity cites for my keywords?
- Is Perplexity traffic even worth chasing?
Perplexity cites pages it can retrieve that directly answer the question asked, so the work is mostly conventional: rank in Google and Bing for the full question, allow PerplexityBot in robots.txt, and put the answer near the top of the page in a form that survives being lifted out. Because Perplexity numbers every citation, you can test all of this yourself and watch who gets picked.
On this page
- Why Perplexity is the engine worth testing against
- PerplexityBot and Perplexity-User are not the same thing
- The awkward question about how Perplexity retrieves
- So the work is mostly conventional, with one twist
- A citation self-test you can run this afternoon
- Perplexity is a stability signal, not a canary
- The Publishers' Program, and the number you should not repeat
- How much Perplexity traffic is actually worth
- Your Perplexity checklist
- Frequently asked questions
- Perplexity documents two web agents with opposite behavior: PerplexityBot respects robots.txt, while Perplexity-User generally ignores it because a human triggered the fetch — so robots.txt alone cannot fully opt you out.
- PerplexityBot is not used for model training. It exists to surface and link sites in Perplexity results, which makes blocking it a pure visibility loss with no training-data upside.
- One operator analysis argues Perplexity answers mainly from scraped Google and Bing top-ten results rather than its own crawl. It is unconfirmed, but under either model, ranking in conventional search is the prerequisite.
- Numbered citations make Perplexity the cheapest AI-visibility audit available: run ten buying-intent questions monthly and record which domains get cited.
- Semrush's 13-week study found Perplexity's citation mix stayed comparatively stable while ChatGPT's swung violently, making Perplexity a good read on durable content quality and a poor early warning of platform change.
You ran one of your own buying-intent questions through Perplexity. It produced a confident four-paragraph answer with nine numbered citations, and not one of them was you. Two were competitors. One was a Reddit thread.
That is more useful than it feels, because Perplexity shows its working. Every claim carries a footnote and the sources panel lists the domains in order, so you can see who won and read the page that beat you. No other major answer engine makes the scoreboard this legible, which is why it is the right place to run your first honest AI-visibility test.
Why Perplexity is the engine worth testing against
Perplexity's interface is built around attribution: numbered inline citations plus a visible list of the pages it pulled from. Every query becomes a small, repeatable experiment: ask what your customer asks, and the engine names the pages it currently treats as the best answers.
Two caveats keep this honest. Results vary between runs, accounts, and Perplexity's quick versus deeper research modes, so one run proves little. And a win on one phrasing is not topic ownership: the Semrush and Kevin Indig topic-authority study (Jan–Jun 2026, 1,094 categories, 600,000+ citations) found narrow single-query wins reverse frequently, and that real ownership means appearing in four of five related prompts. That study measured ChatGPT, but the caution generalizes. Test ten questions, not one.
PerplexityBot and Perplexity-User are not the same thing
Perplexity documents two web agents that behave in opposite ways. It is the most consequential thing about the platform, and the part most robots.txt advice gets wrong.
PerplexityBot | Perplexity-User | |
|---|---|---|
| What triggers it | Automated crawling, on Perplexity's schedule | A user query that needs a page fetched right now |
| Stated purpose | Designed to surface and link websites in Perplexity search results | Live retrieval of a page a user's question requires |
| Used for model training | No — not used to crawl content for AI foundation models | No — the documentation says the same |
| Respects robots.txt | Yes | No — Perplexity's docs say it generally ignores robots.txt because a user requested the fetch |
| Your actual lever | robots.txt, honored | Server-side controls: authentication, paywall, IP rules. Not robots.txt |
The reasoning is the standard industry argument: robots.txt governs automated crawling, and a fetch a human explicitly asked for looks more like a browser loading a page. OpenAI says the same about ChatGPT-User, Meta about Meta-ExternalFetcher. You are free to find it unconvincing. It does not change the operational fact, stated here in Perplexity's own words.
Since a user requested the fetch, this fetcher generally ignores robots.txt rules.
Perplexity bot documentation, on Perplexity-User
For most sites the right configuration is simple: let both through, explicitly, so nobody blocks them by accident during a WAF cleanup.
# Perplexity's indexing crawler. It is not used for model training;
# blocking it only removes you from Perplexity's results.
User-agent: PerplexityBot
Allow: /
Disallow: /cart/
Disallow: /account/
# Live, user-triggered fetches. Perplexity documents that this agent
# generally ignores robots.txt, so treat anything you write here as a
# statement of intent, not an enforcement mechanism.
User-agent: Perplexity-User
Allow: /
Sitemap: https://example.com/sitemap.xml
The awkward question about how Perplexity retrieves
There is a live argument about whether PerplexityBot matters much for real-time answering. An operator analysis published by Primary Position argues it is close to a red herring: on that account, when you ask a question Perplexity queries Google and Bing search APIs, scrapes the top five to ten results, and synthesizes from those pages rather than from an index its own crawler built.
Treat that as third-party inference from observed behavior, not documentation. Perplexity has not published a retrieval architecture confirming or contradicting it. The claim could be right, partly right, or already out of date.
It barely changes what you do either way. If Perplexity answers from its own crawl, you need PerplexityBot allowed and pages worth indexing. If it answers from scraped Google and Bing results, you need the top ten. Both roads run through conventional search. Ranking badly and being cited well is not a stable position.
So the work is mostly conventional, with one twist
The prerequisite is ranking. Take the ten questions your buyers actually ask — the full-sentence questions, not head terms — and check where you sit in Google and Bing for each. If you are not on page one, that is the project.
The twist is what gets extracted once you are eligible. The only academic work testing content strategies against a live Perplexity system is the Princeton and IIT-Delhi GEO paper (arXiv:2311.09735, KDD '24, August 2024). It reported roughly 41% improvement from adding quotations, 30–40% from adding statistics and citing sources, and 15–30% from fluency optimization, with keyword stuffing neutral to negative. Date it honestly. The work is about two years old, it tested GPT-3.5-turbo and one live Perplexity system, and the authors cautioned that results may need to adapt. Directional evidence, not a playbook.
The direction is consistent with what you can observe in Perplexity's citations, and the moves are cheap:
- Answer the question in the first hundred words, in a form that stays correct when lifted out of the page.
- Use specific numbers, dates, and named sources instead of adjectives. A sentence carrying a figure and an attribution is more quotable.
- Quote primary sources directly and link them. Perplexity is synthesizing from pages — give it clean material.
- Keep one page per question rather than one page covering nine questions.
- Write the URL slug as the natural-language question. Ahrefs' Feb 2025 analysis of 1.4 million ChatGPT prompts found natural-language slugs cited 89.78% of the time versus 81.11% otherwise — ChatGPT data, but it costs nothing to apply.
A citation self-test you can run this afternoon
The routine takes about an hour the first time, twenty minutes a month after that.
- Write ten questions with real buying intent — what a prospect types the week before they buy. Not “best CRM” but “which CRM works for a two-person agency that bills hourly”.
- Run each in a logged-out browser session, in Perplexity's default search mode. No account history, no personalization.
- Record four things per question: whether a page of yours was cited, which competitors were cited and at what number, the full list of cited domains, and which URL of yours was pulled.
- Run the same ten in Perplexity's deeper research mode and record the same four things. The source sets often differ.
- Check the top ten Google and Bing organic results for the same questions, and note the overlap with Perplexity's cited domains.
- Repeat on the same date each month, with identical phrasing, and keep the sheet.
The overlap in step five makes the test diagnostic. Three patterns, three different problems:
- Cited competitors all rank on page one and you do not. A ranking problem. Fix search first; nothing else works until you do.
- You rank on page one but are never cited. A content-shape problem. The page probably buries the answer, hedges it, or spreads it across four sections. Rewrite the top to answer outright.
- The cited sources are third-party — roundups, review sites, forum threads, trade press — rather than vendor sites at all. An earned-media problem. One analysis found roughly 84% of AI citations come from earned media rather than brand-owned pages; an aggregator figure, so directional only. Either way, get covered where the engine already looks.
Running this by hand across ten questions is reasonable. Running it across two hundred questions and four engines monthly is what GetFound3 automates.
Perplexity is a stability signal, not a canary
Semrush's citation study — 230,000+ prompts, 100M+ citations, 13 weeks from 14 July to 12 October 2025, across ChatGPT search, Google AI Mode, and Perplexity — found that Perplexity's citation mix stayed comparatively stable. ChatGPT's did not. Reddit's share of its citations fell from around 60% to around 10% and Wikipedia's from around 55% to under 20%, both around mid-September 2025.
That asymmetry has a practical reading. Because Perplexity's source mix moves slowly, losing citations there usually means something changed on your side: you slipped in search, a competitor published something better, the page went stale. It is a clean read on durable content quality and a poor early warning system. To spot an engine quietly rewriting its sourcing policy, watch ChatGPT, where the swings are big enough to see.
The Publishers' Program, and the number you should not repeat
Perplexity runs a Publishers' Program, announced with an initial cohort of TIME, Der Spiegel, Fortune, Entrepreneur, The Texas Tribune, and WordPress.com. It has three components:
- A share of advertising revenue tied to citation-driven interactions, with citation analytics provided through ScalePost.ai.
- Free access to Perplexity's Online LLM APIs.
- A free one-year Enterprise Pro subscription for the partner's employees.
For almost every business reading this, the program is context rather than strategy: it targets publishers with large audiences and has no self-serve path. Know it exists, note that Perplexity has a commercial relationship with some publishers it cites, and move on.
How much Perplexity traffic is actually worth
Two things. The scale figures come from an aggregator, not Perplexity's own reporting. More importantly, the referral share is falling: StatCounter's March 2026 data puts Perplexity at 7.07% of AI referral traffic, down more than 40% from its 12.07% peak, with Gemini overtaking it at 8.65% and ChatGPT still on 78.16%. Measured by usage instead of referrals, First Page Sage's July 2026 US report puts Perplexity at 2.0% — a different methodology and a much smaller number, worth reporting alongside rather than choosing between.
The base rate matters more than either figure. Chartbeat and Axios data published in March 2026 found chatbot referrals still account for under 1% of total publisher pageview referrals, even after ChatGPT referral traffic grew more than 200% during 2025. Perplexity is not a traffic channel yet. It is the best instrument available for measuring whether your content is winning.
Your Perplexity checklist
- Confirm no
robots.txt, CDN, or WAF rule blocks PerplexityBot. Check server logs, not just the file. - Accept that Perplexity-User fetches pages regardless of robots.txt. If a page must stay out of AI answers, put it behind authentication.
- List your ten highest-intent buyer questions and check Google and Bing rankings. Anything off page one is a ranking project first.
- Run the ten-question self-test in a logged-out session, in both Perplexity modes, and record every cited domain.
- Classify each miss as a ranking, content-shape, or earned-media problem, then fix the largest bucket first.
- Rewrite the top of your two weakest pages to answer the question in the first hundred words, with a number and a named source.
- Diary the same test for the same date next month. Perplexity's stability is what makes month-over-month comparison meaningful.
Frequently asked questions
If I block PerplexityBot, will my site disappear from Perplexity answers?
Blocking PerplexityBot removes you from the crawl that surfaces and links sites in Perplexity results, so it costs you visibility. It does not stop Perplexity-User, the live agent that fetches a page because a person asked a question needing it. Perplexity's own documentation says that agent generally ignores robots.txt, so a block gives you the downside without full exclusion.
Does Perplexity use my content to train its models?
Not according to its documentation. Perplexity states that PerplexityBot is not used to crawl content for AI foundation models, and says the same of Perplexity-User. Its stated job is surfacing and linking websites in Perplexity search results. That is a real difference from crawlers like GPTBot or ClaudeBot, whose documented purpose includes training, and it means blocking PerplexityBot buys you no training-data protection — only lost visibility.
Why does Perplexity cite a review site instead of my product page?
Because the review site answered the comparative question and your product page did not. Perplexity is assembling an answer, and vendor pages rarely contain neutral comparisons, specific numbers, or named alternatives. If third-party pages dominate the citations in your category, that is an earned-media signal: you need coverage on the pages the engine already trusts, not just a better product page.
Do I need an llms.txt file to get cited by Perplexity?
No. Ahrefs studied 137,210 domains using May 2026 data and found 28% published a valid llms.txt file, but 97% of those files received zero requests during the study period. Ahrefs concluded that for showing up in ChatGPT, Perplexity, or AI Overviews, an llms.txt file is largely decoration. Spend the time on rankings and page structure instead.
How often should I re-run the citation test?
Monthly is the right cadence for Perplexity. Semrush's 13-week study found its citation mix stayed comparatively stable, so month-over-month comparisons are meaningful rather than noisy. Weekly testing mostly measures run-to-run variance. Keep the phrasing identical between runs, use a logged-out session, and record the full cited-domain list rather than only whether you appeared.
Can a small business join the Perplexity Publishers' Program?
Realistically, no. The announced cohort was publishers with large audiences: TIME, Der Spiegel, Fortune, Entrepreneur, The Texas Tribune, and WordPress.com. There is no self-serve application path that makes it a practical lever for a mid-market brand. Treat the program as useful context about Perplexity's commercial incentives rather than as a channel you can plan around.
Sources
- Perplexity — PerplexityBot and Perplexity-User documentation
- Perplexity — Introducing the Perplexity Publishers' Program
- StatCounter — Google Gemini overtakes Perplexity (March 2026)
- Semrush — Most-cited domains in AI (230,000+ prompts, 13 weeks)
- Primary Position — operator analysis of Perplexity's crawl and index
- GEO: Generative Engine Optimization (arXiv:2311.09735, KDD '24)
- Ahrefs — Why ChatGPT cites one page over another (1.4M prompts)
- First Page Sage — Top generative AI chatbots by usage share