Ad Hoc Digital

Perplexity, tool by tool

How does Perplexity decide which law firms to cite?

Perplexity searches the web on every question and numbers its citations. Here's how its crawlers, index and source labels work for a law firm.

Santiago Alvarez

By , Founder, Ad Hoc Digital
Last updated

The short answer

If you read one part of this page, read this.

Perplexity cites a law firm when the firm's pages are in its search index and answer the question better than the other pages it finds. It searches the web each time someone asks, writes a summary, and attaches numbered citations to the sources it used.

Perplexity documents its two user agents, how it treats robots.txt, and a new set of source labels. It doesn't document how it chooses between local businesses or how it uses a person's location, and we won't guess at that here.

For a firm, that leaves three jobs: let PerplexityBot through both robots.txt and the firewall, publish pages a careful reader would trust enough to click, and make sure the directories it cites say the right things. We work with law firms across the US and Canada, and also in Australia and the UK.

How it answers

Perplexity runs a live web search for every question and shows numbered citations for what it used.

Perplexity calls itself an answer engine. Its help center says it searches the internet in real time, summarizes what it finds, and includes numbered citations linking to the original sources in each answer.

That's a different starting point from ChatGPT, which decides whether a question needs a search. On Perplexity the search is the default, and the citations are on screen from the first answer. A person can switch web sources off entirely with the "Choose sources" control, but most people asking about a lawyer won't.

Paid users can pick which model writes the answer, and Perplexity's Research mode runs many searches for one question. The model changes the writing. The pages it can cite still come from Perplexity's search, which is the part a law firm can influence.

Because the citations are visible, Perplexity users tend to click them. In practice the cited page is the firm's first impression, so it has to read well cold, without the rest of the site around it.

Perplexity's crawlers

PerplexityBot builds the search index, and blocking it is the one setting that removes a firm.

Perplexity lists two user agents and says each robots.txt setting works independently, with changes taking up to 24 hours to register.

How Perplexity says its crawlers behave (Perplexity docs and Help Center, checked October 11, 2026)
CrawlerWhat Perplexity says it doesFollows robots.txt?What a law firm should do
PerplexityBotSurfaces and links websites in Perplexity's search results; not used to crawl for AI foundation modelsYesAllow it, and allow its published IP ranges in the firewall
Perplexity-UserVisits a page when a user's question needs it and may link it in the answer; not used for crawling or trainingGenerally no, since a person askedNothing to set; don't rely on it to put you in the index
Third-party crawler partnersHelp build Perplexity's search indexPerplexity says their agreements now require itKeep robots.txt open to search crawlers in general
How Perplexity says its crawlers behave (Perplexity docs and Help Center, checked October 11, 2026) Perplexity says it doesn't build foundation models, so content PerplexityBot indexes isn't used for AI model pre-training.

That last note settles the usual worry. The reason firms block AI crawlers is training. Perplexity says PerplexityBot isn't a training crawler, so blocking it costs visibility and protects nothing a firm usually cares about.

If you block it

A blocked site can still appear as a headline and a one-line summary, which is the worst version of you.

Perplexity's help center says PerplexityBot won't index the full or partial text of a site that disallows it, but it may still index the domain, the headline and a brief factual summary.

So blocking doesn't make a firm invisible. It makes it thin: Perplexity can still know the firm exists, without the practice pages that explain what it handles, where, and how the process works.

A competitor whose pages are readable gets cited for the substance.

Perplexity also says it partners with third-party crawlers to help build its index and has updated those agreements so the providers respect robots.txt too. A blanket disallow aimed at "AI bots" can reach further than the firm intended.

The firewall layer

Perplexity tells site owners to allow its bots in the firewall, not just in robots.txt.

Perplexity's crawler page includes step-by-step rules for Cloudflare and AWS firewalls, because a bot-blocking setting can stop PerplexityBot even when robots.txt allows it.

  1. Check robots.txt

    Look for a rule naming PerplexityBot, or a catch-all that disallows every user agent. Either one keeps the firm's text out of the index.

  2. Check the CDN or firewall

    Perplexity's Cloudflare instructions: in Security, then WAF, create a custom rule where the user agent contains PerplexityBot or Perplexity-User and the source IP is in Perplexity's published ranges, with the action set to Allow. Its AWS instructions do the same with IP sets and string matches.

  3. Keep the IP ranges current

    Perplexity publishes its ranges as JSON files and says they change, so a rule built on last year's list can quietly start blocking. It recommends refreshing them automatically.

  4. Confirm with a live request

    When we audit a client site, we check what each AI crawler actually receives from the live site, not only what robots.txt says. On client domains where we manage the CDN, AI search and agent crawlers are set to allowed.

Source labels

Perplexity now marks some cited sites as Government, Academic or Trusted, and says money doesn't buy a label.

When Perplexity cites a source, some citations carry a shield icon and one of three labels. Perplexity rates the whole website, not each page, through what it calls its source review process.

What the labels mean

Government marks official government sites, and Perplexity says it helps with legal questions, financial matters and government services. Academic marks scientific sites. Trusted is a broad label for sites that appear often in results and publish within their own area of expertise.

What the review checks

Perplexity gives example questions: does the site correct its mistakes, does it say who wrote each piece, and does it keep news separate from advertising and opinion. It says partnerships and payments don't affect labels, and that most domains have no label, which isn't a negative judgment.

What it means for a law firm

Perplexity doesn't say labels change which sources get cited; they're shown to the reader. But on legal questions, a firm's page will often sit next to a court, bar or immigration agency site wearing the Government label. Pages that cite those official sources, name the lawyer who wrote or reviewed them, and show a real "Last updated" date hold up in that company.

That's the standard we write client pages to: the answer in the first two or three sentences, official sources linked on the page, and a lawyer's name on it only when that lawyer actually reviewed it. It's also how we build law firm websites.

Local questions

Perplexity doesn't document how it picks local businesses, so test with the city in the question.

We found nothing in Perplexity's crawler docs or help center on how it handles location, maps or local business rankings. Treat any claim about Perplexity's local ranking factors as a guess.

What can be observed is the citation list. Ask Perplexity for a lawyer in your practice and city and the numbered sources show which sites it read: your own pages, directories, review sites, news. Those are the places to fix first, because they're the ones in play for that question.

Testing and what we see

Test Perplexity by reading its citations, and judge it over months, not one answer.

Perplexity is simpler to audit than the other AI tools, because every answer shows its sources. Use that.

  • Ask five questions the way clients phrase them, with and without your city named.
  • Open every numbered citation and note the domain, whether it names your firm, and whether the details are right.
  • Run each question more than once; we record the share of runs that name the firm, not a single screenshot.
  • Repeat monthly with the same questions, and compare against the first run.

Perplexity is one of the tools in the monthly checks we run for AI search clients, alongside ChatGPT and Google. Visits from it show up in analytics as referrals, and we group them with the other AI tools so firms can see the trend. Our guide to measuring AI search traffic shows how.

Our clearest AI search results so far come from ChatGPT: one family law firm now shows up first there. Another firm moved into estate planning and started being found inside ChatGPT and similar tools.

The groundwork behind both is the groundwork Perplexity reads, and a six-month reputation build for a new business law firm is the same kind of work. The action plan for AI visibility lists it step by step, and the full prompt method is in our guide to checking what AI says about your firm.

Common mistakes

Where firms go wrong.

We see these when we check firms' sites and listings against Perplexity's answers.

  1. Blocking PerplexityBot to stop training

    Perplexity says PerplexityBot isn't a training crawler. Blocking it shrinks the firm to a headline and a summary, and protects nothing.

  2. Trusting robots.txt alone

    A firewall rule or bot setting at the CDN can block PerplexityBot even when robots.txt allows it. Perplexity publishes firewall instructions for exactly this reason.

  3. Pages with no author and no date

    Perplexity's source review asks whether a site says who wrote each piece. A practice page with no lawyer named and no update date gives a careful reader less reason to trust it.

  4. Ignoring the citation list

    The numbered sources tell you which directories and articles Perplexity read for your city. Firms that fix those pages are working on the actual material; firms that don't are guessing.

FAQ

Questions lawyers ask us.

Straight answers to the questions that come up most.

Does Perplexity use Google or Bing results?

Perplexity's help center says it searches the web in real time and that it partners with third-party crawlers to help build its search index. It doesn't name a search engine it relies on. What it does document is its own crawler, PerplexityBot, so allowing that bot is the setting a firm controls directly.

If we block PerplexityBot, will Perplexity stop using our content?

Mostly. Perplexity says PerplexityBot won't index the full or partial text of a site that disallows it, but it may still index the domain, headline and a brief factual summary. Perplexity-User, which fetches pages a person asks about, generally ignores robots.txt. For most law firms, blocking costs more than it saves.

Is Perplexity training AI models on our website?

Perplexity says it doesn't build foundation models and that PerplexityBot isn't used to crawl content for them. The models that write its answers come from other companies. If your concern is training, the crawlers to look at are the ones other AI companies run for that purpose, each named separately.

Can we get a Trusted label for our firm's website?

There's no application process described. Perplexity says labels come from its source review process alone, rate whole websites, and aren't affected by payments or partnerships. Most domains have no label, and Perplexity says that isn't a negative judgment. Website owners can contact Perplexity support with questions about labels.

Does Perplexity show reviews or star ratings for law firms?

Perplexity doesn't document a local business or review feature for law firms. It cites the pages its search returns, and for lawyer questions those often include directories and review sites. Check the citations for your city and practice to see which ones it read, then make sure your listings there are accurate.

How quickly will Perplexity notice changes to our site?

Perplexity says robots.txt changes can take up to 24 hours to take effect. It doesn't publish how often PerplexityBot recrawls a page. In our experience, getting named in AI answers takes months of steady work, and nobody outside the company can promise when it happens.

Is Perplexity worth the effort for a small firm?

It's rarely separate effort. Allowing PerplexityBot is a one-time check, and everything else it rewards (clear pages, cited sources, accurate listings) also helps in ChatGPT and Google. Our guide on where AI tools find lawyers compares the sources, and our ChatGPT guide covers its differences.

Where do we start?

Check robots.txt and your firewall for PerplexityBot, ask Perplexity five client-style questions, and write down every cited domain along with what each one says about you. Fix the wrong listings first, since those are already in play. To have us read the citations for your own practice and city, schedule a consultation.

Santiago Alvarez

Written by

Santiago Alvarez

Founder of Ad Hoc Digital. Leads strategy and works directly with every client firm on AI search, Local Services Ads, Google Ads and Meta ads.

More about Santiago

Want a second pair of eyes on this?

Book a free 30-minute call. Tell us how cases come in today, and we'll tell you straight what we'd change, and whether we can help.