The short answer
Ranking in Google and being cited in an AI answer are different jobs, and there are five usual reasons a firm does the first and not the second: robots.txt blocking AI crawlers, missing or broken schema, content that isn't structured as questions and answers, entity ambiguity across the web, and no third-party reinforcement. Fix them in that order, because the later ones don't matter while the first is broken.
Here's a combination that surprises people. A firm ranks on page one for its three highest-value queries, has done for a couple of years, and pulls steady organic traffic from all three. Then someone runs those same queries with AI Overviews enabled, or types them into ChatGPT or Perplexity, and the firm doesn't appear anywhere. Competitors ranking below it in classic results are cited in the answer that sits above the blue links.
You can check this on your own site in about five minutes, and it's worth doing before reading any further: take your three best queries, run them in Google with AI Overviews on, then in ChatGPT and Perplexity, and note who gets named.
This is the most counterintuitive pattern in modern legal search. A firm can be doing classic SEO correctly and still be invisible to that layer entirely. Ranking in Google is necessary and no longer sufficient. The broader case for why these are separate disciplines lives in AEO vs SEO for law firms.
Five reasons explain why a top-ranking firm misses AI citations. The diagnostic is fast (under an hour per page). The fixes range from "edit your robots.txt and ship in twenty minutes" to "rewrite your service pages and ship in two weeks." Here's the order to run them in.
Reason one: your robots.txt is blocking the AI crawlers
Run this check first because it's the single most common cause and the cheapest fix.
Visit yourdomain.com/robots.txt. Look for any line that reads User-agent: GPTBot, User-agent: ClaudeBot, User-agent: PerplexityBot, or User-agent: Google-Extended, followed by Disallow: /. If any of those blocks exist, you are explicitly telling ChatGPT, Claude, Perplexity, or Google's AI training crawler that they cannot read your site.
Blocked crawlers are common, and usually nobody at the firm made the decision. The rules get added by a previous agency "to protect content from AI scraping." The intent is reasonable. The effect is that the AI engines can't see the site, and therefore can't cite it. Check yours before assuming this isn't you — it takes thirty seconds and it is the single most common cause on this list.
The diagnostic: open the robots.txt in a browser. Read each block. If you don't recognize a User-agent rule, find out what it blocks before assuming it's harmless. WordPress security plugins (Wordfence, iThemes Security), Cloudflare Bot Management, and several "AI protection" plugins all add these rules automatically. Website builders add their own version of this: GoDaddy and Wix configurations sometimes ship these blocks by default, one of several reasons website builders leave law firms nearly invisible in AI search.
The fix: a robots.txt that explicitly allows the four engines you care about looks like this.
User-agent: GPTBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Google-Extended
Allow: /
User-agent: *
Allow: /
You don't strictly need to allow them explicitly if the default is Allow, but the explicit declaration makes the intent obvious and catches any plugin that tries to override it later.
This is the cheapest fix on the list by a wide margin — a text file edit — and it gates everything else. Nothing further in this article can work while the crawlers are locked out, so do it first even if you plan to do nothing else for a month.
Reason two: no schema markup, or broken schema
The second most common cause, and the one with the biggest compounding return when fixed.
AI engines weight structured data heavily when deciding which firms to cite. A site with valid LegalService, FAQPage, and Person schema reads as a recognized entity. A site with no schema or with broken schema reads as ambiguous, and ambiguous sources don't get cited.
The diagnostic: paste your homepage URL into the Google Rich Results Test. Then paste your three highest-priority practice area pages. Look at what comes back. If the result is "No items detected" or shows only generic WebPage / Article schema generated by Yoast or Rank Math defaults, you don't have the schema layer that AI citation requires.
The seven types we recommend for law firm sites are covered in detail in the seven schema types guide. The short version: LegalService on the homepage, Person on attorney bios, Service on practice areas, FAQPage on Q&A content. Without these four at minimum, AI engines treat your firm as an unknown entity rather than a credentialed practitioner.
The usual culprit is an SEO plugin left on default settings, generating WebPage and Article schema on every page. Neither carries any signal that the business is a law firm, who its attorneys are, or what it practices. It isn't wrong, exactly. It just says nothing. LegalService and Person blocks are what carry that meaning.
Reason three: your content isn't structured as questions and answers
Even with schema fixed, AI engines need extractable answers. The bulk of AI Overview citations come from passages that read "Here's the answer to that specific question" rather than from prose that wanders through context.
The diagnostic: open your three most important pages. Look at the H1 and the first paragraph. Does the H1 read as a question someone would type into ChatGPT? Does the first paragraph deliver a direct answer in 40 to 120 words?
Most law firm sites fail this check. The H1 reads "Houston Personal Injury Lawyer" (a label, not a question) and the first paragraph reads "At Hartman & Reyes, we have over 25 years of combined experience..." (a credential statement, not an answer).
For pages targeting AI citation, the structure that works is this. H1 phrased as a question ("How long do I have to file a personal injury claim in Houston?"). First paragraph as the direct answer ("Two years from the date of the injury under Texas Civil Practice and Remedies Code Section 16.003, with limited exceptions for minors and cases against government entities."). Then expand into the full context.
This doesn't mean your homepage has to be a question. It means that the deep pages where citation actually happens, the service pages, the FAQ pages, the resource articles, need to follow this shape. The build a knowledge base post walks through the editorial model in detail, and the get cited by ChatGPT playbook covers the exact HTML pattern.
Rewriting a set of service pages this way runs about three weeks for a typical firm, most of it deciding what question each page is really answering. It is the slowest of the first three fixes and the one that keeps paying after the schema work has done its job.
Reason four: entity ambiguity
This one is harder to diagnose and slower to fix. AI engines have to disambiguate your firm from every other firm with a similar name, location, or focus. Weak entity signals mean the engine never confidently identifies you as the firm being asked about.
The diagnostic: open ChatGPT and type "Who is Hartman & Reyes PLLC in Houston?" (substitute your firm name). Watch what comes back. If the response is "I don't have specific information about this firm" or worse, hallucinates a different firm with a similar name, you have an entity problem.
Entity confidence is built from a handful of signals that have to agree across the web. The firm's name has to appear identically on the website, the Google Business Profile, the state bar listing, Justia, Avvo, the firm's LinkedIn company page, and a half-dozen legal directories. Inconsistencies (PLLC vs P.L.L.C. vs LLP, "Hartman and Reyes" vs "Hartman & Reyes") split the entity into multiple weak shadows of itself.
The fix is laborious. Audit every appearance of the firm name across the web. Standardize on one version, including punctuation and legal designation. Update Google Business Profile, all bar listings, all directories. This is normally a two-week project across an admin and a junior associate, mostly because the directories are slow to update.
The deeper mechanics live in the entities SEO guide. The short version: AI engines reward consistency disproportionately.
Reason five: no third-party reinforcement
The hardest reason and the one most firms get wrong.
AI engines don't just read your website. They cross-reference what your site says against what other sites say about you. A firm whose own homepage claims it handles personal injury cases will get cited softly. A firm whose homepage claims it handles personal injury cases, AND has matching language on Justia, AND is mentioned in three local news articles, AND has bar profile descriptions that line up, gets cited confidently.
This is the off-site SEO problem reapplied to AI citation. Most firms have a website and a Google Business Profile and maybe a bar listing. That's three sources. The firms getting cited in AI Overviews typically have ten to thirty third-party reinforcements.
The diagnostic: search your firm name in quotes on Google. Count the legitimate third-party mentions in the first three pages of results. Bar associations, news sites, podcast appearances, legal directories with custom-written content (not just a name-and-address listing). Below fifteen real mentions, third-party reinforcement is probably the constraint.
The fix is slow. Real PR, real bar association involvement, real podcast appearances, real publishing in legal trade outlets. The off-site SEO guide covers the playbook in detail.
Be realistic about this one. An established firm with a thin third-party footprint cannot close this gap in a quarter, and anyone promising otherwise is selling something. The honest sequencing is to ship the first three reasons, which are largely on-site and largely within your control, and treat third-party reinforcement as a twelve-month program running underneath. Expect the queries where corroboration is the binding constraint to be the last to move.
The diagnostic in 60 minutes
Walk through this checklist in order. Each step takes under fifteen minutes. By the end you'll know which of the five reasons is your binding constraint.
Step 1: open robots.txt. If GPTBot, ClaudeBot, PerplexityBot, or Google-Extended is disallowed, that's almost certainly the constraint. Stop here, fix it, wait two weeks, retest.
Step 2: run your homepage through the Google Rich Results Test. If LegalService schema is missing or broken, schema is your constraint. Fix the seven types, in priority order.
Step 3: check your three top service pages. H1 a question? First paragraph a direct answer? FAQPage schema present? If any of those is no, content structure is the constraint.
Step 4: ChatGPT identity check. Type "Who is [firm name] in [city]?" If the response is confused or hallucinated, entity ambiguity is your constraint.
Step 5: brand search audit. Quoted firm name in Google. Count real third-party mentions in the first three pages. Below 15, off-site reinforcement is the constraint.
Most firms hit reasons one and two as their binding constraints. The good news: those are also the fastest to fix. The robots.txt change is under an hour. The schema work is under thirty hours. Both produce visible AI Overview movement within four to six weeks.
The order to fix in
If two or more of the five reasons apply, ship the fixes in this order.
One: robots.txt. Highest return for smallest cost. Block any other work until this is right.
Two: schema layer. Compounds with everything downstream.
Three: question-format content on the three most important service pages. Real citations start showing up here.
Four: entity consolidation. Slow but necessary to lock in the gains.
Five: third-party reinforcement. Twelve-month program. Don't start this before the first four are done. The citations won't connect.
Two things are worth setting expectations on. The first three fixes are fast and mostly within your control; the last two are not, and the last one is a year. And the classic organic rankings usually don't move at all through any of this, because they were never the problem. The AI layer is a separate surface, and it is the one you're fixing.
The order, and why it's the order
A robots.txt edit. Gates everything below it — nothing else can work while the engines can't read the site.
LegalService and Person, replacing generic WebPage and Article. Tells an engine what the business is and who practices there.
Question-shaped headings, direct answers up top, statutes cited by section. Gives the engine something quotable.
Consistent name, address, and attorney identity everywhere the firm appears. Slower because it depends on other people's records.
Corroboration from sources you don't own. Starting here before the first four are done wastes it — there's nothing for the citations to connect to.
Send us your URL
The 60-minute diagnostic above will tell you what's broken on your site. If you'd rather have us run it and write up the fix order, send us your URL and three target queries. We'll do the work, hand back a structural audit, and come back with a real quote for what it would take to ship the fixes. The fixes are the same work as our answer-engine optimization service. Free, no card required, no obligation to hire us afterward.
Questions we get about this
-
Why does my firm rank in Google but not appear in AI Overviews?
Because ranking and retrieval are separate mechanisms, and the most common cause is that AI crawlers can't reach you at all. Ranking well proves Googlebot can read your site; it says nothing about whether GPTBot, ClaudeBot, or PerplexityBot are allowed in, and many sites block them by default without anyone deciding to. After access, the usual causes are unstructured content, an entity an engine can't confidently resolve, and nothing outside your own site backing you up. Check them in that order.
-
How do I check whether AI crawlers can access my site?
Read your robots.txt and look for rules naming GPTBot, ClaudeBot, PerplexityBot, Google-Extended, or CCBot. Some platforms and security plugins add those blocks by default, so the absence of a decision is not the absence of a block. Then fetch a key page without JavaScript and confirm your practice areas, city, and phone number are in the raw response. Both checks take minutes and one of them explains most cases.
-
Will adding schema get my firm into AI Overviews?
Not on its own. Schema makes your facts machine-readable, which helps an engine parse what it has already retrieved, but it isn't the mechanism that decides who gets cited. If a crawler is blocked, perfect schema changes nothing; if your entity is ambiguous across the web, schema on one site doesn't resolve it. Ship it because unstructured facts are harder to read and there's no reason to make an engine guess — that's hygiene, not a citation trick.
-
How long does it take to appear in AI answers after fixing this?
Access fixes can be picked up within weeks, since it's a matter of crawlers being allowed in and re-fetching. Structural and entity work is slower, and third-party corroboration is slowest of all because it depends on other people. Nobody can give you a date, and anyone who does is guessing — engines change what they cite without notice. Fix the inputs in order and judge progress on whether the blockers are gone, not on a calendar.
