How Directories Get Cited in AI Answers
What is actually documented about AI Overviews citations: the Search Console switch, the report that only counts impressions, and four crawlers.

Google publishes two pages about appearing in its AI features, and they do not say the same thing.
AI features and your website, last updated 10 December 2025, is emphatic: "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary." Seven months later, the generative AI optimization guide, last updated 10 July 2026, adds a condition: "In addition to the technical requirements for Search, a site must be included in Search generative AI features in Search Console to be eligible for display in generative AI features on Google Search."
That second sentence points at a setting. It exists, it has a default, and most directory operators have never opened it. Everything below is what is written down by the companies that run these systems — Google, OpenAI, Perplexity, Anthropic, Microsoft — with the folklore left out. There is less of it than the GEO industry implies, and the parts that are real are mostly switches rather than tactics.
The one switch that is genuinely binary
In Search Console, under Settings → Search generative AI, there is a control with three states. Google's help page for it describes the default as: "Include my site's links and content in Search generative AI features: Your site's content can appear in Search generative AI features, including showing up as links and helping to ground AI responses in these features. Your site can receive impressions and traffic from these features. This is the default control for all properties."
So you are opted in already. The switch matters in the other direction, and the cost of flipping it is stated plainly: "You won't receive any traffic or impressions from these features." Followed by the sentence that makes the decision for most directories: "Content from other sites will still be available in those features, and it may appear similar to yours."
Three details worth knowing before anyone on your team touches it:
- It is not a ranking signal. Google: "this control isn't used as a ranking or inclusion signal affecting other parts of Search."
- It is not a training control. "This control doesn't affect AI training; to limit training of the models used to generate responses in Search generative AI features, use Google-Extended."
- It propagates in days, not minutes: "Content will be excluded within 1-2 days after the control goes live, but some content may take longer to be excluded due to caching and propagation across Google systems."
There is also an inheritance rule that bites anyone running listings on subdomains: "By default, a property inherits its Search generative AI control from its closest parent that has changed its control to stop inheriting." If your directory lives on directory.example.com and someone excluded example.com two years ago, you inherited that decision without being told.
If your Generative AI performance report is empty, Google's help page names the two causes: not enough impressions, or — "This may be because you've excluded your site from Search generative AI features." Rule out the switch before you spend a sprint on content.
The report exists now, and it counts one thing
Google announced the Search Generative AI performance reports on 3 June 2026, and both the blog post and the help page now carry the same banner: "As of August 31, 2026, we've rolled out these insights to all websites worldwide." If you last looked in June and found nothing, look again.
What it gives you is narrower than most people assume. The metric is impressions — "how many times links to your site were shown to a user in a generative AI feature on Google Search." There is no clicks column, no CTR, no average position. The dimensions are exactly four: Pages, Countries, Dates, Devices. The features covered are exactly two, AI Overviews and AI Mode, and there is no dimension that separates them. Data from Search Labs experiments is excluded.
Four properties of this data change how you should read it on a large catalogue:
| Property | Google's wording | What it means on a 50,000-page directory |
|---|---|---|
| Canonical attribution | "most performance data in this report is assigned to the page's canonical URL, not to a duplicate URL" | Faceted and filtered listing URLs roll up into the canonical. Your facets look invisible even when they aren't. |
| Deduplicated impressions | "if two results from the same site appeared in a generative AI search results feature, they count as a single impression in the chart total" | Two of your listings cited in one answer is one impression. Breadth of citation is not measurable here. |
| Row cap | "the usual data limitations (1,000 row limitation, time period, etc) ... also apply" | You see your top 1,000 URLs. On a catalogue in the tens of thousands, the long tail — the whole point of a directory — is off the page. |
| Preliminary data | "The newest data can be preliminary ... and might change in the next few hours" | Do not build a daily dashboard off yesterday and act on it. |
The honest summary: this report tells you whether you appear, roughly where, and roughly how much. It cannot tell you what that appearance is worth, because the click side of the equation is not published. Google's only claim on that front, from the AI features page, is qualitative — "when people click from search results pages with AI Overviews, these clicks are higher quality (meaning, users are more likely to spend more time on the site)" — with no figure attached. Treat it as a vendor statement, because that is what it is.
Why a directory can show up at all
The mechanism Google documents is worth understanding, because it explains why aggregators are structurally well placed and why one common directory tactic is explicitly a spam risk.
Two definitions from the optimization guide. Grounding: "Retrieval-augmented generation (RAG): A technique (also known as grounding) used to improve the quality, accuracy, and freshness of AI responses by relying on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index." And the retrieval pattern: "Query fan-out: A set of concurrent, related queries generated by the model to request more information and fetch additional relevant search results to address the user's query."
The AI features page draws the consequence: fan-out lets Google "display a wider and more diverse set of helpful links associated with the response than with a classic web search."
A directory built around a real taxonomy — category, region, city, listing — already has a page for a large share of the sub-queries a fan-out will generate. That is not a new insight so much as the old argument for designing directory URLs around the taxonomy collecting a second dividend.
The trap is on the next line of the same guide: "While it might be tempting to create separate content for every possible variation of how people might search (for example, by focusing on other queries that people have asked, or fan-out queries), doing so primarily to manipulate rankings or generative AI responses in Google Search violates Google's scaled content abuse spam policy."
That is the difference between a city page backed by listings and a city page backed by a string replacement. It is also why DirectoryLaunch's local SEO pages emit noindex below a minListingsToIndex threshold instead of publishing every combination the data can spell: a page with two listings on it is a fan-out target you cannot defend.
The eligibility floor is boring and frequently broken
Both Google pages agree on the actual entry requirement: "a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements."
Two failure modes, both common on directories.
The first is indexing. If Google has your listing URLs filed under Discovered — currently not indexed, the AI question is moot; there is nothing in the index to retrieve. This is the ordinary crawl-budget problem that large catalogues have always had, and it now costs you two surfaces instead of one.
The second is snippet controls. The AI features page names them together: "To limit the information shown from your pages in Search, use nosnippet, data-nosnippet, max-snippet, or noindex controls." Directories accumulate these by accident — a data-nosnippet wrapper placed around aggregated third-party descriptions during some long-forgotten legal review, or a global max-snippet set low to discourage scraping. Those decisions were made about SERP snippets. They now also govern whether the page can be used to ground an answer.
Also from that page, for whoever eventually asks why the change didn't take: "crawling can take anywhere from several days to several months, depending on how often our systems determine a page needs to be refreshed."
Four assistants, four different control surfaces
Google is one of several systems that might cite a listing page, and the others expose their controls through robots.txt rather than a console. The important structural fact, which almost no boilerplate robots.txt reflects, is that the bot that decides whether you get cited is a different bot from the one that collects training data, and every vendor here says you can treat them separately.
| System | Bot that governs citation | Bot for model training | Documented cost of blocking the citation bot |
|---|---|---|---|
| ChatGPT | OAI-SearchBot | GPTBot | "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links." |
| Perplexity | PerplexityBot | none — "Perplexity does not build foundation models" | Blocked pages may still yield "the domain, headline, and a brief factual summary." |
| Claude | Claude-SearchBot | ClaudeBot | "prevents our system from indexing your content for search optimization, which may reduce your site's visibility and accuracy in user search results." |
| Google Search AI | Googlebot + the Search Console control | Google-Extended | "You won't receive any traffic or impressions from these features." |
Sources, in order: OpenAI's crawler documentation, Perplexity's crawler documentation and its help-centre article "How does Perplexity follow robots.txt?" (last updated 16 July 2026; the help centre blocks automated fetches, so read it in a browser), and Anthropic's crawler page, last updated 7 April 2026.
OpenAI states the separation explicitly: "a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training OpenAI's generative AI foundation models." Changes are not instant there either — "it can take ~24 hours from a site's robots.txt update for our systems to adjust." One useful side effect for measurement: "ChatGPT automatically includes the UTM parameter utm_source=chatgpt.com in referral URLs."
There is a third category on every vendor: the user-triggered fetcher, which behaves differently. OpenAI on ChatGPT-User: "Because these actions are initiated by a user, robots.txt rules may not apply." Perplexity on Perplexity-User: "Since a user requested the fetch, this fetcher generally ignores robots.txt rules" — while the separate Perplexity help article cited above says "PerplexityBot only crawls content in compliance with robots.txt." Both pages are live and the two statements are scoped to different agents; nobody has reconciled them publicly. Anthropic is the outlier that makes no such carve-out, stating flatly that "Anthropic's Bots respect 'do not crawl' signals by honoring industry standard directives in robots.txt."
Microsoft is the one place with a citation count rather than an impression count. Its AI Performance report in Bing Webmaster Tools, announced 10 February 2026 as a public preview, promises "visibility into which URLs are referenced and how citation activity changes over time," plus a "Grounding queries" dimension showing "the key phrases the AI used when retrieving content." Microsoft is careful about what that means: "This reflects how often pages are cited, not page importance, ranking, or placement."
Freshness, and what IndexNow does and does not buy you
A directory's content changes constantly — new listings, closed businesses, changed hours. The only documented lever on freshness across several of these systems is IndexNow: submit up to "10,000 URLs per post," using a key of "a minimum of 8 and a maximum of 128 hexadecimal characters," to any one participating endpoint, and the submission is shared with the others.
Be precise about the benefit. Microsoft's claim is that IndexNow "helps ensure that AI systems reference the most current version of a page when generating answers." That is a freshness claim, not an eligibility claim — no official page says IndexNow gets you cited. And it is not free: IndexNow's own FAQ says "Every URL submitted through IndexNow counts toward your site's crawl quota," recommends waiting "at least 5 minutes between updates before resubmitting," and states plainly that "Submitting a URL does not guarantee immediate indexing."
For a catalogue that publishes in batches, this is a good trade. For one that fires a submission on every field edit, it is a way to spend crawl budget on nothing.
What you have that a model does not
Everything above is plumbing. The guide's own ranking of what matters puts content first: "Creating content that people find unique, compelling, and useful will likely influence your website's presence in generative AI search in the long run more than any of the other suggestions in this guide."
The distinction it draws is the one that should shape a directory's roadmap: "Commodity content (for example, something like '7 Tips for First-Time Homebuyers') is often based on common knowledge, which could originate from anyone... In contrast, non-commodity content... provides unique expert or experienced takes that go beyond common knowledge."
A model can already write the description of any business on your site. What it cannot generate is the part of your catalogue that only exists because you collected it: verified opening hours, a price band you gathered from twelve suppliers in one city, which of forty vendors actually serves a given postcode, who was removed and why. If your listing pages are re-worded copies of the businesses' own about-pages, you are commodity content in a system explicitly built to prefer the other kind — and that holds whether or not any AI feature ever existed.
One more line from the guide that directories under-read: "Where appropriate, generative AI responses can include product listings, product information, and information about local businesses. Using products like Merchant Center... and Google Business Profiles can help your products and services to be visible in both AI responses and other Google Search results." That is about your own business, not your listings — but if you sell paid placements or lead subscriptions, it is a surface you are probably not using.
What to do tomorrow morning
Ordered by effort, not by ambition. The first three take under an hour between them.
- Open Search Console → Settings → Search generative AI. Confirm the property is set to include, and confirm what it inherits if your listings run on a subdomain. This is the only documented on/off switch in the entire subject.
- Open the Generative AI performance report. It reached all properties on 31 August 2026, so an empty report now means either genuinely low impressions or the switch above. Note which page types appear — and remember the 1,000-row cap before you conclude your long tail is absent.
- Read your own
robots.txtagainst the table above. If a past decision to block AI training also blockedOAI-SearchBot,PerplexityBotorClaude-SearchBot, you removed yourself from three assistants' answers to protect against a different thing entirely. Every one of those vendors documents that the two are separable. - Grep your templates for
nosnippet,data-nosnippetandmax-snippet. Find out who added them and why. If the reason was scraping, the cost is now larger than the reason. - Fix indexing before anything else. Pages that were never indexed cannot be retrieved to ground an answer, and no amount of AI-specific work substitutes for that.
- Pick one category and make it non-commodity. One field on every listing that nobody else in your niche has collected. That is the only item on this list with no ceiling and no vendor between you and the result.
And a caveat to carry into whatever meeting this comes up in: Google publishes impressions, Microsoft publishes citations, and neither publishes what either is worth. Anyone quoting you a conversion rate from AI answers is quoting an estimate. Build for the surface, measure what is actually reported, and keep your revenue attribution somewhere that counts money.