AEO: how to get cited by ChatGPT, Gemini and Perplexity
Getting cited by an answer engine is neither magic nor a hidden trick. Here is what actually matters, engine by engine, and how to check it on real answers rather than on impressions.
In short
To be cited by ChatGPT, Gemini or Perplexity, their crawlers must first be able to read your pages, and those pages must answer real questions clearly, in the first sentence. Add an unambiguous entity (one name, consistent facts), verifiable sources and a presence on the sites these engines already cite. Then measure on real answers, repeated over time, not on a single screenshot.
On this page
- What an answer engine does
- Answers grounded in a web search
- Answers drawn from the model's knowledge
- Let the right crawlers in
- Write pages that answer first
- A clear, consistent entity
- Content worth citing
- Be present where engines draw from
- llms.txt and markup: what to expect
- Measure on real answers
- A five-step method
- Frequently asked questions
What an answer engine does
An answer engine does not return a list of ten links: it writes an answer. Your brand is in it — named, recommended, or cited with a link to your page — or it is not. Answer Engine Optimization (AEO) covers everything that raises your odds of being in that answer. It is not a discipline separate from SEO: it is SEO applied to an interface that summarises instead of listing.
To act in the right place, separate two families of answers, because they do not draw their information from the same place.
Answers grounded in a web search
Some answers rest on a search run at the moment of the question. Perplexity searches the web and links its sources; ChatGPT's search features rely on the pages read by the OAI-SearchBot crawler; Google's AI Overviews and AI Mode rely on the Google Search index. Google describes two techniques: retrieval-augmented generation (RAG), which retrieves pages from the index to ground the answer, and query fan-out, which issues several related searches to cover subtopics (Google's AI optimization guide).
For this family the rule is simple: a page the engine cannot read, or does not rank well, cannot be cited. SEO fundamentals apply directly.
Answers drawn from the model's knowledge
Other answers come from what the model learned during training, with no search at question time. Your presence there depends on what was published about you, on your site and elsewhere, before its knowledge cut-off. A page published yesterday does not weigh there yet. This is where the consistency of your entity and your presence in third-party sources matter most.
| Family | Where the information comes from | Main lever |
|---|---|---|
| Search-grounded answer | Pages found and read at question time | Crawler access, well-ranked pages that answer |
| Model answer | What the model learned before its knowledge date | A consistent entity, lasting presence in third-party sources |
Let the right crawlers in
The first check, and the most often forgotten: do your robots.txt, CDN or firewall block the answer engines' crawlers? Publishers list their agents and what each one does; read those lists, because every agent is a separate decision.
- OpenAI separates
OAI-SearchBot, used to surface websites in ChatGPT's search features, fromGPTBot, which crawls content that may be used to train its models. The two settings are independent: you can allow one and disallow the other. OpenAI notes that a robots.txt change can take about 24 hours to be picked up (OpenAI crawlers). - Perplexity states that
PerplexityBotsurfaces and links websites in its results and is not used to crawl content for AI foundation models.Perplexity-Useracts on a user's request and, for that reason, generally ignores robots.txt rules (Perplexity crawlers). - Google: AI Overviews and AI Mode go through Googlebot and the Search index. To be shown as a supporting link, a page must be indexed and eligible for a snippet; there are no additional technical requirements (AI features and your website).
# Answer engines that search the web: allowed
User-agent: OAI-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
# Training crawler: an independent decision
User-agent: GPTBot
Allow: /In NeoRank, the Robots directives view reads your live robots.txt for eight AI crawlers, and the simulator tests a specific path. Details are in the documentation on AI crawler access.
Write pages that answer first
An engine more readily reuses a passage that stands on its own. In practice:
- The answer first. One or two sentences that answer fully, then the detail, the nuances and the exceptions.
- Subheadings worded as your customers' questions, with the answer right below.
- Numbered lists for procedures, tables for comparisons. Structure makes each element extractable.
- Self-contained passages: avoid “as seen above”. Each section should make sense alone.
Google does point out that you do not need to chop content into tiny pieces or rewrite it for AI: its systems understand a page that covers several topics. Structure serves the reader first; it helps engines too. The direct answer blocks page details what NeoRank records on every crawled page.
A clear, consistent entity
An engine can only recommend what it identifies without ambiguity. If your brand is spelled three ways, if your offer is described differently on your site and on your external profiles, the answer will reflect it.
- Use the same name everywhere: title, structured data, legal notice, external profiles.
- Say in one sentence what you do, for whom and where, on the home and About pages.
- Describe the organisation in JSON-LD of type Organization, with
sameAspointing to your official profiles. Our JSON-LD guide shows how.
Content worth citing
Google puts it plainly in its AI optimization guide: unique, useful content grounded in real experience will likely influence your presence in generative search more than any of its other suggestions. A summary of what already exists adds nothing a model could not produce itself; see also its guide to helpful, reliable, people-first content.
- Back your claims with dated facts and named sources.
- Publish what only you know: method, first-hand experience, terms, limits.
- Keep the deciding facts current: terms, areas served, lead times, integrations.
Be present where engines draw from
Engines also cite third-party sites: publications, directories, comparison pages, forums. Spotting the ones cited on your questions without your site tells you where a legitimate mention would bring you into the answer. But Google warns that seeking inauthentic mentions is not as helpful as it may seem, because its systems favour quality content and block spam. Aim for real, earned mentions.
llms.txt and markup: what to expect
The /llms.txt file is a proposed convention: a Markdown file that introduces the site and its key pages to language models. Google states that Google Search does not use it: it neither harms nor helps rankings, and there is no special schema.org markup for AI features. It remains fine to publish it for other services that read it. Treat it as a complement, never as a guaranteed lever.
Measure on real answers
The only honest measurement is to ask the engines real questions, store the answers and look for your brand in them. NeoRank calls this an observation: each tracked question is asked to each verified engine, and every answer is kept.
- Mention: your name or domain appears in the answer's text.
- Citation: a URL on your domain is written in the answer. It is matched by domain, not verified page by page (mentions and citations).
- Share of voice: your mentions divided by your mentions plus tracked competitors' mentions, over the last three runs. With no tracked competitor it is 100 as soon as you are named, so it is not comparative (share of voice).
Google's AI Overviews are not queried by NeoRank, since Google offers no API for them: how to measure them is covered in our post on what you can measure of AI Overviews.
A five-step method
- Track 10 to 25 real questions: brand, category (“which tool for…”), comparisons, objections (track a prompt).
- Run an observation and note, for each question, who is cited instead of you and with which page.
- Open the resolution of the questions where you are not cited: the page cited instead of yours shows the content the engine expected.
- Fix: crawler access, a page that answers, the entity, third-party sources.
- Run an analysis then an observation again, and compare over the last three runs.
Frequently asked questions
- Does AEO replace SEO?
- No. Search-grounded answers rest on crawled, well-ranked pages, and Google treats optimising for its AI features as SEO. AEO adds particular attention to the answer itself, the entity and third-party sources.
- Do I need to allow GPTBot to appear in ChatGPT?
- Not necessarily. According to OpenAI, ChatGPT's search features depend on OAI-SearchBot, while GPTBot concerns training. The two settings are independent and are decided separately in robots.txt.
- Is an llms.txt file enough to get cited?
- No. It is a proposed convention that Google Search does not use. It may serve other services, but it replaces neither pages that answer, nor open crawler access, nor credible sources.
- How long does it take to get cited?
- It depends on the family of answer. A search-grounded answer can reflect a page as soon as it is crawled and ranks well; a model answer waits for the next update of its knowledge, with no guaranteed date.
See where your site stands
See where your site stands today, on Google and in AI answers: run the free analysis.
Start the free analysisSources
- OpenAI — Overview of OpenAI crawlers
- Perplexity — Perplexity crawlers
- Google Search Central — Optimizing your website for generative AI features
- Google Search Central — AI features and your website
- Google Search Central — Creating helpful, reliable, people-first content
- Google Search Central — Introduction to robots.txt
- llms.txt — proposal
- Schema.org — Organization