How to Get Cited in ChatGPT and Perplexity: AEO Checklist
To get cited in generated answers by ChatGPT and Perplexity, traditional SEO technical setups are no longer enough. Discover our operational AEO checklist to configure your crawlers, structure your pages, and measure your share of voice.

In short
To appear in answers generated by ChatGPT and Perplexity, you must open your robots.txt file to dedicated web crawlers (OAI-SearchBot, PerplexityBot), structure your content with immediate concise answers, integrate validated structured data (Schema.org markup), and monitor your citations using a dedicated AEO measurement tool.
On this page
- Step 1: Allow AI crawlers in your robots.txt file
- Official user-agents for OpenAI and Perplexity
- Step 2: Write direct, structured answers
- Formatting content for automated extraction
- Step 3: Implement appropriate structured data
- The FAQPage schema and its role in AEO
- Step 4: Evaluate the llms.txt file (proposed standard)
- Structure and utility of the llms.txt file
- Step 5: Measure and track your AI citations
- Analyzing share of voice and brand mentions
- Our analysis: limitations and uncertainties of AEO in 2026
- What remains to be confirmed
- Frequently asked questions
Step 1: Allow AI crawlers in your robots.txt file
The very first requirement to appear in generative AI search engines is to grant their web crawlers access to your website content. Many webmasters inadvertently block recent user-agents inside their robots.txt file. To appear in ChatGPT Search and Perplexity, specific crawling permissions are required. You can check your technical crawl settings and accessibility with a site audit.
Official user-agents for OpenAI and Perplexity
OpenAI relies primarily on OAI-SearchBot for real-time web search retrieval and GPTBot for model retraining or general content collection [R10]. On its end, Perplexity uses the PerplexityBot crawler to browse and index web pages live [R11]. If your robots.txt file contains a Disallow rule targeting these agents, your pages remain invisible to these answer engine models.
User-agent: OAI-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /It is wise to verify your site configuration with an automated technical audit to avoid any unintentional blocking of user-agents or crawling paths.
Step 2: Write direct, structured answers
Answer engines like ChatGPT and Perplexity do not seek to display a list of blue links, but rather to synthesize direct, immediate information for the user. To capture these direct citations, your editorial structure should follow the inverted pyramid method. Understanding how to format your text for synthetic retrieval is a foundational concept in AEO strategy.
Formatting content for automated extraction
Place a clear, concise answer of 2 to 4 sentences right at the start of your articles or directly beneath your level 2 headings (##). Large language models can easily identify and extract these self-contained blocks of information. You can then flesh out your topic with contextual details, bulleted lists, and numbered action steps.
- Write a self-contained concise paragraph immediately following each subheading.
- Use precise, factual terminology rich in named semantic entities.
- Avoid vague introductory fluff that delays delivering the core facts.
The fundamentals of AEO and synthetic visibility rely heavily on this fundamental clarity within the source text.
Step 3: Implement appropriate structured data
Schema.org markup provides an explicit semantic layer that AI models leverage to validate the precise nature of your data. Implementing structured markup facilitates the clear identification of entities, organization details, and frequently asked questions across your domain. By delivering unambiguously labeled data formats to automated scrapers, your web pages become far easier for AI language models to interpret correctly without relying solely on heuristic text extraction.
The FAQPage schema and its role in AEO
The Schema.org FAQPage markup allows you to pair each specific question with a concise, factual answer. Even as the display of rich snippets in Google SERPs has evolved over time, JSON-LD structured data remains a preferred reading format for language model web crawlers. When automated crawlers parse a page equipped with JSON-LD, they can directly link the precise inquiry entity to its corresponding answer payload, reducing parsing errors during information retrieval.
{
"@context": "https://schema.org",
"@type": "FAQPage",
"mainEntity": [{
"@type": "Question",
"name": "How to configure robots.txt for ChatGPT?",
"acceptedAnswer": {
"@type": "Answer",
"text": "You need to allow the OAI-SearchBot user-agent using the Allow: / directive in your robots.txt file."
}
}]
}Adding structured markup ensures that AI engines unambiguously identify answer pairs across your domain's content. Integrating explicit Schema.org entities across your templates helps bots process key concepts without hallucinating facts.
Step 4: Evaluate the llms.txt file (proposed standard)
The llms.txt file is an emerging proposed standard designed to provide language models with a curated list of clean Markdown documents [R18]. Located at the root of a website, this file helps AI agents identify the most relevant site resources without having to parse complex raw HTML pages, navigation menus, or client-side JavaScript applications. By maintaining a clean index of Markdown text files, site owners can offer language model agents a streamlined map of core assets.
Structure and utility of the llms.txt file
Although this format remains a proposed standard rather than a strict requirement enforced by every AI vendor [R18], setting it up provides a positive signal to the broader AI ecosystem. The file typically contains a title, a short site overview, and direct links to essential Markdown documents. Implementing a clean directory file reduces computational overhead for file readers and automated scrapers scanning your domain's content hierarchy.
# Site Name
> Brief summary of documentation and core site resources.
## Main Documentation
- [Article Title](https://example.com/page.md): Concise description of the content.This delivery mechanism streamlines content ingestion for autonomous agents seeking clear, directly actionable documentation. Providing organized Markdown resources guarantees that automated crawlers receive unpolluted factual text directly from the source.
Step 5: Measure and track your AI citations
Answer Engine Optimization cannot move forward without precise performance measurement across each generative engine. Unlike traditional search engines, your presence in ChatGPT or Perplexity cannot be evaluated solely by monitoring rank position on a single target keyword. To monitor your actual brand share of voice and generated references, you can leverage dedicated tools for AI visibility.
Analyzing share of voice and brand mentions
To evaluate whether your optimization efforts are working, you must track your brand visibility and URL citations across generated answers. You can use a per-engine AI visibility tracking module to monitor how your core search queries perform.
- Define strategic core prompts that accurately represent your business activity.
- Track specific URL citations and brand mentions generated across ChatGPT and Perplexity.
- Analyze overall sentiment and factual accuracy in the answers output by AI engines.
This ongoing monitoring process enables you to refine your content strategy wherever AI response engines still lack technical precision.
Our analysis: limitations and uncertainties of AEO in 2026
According to our analysis, optimizing for answer engines still presents several areas of uncertainty. Unlike traditional SEO, which has stabilized over many years, decision-making algorithms within ChatGPT and Perplexity are evolving rapidly. Because generative engines periodically update how they select live web sources versus historical training sets, site owners must maintain realistic expectations regarding citation stability.
What remains to be confirmed
Our hypothesis is that the actual performance impact of llms.txt files will depend heavily on widespread industry adoption by major AI ecosystem players [R18]. Furthermore, precisely measuring outbound referral traffic driven by AI answer citations remains challenging in conventional web analytics platforms. Until tracking capabilities fully standardize, balancing traditional search visibility with answer engine optimization techniques remains essential for risk mitigation.
It is plausible that the relative weight assigned to direct citations will change as generative models integrate newer real-time fact-checking techniques. Maintaining continuous testing alongside rigorous data monitoring remains the only reliable method for adapting your AEO strategy.
Frequently asked questions
- Which crawlers must be allowed for ChatGPT and Perplexity?
- To permit indexing by OpenAI and Perplexity, you must configure your robots.txt file to allow OAI-SearchBot, which handles dedicated OpenAI search retrieval, and PerplexityBot, which manages real-time web index crawling for Perplexity. Allowing both user-agents ensures your site pages remain accessible to live generative search systems.
- Is the llms.txt file mandatory to get cited?
- No, the llms.txt file is an emerging proposed standard rather than a mandatory requirement. Implementing it helps autonomous AI agents digest your core documentation efficiently in clean Markdown, but its absence will not automatically block your website from being indexed or cited by ChatGPT or Perplexity.
- How can I track if my site is cited by Perplexity or ChatGPT?
- You can measure brand mentions and direct URL citations by using specialized AI visibility platforms designed to track synthetic answers. These systems continuously monitor target user prompts across models like ChatGPT and Perplexity, calculating your brand's share of voice, citation placement, and overall sentiment in real time.
- What is the core difference between SEO and AEO?
- SEO focuses primarily on earning higher positions in search engine result pages to drive direct organic clicks. In contrast, AEO (Answer Engine Optimization) aims to structure content so that generative AI models extract, synthesize, and explicitly cite your domain as the source within direct conversational answers.
See where your site stands
See where your site stands on Google and in AI answers: run the free analysis.
Start the free analysis