Answer Engine Optimisation (AEO) & AI Search: The Complete Guide
The rules of search have changed. Over 40% of Google searches now display an AI-generated answer before any blue links. Learn how to optimise for Google AI Overviews, Perplexity, ChatGPT, and Gemini, and get your content cited where it counts most.
What is Answer Engine Optimisation (AEO)?
Answer Engine Optimisation (AEO) is the practice of structuring, formatting, and signalling content so that AI-powered answer engines, including Google AI Overviews, Perplexity, ChatGPT, and voice assistants, select it as a cited source inside generated answers. Unlike traditional SEO, the goal is not just to rank on page one, but to become the answer itself.
The emergence of answer engines represents the most significant shift in search behaviour since the introduction of the universal search results page in 2007. When a user asks Google "what is the best framework for building a REST API in Python?", they no longer necessarily see ten blue links and choose one. Instead, they receive a synthesised, paragraph-form answer at the top of the page, drawn from multiple sources, which may or may not be visible to the user.
This has profound implications for how content generates value. A page that earns a citation in a Google AI Overview can receive significant referral traffic even if it doesn't rank first in organic results, and a page that ranks #1 organically may see a dramatic drop in click-through rate if an AI Overview answers the question before the user ever scrolls to the results.
AEO builds on the foundations of traditional SEO, technical health, E-E-A-T signals, topical authority, but adds specific new requirements:
- Question-first structure: Content should anticipate the exact questions users ask, placing a clear, direct answer within the first two paragraphs of each section.
- Passage-level quality: Because AI systems often extract specific passages rather than summarising whole pages, each paragraph should be able to stand alone as a credible, accurate answer.
- Structured data signals: FAQPage, HowTo, Speakable, and Article schema markup help AI systems understand which sections of a page contain answers.
- Topical completeness: Pages that address every sub-question related to a topic, not just the headline question, are more likely to be cited across multiple fan-out sub-queries.
- Authority signals: Original research, data citations, named authors with credentials, and links from trusted sources all increase the likelihood of an AI system treating your content as authoritative.
Use our free AEO Score Checker to audit any URL across 18 AEO signals, including passage density, structured data completeness, readability, and E-E-A-T indicators. The tool generates a prioritised list of improvements you can make today.
AEO is not a replacement for SEO, it is an extension of it. Pages optimised for answer engines also tend to perform better in traditional search because they are clearer, better-structured, and more authoritative. The discipline demands a higher standard of content creation, one where vague generalities are replaced by specific, verifiable, passage-ready answers.
How does Google AI Overviews decide what to show?
Google AI Overviews uses Google's existing search quality infrastructure, including its Knowledge Graph, passage ranking (MUM/Gemini), and E-E-A-T signals, to synthesise answers. It primarily cites pages that already rank well for the query or related queries, prioritising sources with strong Experience, Expertise, Authoritativeness, and Trustworthiness signals, clear passage-level answers, and relevant structured data.
Google AI Overviews (formerly Search Generative Experience, or SGE) is powered by a version of Google's Gemini model, integrated tightly with Google's existing search infrastructure. Crucially, AI Overviews do not search the live web in the way Perplexity or Bing Copilot do, they operate on Google's pre-indexed, pre-ranked corpus.
This means the first prerequisite for appearing in an AI Overview is simply being indexed and ranking reasonably well for queries related to your topic. Google has confirmed that AI Overviews pull predominantly from top-ranking pages, not exclusively from #1, but typically from the top 5 to 10 results for the query and closely related sub-queries.
Beyond that, the selection process appears to weight several factors:
- Passage-level relevance: Google's MUM and Gemini models can evaluate individual passages, not just pages. A page that ranks #8 but contains a highly relevant, clear passage may be cited over a #2 page that buries its answer in dense prose.
- E-E-A-T alignment: Google's quality rater guidelines heavily influence AI Overview source selection. Authors should have demonstrable expertise; content should cite sources, include data, and show first-hand experience where relevant.
- Structured data: FAQPage schema, HowTo schema, and Article schema help Google's systems understand the intent and structure of content. Speakable schema specifically signals which passages are designed to be extracted and read in isolation.
- Content freshness: For rapidly evolving topics, Google AI Overviews prefer recently updated content. Date-stamping your content and keeping it current is essential in fast-moving niches like AI, technology, and health.
- SafeSearch alignment: AI Overviews avoid citing content that is YMYL (Your Money, Your Life) without exceptional authority signals. Medical, legal, and financial content is subject to far stricter sourcing requirements.
Use the AI Overview Checker to test whether your page, or any competitor page, currently appears in Google AI Overviews for a given keyword. The tool runs a live SERP check and reports citation status.
One important nuance: AI Overviews do not appear for every query. Google currently triggers them most frequently for informational, question-format queries, particularly how/what/why questions, comparison queries, and multi-part queries. Navigational and transactional queries (e.g. "buy running shoes", "Gmail login") rarely trigger AI Overviews. This means AEO effort is most valuable for the informational content at the top of your funnel, not for product or conversion pages.
What is GEO and how is it different from AEO?
GEO (Generative Engine Optimisation) is the broader discipline of optimising for all generative AI surfaces, including ChatGPT, Claude, Gemini, Perplexity, and Microsoft Copilot, not just answer engines attached to search. AEO focuses specifically on answer engines (Google AI Overviews, voice search). GEO encompasses brand presence across the training data, web, and live retrieval systems that inform every major LLM.
The distinction between AEO and GEO is more than semantic, it reflects fundamentally different mechanisms by which content reaches users through AI systems.
AEO targets answer engines that operate on top of real-time or near-real-time search indexes. When a user asks Google Assistant a question, or when an AI Overview appears in search results, the system is retrieving from a live (or very recently indexed) corpus. Content optimised for AEO needs to be discoverable through normal search crawls and pass quality filters at index time.
GEO is more complex because it involves two distinct channels:
- Training data inclusion: The underlying knowledge of models like GPT-4, Gemini Ultra, and Claude comes from their training data, massive web crawls performed months or years before the model was deployed. Getting content into training data means being indexed and highly valued by the web at the time of the training crawl. This is largely a function of domain authority, content quality, and being cited by other authoritative sources, classic authority-building SEO.
- Live retrieval augmentation (RAG): Many modern AI chat interfaces use Retrieval-Augmented Generation (RAG) to supplement their trained knowledge with live web retrieval. ChatGPT with Browsing enabled, Perplexity, and Bing Copilot all use RAG. For these systems, content discoverability through normal search channels matters, but so does crawl accessibility for the specific bots each system uses.
Google AI Overviews, voice search, featured snippets, structured data, E-E-A-T, FAQPage schema, Speakable schema, passage-level optimisation.
Brand mention building across authoritative domains, LLM-friendly writing style, citation-worthy original research, open web presence, RAG accessibility via robots.txt.
In practice, the two disciplines share most of their tactical foundations: write clearly, answer questions directly, build topical authority, cite your sources, earn mentions from credible sites. Where they diverge is in measurement, AEO success is measurable through SERP monitoring tools, while GEO success requires actively querying AI platforms and monitoring brand mention frequency across chat interfaces.
A complete AI search strategy addresses both AEO and GEO simultaneously. The foundational content quality work serves both. The differentiation comes in the distribution layer: earning editorial citations, press coverage, and expert quotes from high-authority domains serves GEO; structured data implementation and passage-level formatting serves AEO.
What is query fan-out and how do LLMs use it?
Query fan-out is the process by which an AI search system decomposes a single user query into multiple sub-queries, retrieves answers for each independently, and then synthesises those answers into a single coherent response. A query like "how to fix slow WordPress site" might fan out into 4–6 sub-queries covering hosting, caching, image optimisation, plugin bloat, and database performance. Content covering the broadest range of these sub-queries is most likely to be cited.
Query fan-out is one of the most important, and least understood, mechanisms in AI search. It was introduced publicly by Google as part of its MUM (Multitask Unified Model) research and has become central to how both AI Overviews and Perplexity retrieve information.
Here's how it works in practice: when a user submits a query, the AI system doesn't treat it as a single retrieval problem. Instead, it uses its language model to identify all the component questions embedded within the query. Each component becomes a separate retrieval job. The system then fetches the best available answer for each sub-query from its index, and the language model synthesises these retrieved passages into a single response.
Consider the query: "What's the best diet for someone with type 2 diabetes who also exercises regularly?" A fan-out system might decompose this into:
- What foods should type 2 diabetics eat?
- What foods should type 2 diabetics avoid?
- How does exercise affect blood sugar in type 2 diabetes?
- What macronutrient ratios are recommended for diabetic athletes?
- What does the research say about low-carb diets and type 2 diabetes?
A single page that addresses all five sub-questions is far more valuable to the AI system than five separate pages, each addressing one. The AI system can retrieve everything it needs from one source, making citation more likely and citation of a larger proportion of your content more likely.
When planning content, use the PAA Extractor to identify the full set of "People Also Ask" questions around your target topic. These PAA questions are a strong proxy for the sub-queries a fan-out system will generate. Content that answers all PAA questions comprehensively, not just the head question, is structurally better positioned for AI citation.
Query fan-out also explains why topical authority, not just individual page quality, is critical for AI search performance. An AI system that encounters your domain frequently while retrieving answers to many sub-queries within a topic builds a higher degree of implicit trust in your content. Sites that dominate a topic cluster tend to be cited disproportionately across all related queries, not just the ones they rank #1 for.
This creates a compound effect: the more comprehensively you cover a topic at the cluster level, the more often fan-out retrieval encounters your content, the more citations you receive, which in turn reinforces your topical authority. AEO rewards the same comprehensive, entity-complete approach to content that has long been the best practice in semantic SEO.
How do LLMs choose which sources to cite?
LLMs in retrieval-augmented mode choose sources based on: semantic relevance of the retrieved passage to the sub-query, the domain authority and trustworthiness of the source as evaluated at crawl time, the freshness of the content for time-sensitive queries, and the structural clarity of the answer (short, direct, passage-level answers are preferred over long narrative prose). Writing style, citation habits, and author credibility signals also play a role.
This question gets to the heart of GEO strategy. When a retrieval-augmented AI system, like Perplexity, ChatGPT with web search, or Bing Copilot, goes to fetch sources for a query, it is not simply reproducing the organic search ranking. It is running its own relevance assessment on the retrieved passages, and that assessment has specific characteristics.
Based on published research and observed behaviour across these platforms, here is what appears to drive source selection:
1. Passage-level semantic match. Modern RAG systems don't retrieve whole pages, they retrieve chunks (typically 256–512 token chunks) of content that have been embedded and stored in a vector database. The chunk that most closely matches the semantic embedding of the sub-query wins the retrieval slot. This means a single well-written paragraph that perfectly answers a specific question can get a page cited even if the rest of the page is only tangentially related.
2. Source authority at crawl time. Most RAG systems pre-compute a trustworthiness score for sources based on their domain authority, inbound link profile, and historical citation patterns across the web. High-authority domains get a systematic boost in retrieval ranking, independent of individual passage quality. This is why established publishers, university research pages, and widely-linked explainers dominate AI citations.
3. Writing clarity and directness. Several studies of AI citation patterns, including the Princeton/Georgia Tech GEO research paper, have found that content written in a more direct, declarative style is cited more frequently. Content that buries answers in qualifications, hedges every statement, or requires significant context to understand is less likely to be retrieved as a clean, citable passage.
4. Explicit citations within content. The GEO research found that adding explicit citations and statistics to content increases citation frequency by AI systems. This seems counterintuitive but makes sense: AI systems are calibrated to prefer content that itself signals epistemic rigour, content that shows its work.
5. Content freshness. For time-sensitive topics, recency signals matter significantly. Perplexity in particular strongly prefers freshly indexed content for queries about current events, product releases, or evolving research areas.
| Factor | Google AI Overviews | Perplexity | ChatGPT (Browse) |
|---|---|---|---|
| Domain authority | Very high weight | High weight | High weight |
| Passage relevance | Very high weight | Very high weight | High weight |
| Structured data | High weight | Low weight | Low weight |
| Content freshness | Medium weight | Very high weight | Medium weight |
| Organic rank | High weight | Medium weight | High weight |
| Author E-E-A-T | Very high weight | Low weight | Low weight |
Tools to Accelerate Your AI Search Performance
All tools are free. No signup required for most. Powered by live data.
AI Overview Checker
See if any page appears in Google AI Overviews for your keywords.
Live SERP DataAEO Score Checker
Audit any URL across 18 AEO readiness signals and get a prioritised fix list.
Free AuditPAA Extractor
Extract all People Also Ask questions for any keyword, your fan-out roadmap.
No SignupSchema Generator
Generate FAQPage, Speakable, Article, and HowTo schema for any content.
12 TypesHow to optimise for Perplexity, ChatGPT, and Gemini
Optimising for Perplexity, ChatGPT, and Gemini requires: allowing their crawlers in robots.txt, ensuring fast page load and clean HTML structure, writing in a clear question-and-answer format with cited statistics, building brand mentions across authoritative external sources, and keeping content fresh. Each platform has different retrieval mechanisms, so no single tactic covers all three, a platform-aware approach is needed.
Each major AI answer platform has a different architecture, and effective optimisation requires understanding those differences.
Perplexity operates primarily as a real-time RAG system. It crawls the web using PerplexityBot (check your server logs for this user agent) and retrieves content during query execution. Perplexity's source selection is heavily influenced by the top organic results on Google and Bing for related queries, if you don't rank in the top 10 for related queries, you are unlikely to be retrieved. Beyond that, Perplexity strongly prefers:
- Content with explicit data points, statistics, and source citations
- Comparative content (vs. articles, comparison tables, pros/cons lists)
- Concise, structured writing with clear headings
- Fast-loading pages with minimal JavaScript rendering requirements
- Fresh content, Perplexity actively penalises stale pages for time-sensitive queries
ChatGPT with Browse (and GPT-4o with web search) uses Bing's index as its primary retrieval layer, supplemented by OpenAI's own crawl via GPTBot. Pages that perform well in Bing organic results are systematically advantaged. ChatGPT's source selection also appears to be influenced by the diversity of perspectives on a topic, it often cites 3–5 distinct sources rather than drawing heavily from one. Writing in a uniquely differentiated voice, with original data or angles, makes you more likely to be the source selected for your specific angle rather than being crowded out by a mainstream publisher.
Before anything else, verify that your robots.txt does not block PerplexityBot, GPTBot, GoogleBot, or Googlebot-Extended. These crawlers need explicit permission on some configurations. Blocking them means you are invisible to those platforms' retrieval systems, no amount of content quality will compensate for a crawl block.
Gemini (Google's AI assistant, distinct from AI Overviews) uses a blend of Google's Knowledge Graph, Google Search index, and real-time retrieval. Gemini tends to be more conservative in source citation than Perplexity, preferring high-authority publishers and Google's own properties (YouTube, Google Scholar, government sites). For most web publishers, the best path to Gemini citations runs through Google AI Overviews, the quality signals Google values for one tend to correlate closely with the other.
Microsoft Copilot (formerly Bing Chat) is deeply integrated with Bing's organic search index. Strong Bing SEO performance, which often mirrors Google SEO but with more weight on exact-match signals and less sophisticated semantic understanding, is the primary lever. Copilot also strongly prefers content from sites that are verified in Bing Webmaster Tools.
The universal principles across all platforms: be crawlable, be fast, be authoritative, be specific, and be cited by others. A site with genuine topical authority that is well-known within its niche will naturally appear across multiple AI platforms without needing to optimise for each individually.
What is Speakable schema and does it help AEO?
Speakable schema (speakable: SpeakableSpecification) is a structured data markup that tells Google's systems which specific sections of a page contain the most important, self-contained answers, intended for voice assistant playback. It now also serves as a signal for AI search systems to identify high-priority passages for citation. While not confirmed as a direct ranking factor for AI Overviews, adding Speakable to key passages is considered best practice for AEO.
Speakable schema was introduced by Schema.org as part of the News schema vocabulary, initially targeted at news publishers who wanted Google Assistant to be able to read key sections of their articles aloud in response to voice queries. It has since expanded beyond news and is now recommended for any publisher who wants to signal authoritative passage-level content to AI systems.
The schema works by pointing to specific CSS selectors or XPath expressions within a page, identifying the text blocks that are best suited for spoken delivery, and by extension, for isolated passage extraction:
{
"@type": "Article",
"speakable": {
"@type": "SpeakableSpecification",
"cssSelector": [
".answer-box p",
"h1",
".faq-answer p"
]
}
}
The sections you mark with Speakable should share several characteristics: they should be self-contained (understandable without surrounding context), clearly written (no jargon, acronyms, or shorthand), directly answering a specific question, and no longer than 200–250 words. This maps closely to what AI retrieval systems look for when selecting passages for citation.
There is ongoing debate in the SEO community about whether Speakable schema has a measurable effect on AI Overview citation rates. Google has not made an explicit statement linking the two. However, the selection criteria for Speakable-worthy passages align closely with the selection criteria for AI citation-worthy passages, so even if the schema markup itself is not a signal, the practice of writing clearer, more extractable passages demonstrably improves AI search performance.
Add Speakable to 3–5 key passages per page: the hero summary, the answer box at the start of each major section, and the key conclusion. Avoid marking every paragraph, selectivity signals to Google's systems that the marked passages are genuinely the most important. Use the CSS selector method rather than XPath for easier maintenance.
For best results, combine Speakable schema with other AEO-oriented structured data. An Article + FAQPage + Speakable combination gives AI systems multiple signals about your content's structure, intent, and key passages. This multi-schema approach is what professional AEO implementation looks like in 2025.
What does a complete AEO content strategy look like?
A complete AEO strategy combines five pillars: (1) topical authority mapping, building a comprehensive cluster of content covering every sub-question in your niche; (2) passage-level optimisation, structuring each section with a direct answer box, clear heading, and self-contained paragraphs; (3) structured data, implementing Article, FAQPage, HowTo, and Speakable schema; (4) E-E-A-T signals, author pages, source citations, original research; and (5) distribution, earning brand mentions and citations from high-authority external sources.
A well-executed AEO strategy is not a set of individual tactics, it is a systematic content architecture designed from the ground up to be answer-engine-ready. Here is how to build one.
Pillar 1: Topical Authority Mapping. Start by mapping the complete question space around your core topic using a tool like the PAA Extractor. Extract every "People Also Ask" question for your head term and its closest related queries. Cluster these questions by intent and sub-topic. Every cluster that has significant search volume should become its own page. Your pillar hub page (like this one) should link to all cluster articles, and all cluster articles should link back. This architecture makes topical authority legible to both Google's crawlers and AI retrieval systems.
Pillar 2: Passage-Level Optimisation. Every section of every page should follow the same structure: an H2 phrased as the question the section answers, followed immediately by a direct answer box (1–3 sentences that can stand alone as a citation), followed by supporting detail. Never bury the answer. Write as if the AI system will extract only your first two paragraphs, because frequently, it will.
Pillar 3: Structured Data. At a minimum, every informational page should carry Article schema with author, publisher, datePublished, and dateModified fields. Add FAQPage schema for any Q&A content. Add Speakable to your key passages. Add HowTo for any step-by-step content. Use Google's Rich Results Test to validate your implementation before publishing.
Pillar 4: E-E-A-T Signals. Build an author page for every named contributor, linking to their professional profile, relevant credentials, and published work. Add "About This Article" sections that explicitly state the author's relevant experience. Cite your sources with inline links to primary research. Include original data where possible, your own surveys, case studies, test results, as original data is cited by AI systems at a significantly higher rate than secondary sources.
Pillar 5: Authority Distribution. None of the above technical work matters if your domain is not trusted by AI systems. Trust is built through external signals: being cited by high-authority publications, earning editorial links, appearing in Wikipedia (or being the subject of Wikipedia citations), and having your brand mentioned in diverse, credible contexts across the web. A link-building strategy focused on topically relevant, high-DR domains is the GEO component that amplifies all your on-page AEO work.
Ready to Measure Your AEO Readiness?
Get a detailed audit of any page across 18 AEO signals, structured data, passage quality, E-E-A-T, and more, in under 30 seconds.
Measurement is the final component of a complete AEO strategy. Track your AI Overview citation rate for target keywords using the AI Overview Checker. Monitor Perplexity and ChatGPT manually by querying your target keywords monthly and recording which sources they cite. Track brand mentions using Google Alerts and tools like Brand24. Over a 6–12 month window, you should see a clear correlation between topical authority growth and citation frequency across all major AI platforms.
AEO is a long game. The sites that dominate AI search citations today are predominantly those that have spent years building genuine topical authority, through consistent, high-quality content creation and systematic link-building. The good news is that the compounding effect of topical authority means every piece of cluster content you publish, every structured data implementation you make, and every high-authority mention you earn makes your next citation marginally more likely. Start now, measure consistently, and iterate.
Frequently Asked Questions About AEO & AI Search
SEO (Search Engine Optimisation) focuses on ranking in traditional blue-link search results, where users see a list of pages and choose which to click. AEO (Answer Engine Optimisation) focuses on getting your content directly cited inside AI-generated answers, in Google AI Overviews, Perplexity, ChatGPT, and Gemini, where the AI system synthesises an answer on your behalf and may or may not provide a clickable citation.
The key practical difference: in SEO, your goal is to appear on the list. In AEO, your goal is to become the answer. The tactics overlap significantly, both require high-quality content, strong E-E-A-T signals, and technical health, but AEO adds specific requirements around passage-level formatting, structured data, and question-first content architecture.
Google AI Overviews primarily draws from pages that already rank in the top 5–10 results for the query and closely related sub-queries. From that pool, it selects sources based on: passage-level relevance to the specific sub-question being answered, E-E-A-T (Experience, Expertise, Authoritativeness, Trustworthiness) signals, structured data such as FAQPage and Speakable schema, and content freshness for time-sensitive queries.
Ranking well in organic search is a prerequisite, not a guarantee, of AI Overview inclusion. Pages that rank well but have unclear structure, no structured data, and buried answers are often passed over in favour of lower-ranked pages with better passage-level answers.
Query fan-out is the process by which AI search systems decompose a single user query into multiple sub-queries and retrieve answers for each independently. For example, "best protein powder for building muscle" might fan out into sub-queries about protein types, dosage, timing, and brand quality.
For AEO, this matters enormously because content that covers the breadth of these sub-queries, rather than just the headline topic, is far more likely to appear in synthesised AI answers. A page addressing only the head question will be passed over for sub-questions answered by more comprehensive competitors. Mapping the full fan-out using PAA data and comprehensively covering all sub-questions is the core content strategy implication of query fan-out.
Speakable schema was originally designed for voice assistants to identify the most important passages of an article to read aloud. It now serves a secondary function in AI search: it signals to Google's systems which text blocks contain concise, authoritative, self-contained answers.
Google has not confirmed Speakable as a direct ranking factor for AI Overviews. However, the characteristics of good Speakable passages, short, direct, self-contained, clearly answering a question, align exactly with what AI retrieval systems prefer. Implementing Speakable correctly forces you to write better AEO-optimised passages, which benefits citation rates regardless of whether the schema markup itself is read as a signal.
The key architectural difference: Google AI Overviews operates on Google's pre-built search index, while Perplexity retrieves content in near-real-time using its own crawler (PerplexityBot). This means Perplexity is more sensitive to content freshness and can cite pages that are newly published, whereas Google AI Overviews are slower to reflect recent content changes.
Perplexity also places less weight on E-E-A-T and structured data signals than Google does. Its source selection skews toward pages that rank well in Google and Bing, with strong preference for content that is factual, data-rich, and clearly structured. Perplexity is also more transparent about its sources, making it easier to monitor your citation rate by querying your target keywords directly.
GEO (Generative Engine Optimisation) is the practice of optimising content to be cited, quoted, or synthesised by large language models across all AI-powered surfaces, including ChatGPT, Perplexity, Google Gemini, Claude, and Microsoft Copilot. Where AEO focuses on answer engines directly attached to search (Google AI Overviews, voice search), GEO is the broader discipline covering all generative AI surfaces.
GEO strategies include building brand mentions across authoritative external sites (to appear in LLM training data and retrieval indexes), writing in a citable, attribution-friendly style, creating original research and data that AI systems prefer to cite, and ensuring all major AI crawlers are permitted in your robots.txt. GEO operates on a longer time horizon than AEO because training data inclusion is a slow, accumulative process.
Explore the Full AI Search Cluster
Every sub-topic covered in detail, pick the guide that matches your next priority.
Google AI Overviews: How to Get Cited in 2026
AI Overviews appear on 14% of queries and cut CTR by up to 60%. The content signals, entity factors, and tracking approach that determine whether you get cited.
Read article →