Research
Answer engine optimization, explained without the acronym soup
Answer engine optimization is the work of being quoted inside an AI-generated answer rather than listed underneath it. The field arrived with three competing acronyms and a great deal of advice that has never been tested. This separates the terms, explains the mechanism that decides whether you get cited, and ranks the levers by how much evidence actually stands behind each one.
Key takeaways
- As sold today, AEO, GEO and LLMO describe the same work. The acronym usually tells you who is selling, not what is done.
- Cite Sources, Quotation Addition and Statistics Addition measured a 30–40% relative improvement in the one controlled experiment.[1]
- Only 38% of AI Overview citations come from the organic top ten, so ranking first no longer guarantees being quoted.[3]
- The widely repeated “FAQ schema gives +44% citations” figure has no traceable source,[6] and the nearest controlled test found no uplift from schema on two engines and a decline on the third.[4]
What answer engine optimization means#
Answer engine optimization is the practice of making a page more likely to be used, and credited, when an engine writes an answer instead of returning a list of links. The unit of success is not a position. It is whether a sentence of yours ends up inside the answer with your name attached.
AEO, GEO and LLMO are the same job#
Generative engine optimization (GEO) is the term the academic literature uses. Answer engine optimization (AEO) predates generative systems — it came out of the featured-snippet and voice-answer era, and in its original sense covers answer surfaces that involve no language model at all. LLM optimization (LLMO) is the newest. In current usage the three are sold as the same activity: influencing what a retrieval-plus-generation system quotes. That overlap is a fact about the market, not a finding from any of the research below.
The practical note is about search behaviour rather than substance. In our own pull of Google Ads keyword-planner data for the United States, “generative engine optimization” and “answer engine optimization” both carry real monthly search volume, while the bare abbreviation “GEO” collides with geography, a prison operator and a Pakistani news channel.[7] If you are writing to be found, spell the term out.
How an answer gets assembled#
Some engines run their own web index — Google and Brave among them — while others send a rewritten version of your question to an index they do not own. Either way the shape is the same: retrieve a set of candidate pages, then generate an answer from the passages pulled. Two consequences follow.
First, inclusion in the underlying index is a hard gate. A page absent from the index an engine reads cannot be cited regardless of quality. Second, what gets quoted is a passage, not a page. Neither point is a measured finding — they follow from how retrieval-plus-generation works — but they explain why a self-contained paragraph is easier to lift than an argument that only makes sense read in full.
Ranking and citation have measurably separated. Ahrefs found only 38% of Google AI Overview citations now come from pages in the organic top ten, down from about 76% in July 2025.[3]
The levers, ranked by evidence#
The one designed experiment with published per-tactic effect sizes we could find is GEO: Generative Engine Optimization (Aggarwal et al., KDD 2024), which evaluated nine methods on a benchmark of queries across multiple domains.[1]Its top three — Cite Sources, Quotation Addition and Statistics Addition — achieved, in the paper's words, “a relative improvement of 30-40% on the Position-Adjusted Word Count metric and 15-30% on the Subjective Impression metric.” The paper also reports that efficacy varies by domain.
Observational work on live engines points in a compatible direction. The GEO-16 audit of 1,100 URLs behind 1,702 citations across three production engines reported an odds ratio of 4.2 for citation against its overall composite quality score.[2] Its strongest per-signal correlations sat in metadata and freshness, semantic HTML and structured data, the same signals the controlled schema test below found no citation lift from. Its scope is 70 product-intent prompts against English-language B2B SaaS pages, and it is observational, so read it as association worth testing, never as evidence those signals cause citation.
Keep the two apart. The experiment tested changes to visible text and found that checkable material — sources, quotes, numbers — raised a page's visibility inside the generated answer. GEO-16 is observational and found its strongest associations elsewhere, in metadata and freshness, semantic HTML and structured data. One is a tested lever; the other is a correlation we do not sell.
Claims with nothing behind them#
“FAQ schema gives you 44% more citations.” This figure circulates widely and is usually attributed to BrightEdge. It does not trace to any BrightEdge publication. An independent review that went looking for the study reached the same conclusion and found where the number came from: BrightEdge did publish a 44% statistic, but it is about AI Overviews being 44% more likely to criticise brands than ChatGPT. The figure changed hosts.[6]The nearest controlled test — 1,885 pages that added schema against ~4,000 matched controls — found no uplift it could distinguish from zero on ChatGPT or AI Mode, and an estimated −4.6% on Google AI Overviews.[4]That study pooled schema types rather than isolating FAQ specifically and says individual types may differ, so read it as “the 44% figure has no source, and the best available test found no uplift”, not as a measurement of FAQ schema alone.
“Repeat the target phrase to signal relevance.”Keyword Stuffing was one of the nine methods tested. The paper's conclusion: “simple methods like Keyword Stuffing traditionally used in SEO don't perform well.”[1]
“We improved your AI visibility by X%.”Without a control and repeated runs this is unfalsifiable. The sources an engine cites overlap by only 34–42% between consecutive days,[5] which is more than enough run-to-run movement for a single before-and-after reading to show a swing that nothing on the site caused.
Related: what the research measures about LLM SEO.
Want the evidence-ranked fix list for your site?
Sources
- Aggarwal et al., “GEO: Generative Engine Optimization” — arXiv:2311.09735, KDD 2024. 10,000 queries drawn from 25 domains.
- GEO-16 framework— arXiv:2509.10762. 70 prompts, 1,702 citations, 1,100 URLs, three production engines.
- Ahrefs, “38% of AI Overview Citations Pull From The Top 10” — 863K keyword SERPs, 4M AI Overview URLs, March 2026. Vendor study.
- Ahrefs schema study— matched difference-in-differences, 1,885 treated pages vs ~4,000 controls, Aug 2025–Mar 2026. Vendor study.
- “Don't Measure Once” — arXiv:2604.07585, on run-to-run variance in AI answer sources.
- Cheung, “Should You Bother With Schema Markup for AI Search?” — independent review of ten direct sources on schema and AI citation, ordered by methodological control. Traces the BrightEdge 44% figure to a real statistic about a different subject.
- Noetio, Google Ads keyword-planner data for the United States, pulled September 2026 via the DataForSEO search-volume endpoint. Our own measurement, not a third-party study.
Three of these are papers with disclosed method, two are vendor studies, one is an independent evidence review, and one is our own keyword pull; each is labelled as such. Where a widely repeated figure has no source, that is stated rather than omitted.