What LLM SEO is, and what the research actually measures

LLM SEO is the work of getting a language model to cite you when it answers a buyer's question. It is a young field with a lot of confident advice and very little controlled evidence. We could find one designed experiment with published effect sizes, a handful of large-scale observational studies, and a long tail of vendor claims that do not survive checking. This is what each of them actually found.

Published 2026-08-285 sources, papers and large-scale studies7 min read

Key takeaways

  • The controlled study we could find measured a 30–40% relative improvement for Cite Sources, Quotation Addition and Statistics Addition on its primary metric.[1]
  • Keyword stuffing did not perform well in that study. Repeating the phrase is not the lever.[1]
  • Schema markup was tested on 1,885 pages against 4,000 matched controls. Citations moved less than measurement error on two engines and down 4.6% on the third.[4]
  • Day-to-day source overlap in AI answers is 34–42%, so a single before-and-after reading cannot tell a real effect from noise.[3]

What LLM SEO means#

LLM SEO, also written as LLM optimisation, is the practice of making a page more likely to be retrieved and cited by a large language model when it answers a question. The same work travels under generative engine optimisation (GEO) and answer engine optimisation (AEO). The names differ; the mechanism is the same one: an engine retrieves a set of sources, then writes an answer that quotes some of them.

The distinction that matters is not the label. It is that you are optimising to be quoted inside an answer, not to occupy a position in a list of links.

How it differs from search engine optimisation#

Classic SEO earns a ranked position; the click follows. An LLM does not hand out positions. It assembles an answer and attributes parts of it. That changes which properties of a page matter.

The gap is measurable. Ahrefs found that only 38% of Google AI Overview citations now come from pages ranking in the organic top ten, down from about 76% in July 2025, across 863,000 keywords.[5] Ranking well and being cited well have come apart.

What the research measures#

The strongest evidence is a designed experiment rather than a correlation study. GEO: Generative Engine Optimization (Aggarwal et al., KDD 2024) evaluated nine methods on GEO-bench, a benchmark of user queries across multiple domains. In the paper's own words, the top three — Cite Sources, Quotation Addition and Statistics Addition — “achieved a relative improvement of 30-40% on the Position-Adjusted Word Count metric and 15-30% on the Subjective Impression metric.”[1]The paper adds that its best methods “improve upon baseline by 41% and 28%” on those two metrics respectively. The paper also notes efficacy varies by domain.

A second study looked at production engines rather than a research pipeline. The GEO-16 framework audited 1,100 URLs behind 1,702 citations across Brave, Google AI Overviews and Perplexity, and reported an odds ratio of 4.2 for citation against its overall composite quality score.[2] Its strongest per-signal correlations were metadata and freshness, semantic HTML and structured data. The odds ratio belongs to the composite score, not to any one of those signals. Its scope is narrower than the headline suggests: 70 product-intent prompts against English-language B2B SaaS pages, and the design is observational, so it establishes association rather than cause.

The two support different things, and it is worth keeping them apart. The experiment is evidence that named sources, real quotes and actual numbers raise how much of the answer is drawn from the page, and how prominently. GEO-16 found its strongest associations somewhere else: metadata and freshness, semantic HTML, structured data. What they have in common is only the direction of travel, not the mechanism. Treat the first as the tested lever. The second is a correlation from one observational, single-vertical study whose signals the controlled schema test below found no citation lift from, so do not sell metadata, freshness or structured data as levers.

What does not work#

The same experiment tested Keyword Stuffing — adding more of the query's words to the page, the classic SEO move. The paper's finding is blunt: “simple methods like Keyword Stuffing traditionally used in SEO don't perform well.”[1] If a piece of advice amounts to saying the phrase more often, the one experiment that measured it found nothing there.

Schema markup is the more surprising result. Ahrefs tracked 1,885 pages that added schema against roughly 4,000 matched controls between August 2025 and March 2026. AI citations barely moved: +2.2% on ChatGPT and +2.4% on AI Mode, both statistically indistinguishable from zero, and −4.6% on Google AI Overviews, an estimated decline rather than a null.[4]Two caveats the study itself carries, in its own words: it “pooled all schema types together” rather than isolating any one, and it studied “pages that were already being cited heavily by AI”. So it is evidence against schema as a large citation lever. It is not proof that schema does nothing, and on two of the three engines the honest reading is that the study could not separate a small positive effect from zero. It is SEO hygiene, not a citation lever.

Why one measurement lies#

AI answers are not stable. Don't Measure Once found overlap in the sources an engine cites between consecutive days of only 34–42%.[3] Two readings a day apart can differ by more than half without anything on your site changing.

The practical consequence is blunt: a single before-and-after screenshot cannot distinguish a real improvement from normal variance. Any claim of a percentage lift that rests on one measurement, including any agency's, is indistinguishable from noise. Repeated runs against the same prompt set, reported per engine with a confidence interval, are the minimum that means anything.

Related: the same work under a different acronym.

Want to know which of these your site is missing?

Sources

  1. Aggarwal et al., “GEO: Generative Engine Optimization” — arXiv:2311.09735, KDD 2024. 10,000 queries drawn from 25 domains. Designed experiment with per-tactic effect sizes. Measured on a research retrieval pipeline, validated against Perplexity, so treat it as a strong proxy rather than a direct production measurement.
  2. “AI Answer Engine Citation Behavior: the GEO-16 Framework” — arXiv:2509.10762. 70 prompts, 1,702 citations, 1,100 URLs audited across three production engines. Observational, not causal.
  3. “Don't Measure Once: Measuring Visibility in AI Search” — arXiv:2604.07585. Establishes that single-run measurement of AI visibility is unreliable.
  4. Ahrefs, “We Tracked 1,885 Pages Adding Schema”Matched difference-in-differences against ~4,000 controls, Aug 2025–Mar 2026. Vendor study, disclosed method.
  5. Ahrefs, “38% of AI Overview Citations Pull From The Top 10” — 863K keyword SERPs, 4M AI Overview URLs, March 2026. Vendor study.

Each claim names its source with sample size and method; papers are linked directly. Three of the five are papers with disclosed method; the two vendor studies are labelled as such.