How to tell whether generative engine optimization actually worked

The awkward fact underneath every AI-visibility report is that the thing being measured moves on its own. Ask an engine the same question two days running and it will cite a substantially different set of sources. That makes most of what gets sent to clients unfalsifiable, and it makes a small number of specific demands worth making before you sign anything.

Published 2026-09-074 sources: two papers, two of our own studies5 min read

Key takeaways

  • The sources an engine cites overlap only 34–42% between consecutive days.[1] Two readings can differ by more than half with nothing changed.
  • So a single before-and-after screenshot cannot distinguish work from noise, whoever produced it, including us.
  • Demand the same prompt set, run repeatedly, reported per engine with the run count. Those four together are what make a claim checkable.
  • A percentage with no run count attached is not a result. It is a reading.

Why one reading proves nothing#

The measurement problem here is not sloppiness, it is the surface. A study of run-to-run stability found that the set of sources an engine cites overlaps only 34–42% between consecutive days for the same prompt, and brand sets 45–59%.[1] The authors put the conclusion plainly: visibility has to be characterised as a distribution, not a single-point outcome.

That has a direct consequence for reporting. If roughly half the cited sources turn over between Tuesday and Wednesday, then a report comparing one day in March against one day in June is mostly measuring which day it landed on. Any number produced that way can be moved by picking different dates, which is why we treat single-run before-and-afters as unfalsifiable rather than merely weak.

The five things to demand#

1. The prompt set, in writing, before work starts. Which buyer questions are tracked, how many, and whether they change during the engagement. A prompt set edited mid-flight makes before and after incomparable, and it is the easiest way to manufacture an improvement without doing any work.

2. Runs per prompt, per engine, with the count in the report.Not “we check regularly”. A number. Repeated sampling is the only thing that separates a real move from the variance above, and the count is what lets you judge whether it did.

3. Per-engine results, never a blended score.Engines disagree sharply. In our own CRM study one brand was named in 100% of grounded Gemini responses and 90% of Perplexity's, while another sat at 70% on ChatGPT and 40% on both others.[3] A single blended figure hides exactly the differences you would act on.

4. The same prompts before and after. Obvious, routinely not done. If the after-report contains prompts the before-report did not, the comparison is decorative.

5. What changed on the site, itemised. Measurement without shipped work explains nothing. Ask which pages were written, what was fixed, and where coverage was placed, so a movement can be attributed to something rather than asserted.

What a weak report looks like#

A single percentage with no denominator.“Visibility up 40%” is not a result without the prompt count, the run count and the engine breakdown behind it.

Screenshots of chat windows. One conversation is one sample from a distribution that turns over daily. A screenshot proves the answer existed once.

A score that only goes up. Real measurement produces down weeks. Our own study had one: Perplexity citations ran 50.4%, 60.7%, 66.3%, then eased to 64.8% in the final bucket, on 1,558, 1,706, 1,708 and 650 runs.[4] A report with no down weeks is either lucky or curated.

Tactics sold on correlation as though they were causes.The one controlled experiment in this field tested nine methods and found its top three — citing sources, adding quotations, adding statistics — produced a 30–40% relative improvement on its primary metric.[2] Most other advice in circulation is observational, and should be sold that way.

Should my SEO agency already handle this?#

Some do it well. The question worth asking is not whether they offer it but whether they measure it, because the two skills separate: ranking is a position you can look up, and citation is a distribution you have to sample. An agency that reports AI visibility from a rank tracker is reporting the wrong surface.

Use the five demands above as the test. If your existing agency can answer all five, there is no reason to move. If the answer to the run count is a shrug, you now know what to ask for. Our own protocol, including the sampling schedule and the confidence intervals, is public on the methodology page, and what this costs is written up separately.

Want a measurement you can check? Ours ships with the run counts attached.

Sources

  1. “Don't Measure Once: Measuring Visibility in AI Search” — arXiv:2604.07585. Source sets overlap 34–42% between consecutive days, brand sets 45–59%; visibility should be treated as a distribution.
  2. Aggarwal et al., “GEO: Generative Engine Optimization” — arXiv:2311.09735, KDD 2024. Nine methods on a 10,000-query benchmark drawn from 25 domains; top three at 30–40% relative improvement on Position-Adjusted Word Count.
  3. Noetio CRM AI Visibility study— 20 brands, 60 queries across ChatGPT, Gemini and Perplexity, 50 search-grounded and scored, August 2026. Per-engine rates and the validity note are published with it.
  4. Noetio own-domain citation study— a fixed 122-question set tracked on noetio.com from 2026-07-14 to 2026-08-05, weeks starting 13 Jul, 20 Jul, 27 Jul and 3 Aug. The final week is a partial bucket, 650 runs against roughly 1,700 in each of the first three. No control group.

Two papers with disclosed method, plus our own study, whose dataset and validity note are published alongside it.