SaaS SEO is now two jobs.
Most agencies still sell one.

Ranking for “best CRM for small teams” and being the answer when a buyer asks ChatGPT that same question are different problems with different inputs. We do the second one, and we measure it in the only way that means anything: the same prompts, run repeatedly, per engine, before and after.

Book a call · 15 min →See a category we measured →

We scored 20 CRM brands on how often three AI engines name them, and published the whole thing. HubSpot CRM is named in 94% of grounded responses. 7 of the 20 were never named once. The median score across the category is 3.1 out of 100.

Several of those invisible brands are established companies with real revenue and real Google rankings. That is the gap this work closes, and the full study with its method is public, including what we excluded and why.

Method note: every scored response is search-grounded. OpenAI runs gpt-5.4-nano with the web_search tool forced, Gemini runs gemini-3.6-flash with google_search, Perplexity runs sonar. Gemini declined to search on 10 of its 20 calls and those responses are excluded from the denominator rather than counted as an absence, so Gemini rates rest on 10 grounded responses against 20 each for OpenAI and Perplexity. Combined scores pool all 50 grounded responses rather than averaging the three engines equally, which means Gemini carries 20% of the weight in a combined score, not a third. This run replaces a June 2026 study whose OpenAI and Gemini legs were collected without a search tool and therefore measured model memory, not AI search.

01

We ask the engines what a buyer asks

30 category prompts per engagement, in the words a buyer uses: best tool for X, alternatives to Y, is Z any good. Run repeatedly, per engine, not once.

02

You get the gap, not a score

Which prompts name you, which name a competitor instead, and what those answers cite. A visibility score is a thermometer; the fix list is the work.

03

We ship the fixes

Evidence-backed pages, rendering fixes so crawlers read full HTML, and coverage on the third-party pages the engines actually quote. Every change is a pull request or CMS draft you approve.

04

Then we re-run the same prompts

Same prompt set, same engines, before and after. Repeated runs, because a single reading cannot tell a real change from normal variance.

That we can guarantee a citation. Nobody controls what an engine says, and the sources behind an answer overlap only 34–42% between consecutive days, so any agency quoting a single before-and-after percentage is quoting noise. What we commit to is a fix list with named owners, the fixes shipped, and the same prompts re-run so you can see the change or the absence of one. The reasoning is in how to tell a good agency from a reseller and the protocol is on our methodology page.

Book a call · 15 min →What the audit covers →