Companion piece to our first-party dataset: 5,051 AI-search citation polls in 77 days. That post is the findings. This one is the method — so you can run it yourself.
How do I track AI search visibility? Build a list of 15–70 real buyer questions, ask each one to ChatGPT, Gemini and other AI engines on a fixed schedule (weekly by hand or daily via API), record whether your site is cited and which URL, then chart your share of answer per engine over time.
That’s the whole method in one paragraph. The rest of this post is the detail: how to choose the questions, how often to poll, what to log, and what the numbers actually tell you — because we’ve been running exactly this on our own pipeline since May, and most of what we believed before measuring turned out to be wrong.
Why you can’t just ask ChatGPT once and screenshot it
The most common way businesses “check” their AI visibility is to type one question into ChatGPT, see whether they come up, and either celebrate or panic. Both reactions are premature. AI engines vary their answers day to day — the same prompt that cites you on Monday can skip you on Wednesday with no change on your end. In our own polling, once ChatGPT first cited us on a prompt, it kept citing us in only about 44% of subsequent polls of that prompt (median), while Gemini held around 70%. A one-off check can’t distinguish “we’re visible” from “we got lucky today.”
So visibility has to be measured the way you’d measure anything noisy: same questions, repeated on a schedule, logged consistently. Here’s the six-step version we use.
The 6-step method
Step 1: Build a prompt set of 15–70 real buyer questions
Your prompt set is the foundation, and the biggest mistake is filling it with brand searches (“what is [your company]”) instead of buying-intent questions. Nobody who doesn’t already know you asks an AI about you by name. Write prompts the way a buyer who has never heard of you would phrase them:
- Buying-intent phrasing: “best appointment setting agency for financial advisers”, “who should I hire to book sales meetings for my mortgage brokerage” — not “appointment setting definition”.
- Your verticals, spelled out: one or two prompts per niche you serve. Generic prompts get generic answers; specific prompts are where smaller businesses actually get cited.
- Competitor comparisons: “X vs Y”, “alternatives to [category leader]”. These prompts almost always produce a list of names, so there’s a slot to win.
- Local variants if geography matters: “…in Melbourne”, “…Australia”.
Fifteen prompts is enough to start and matches what you can check by hand in under an hour. We keep a registry of 68 — 33 polled on the LLM engines (ChatGPT and Gemini) and the full set on DuckDuckGo’s answer pages; past 70 or so you’re adding polling cost faster than insight. Freeze the list once it’s written — if you keep editing prompts, you can’t compare week 1 with week 8.
Step 2: Poll each engine on a fixed cadence
Pick a schedule and keep it. Two workable tiers:
- Manual (free): once a week, same day, run every prompt through ChatGPT and Gemini in a fresh session (logged out or a clean profile, so personalisation doesn’t skew results). Thirty to sixty minutes for 15 prompts.
- API (small cost, much better data): hit each engine’s API daily or every second day with the same prompts. We poll ChatGPT and Gemini every two days and run a daily check on DuckDuckGo’s answer pages. You don’t need real infrastructure — a scheduled script and a database (or even a CSV it appends to) is enough.
Cadence matters more than volume. Ten prompts polled weekly for three months beats seventy prompts polled once, because the entire point is the trend line.
Step 3: Record cited / not cited — and which URL
For every poll, log four things: date, engine, prompt, and whether your domain appeared in the answer or its citations. If it did, also record which URL was cited. That last field is the one people skip and later regret: knowing that engines cite your case-study page but never your homepage tells you exactly what kind of content earns citations in your niche. A spreadsheet with one row per poll is genuinely fine — resist the urge to build a dashboard before you have data worth dashboarding.
One definitional note: “cited” should mean your domain appears in the answer or its source list for a question you didn’t ask about yourself. Being mentioned when someone asks about you by name is table stakes, not visibility.
Step 4: Track share of answer over time, per engine
Your headline metric is simple: of the prompts you polled this week, what percentage cited you, per engine? That’s your share of answer — we’ve written a full explainer on the share-of-answer metric if you want the deeper version. Report it per engine, never blended: the engines behave so differently that an average hides everything useful. Plot it weekly. Flat is information. A slow climb after publishing new pages is information. A one-week dip is usually just noise.
Step 5: Segment — persistence, and single-engine vs multi-engine
Once you have a few weeks of rows, two cuts are worth the effort:
- Persistence: after an engine first cites you on a prompt, what fraction of later polls of that prompt still cite you? This tells you whether citations “stick” or flicker. In our data the gap was large — roughly 44% persistence on ChatGPT vs roughly 70% on Gemini.
- Engine overlap: for each prompt where you’ve ever been cited, how many engines cite it? Across our 5,051 polls, 74% of cited prompts were cited by one engine only. If your numbers look similar, “AI visibility” isn’t one thing for you either — it’s a per-engine game, and winning ChatGPT says little about Gemini.
Those are our numbers, on our domain, in our niche — treat them as directional, not universal. The point of running this yourself is to get your numbers.
Step 6: Act on the gaps
The un-cited prompts are your content roadmap. For each prompt where you’re never cited, check: do you have a page that directly answers that exact question, with the answer stated plainly near the top? Usually the honest answer is no — you have a services page that’s about the topic, not a page that answers the question. Build the page, then keep polling; the log tells you whether it worked, typically within a few weeks. And keep the pages that do get cited fresh — we’ve covered how often to update content for AI search separately, but stale pages are one of the few gap causes you can fix in an afternoon.
DIY spreadsheet vs API polling vs paid AEO tools
| Approach | Cost | Effort | Data quality | Best for |
|---|---|---|---|---|
| Manual + spreadsheet | Free | 30–60 min/week for ~15 prompts | Weekly resolution; some personalisation risk; fine for trends | Any business starting out; validating whether this matters for you |
| DIY API polling | API usage fees (typically small at this scale) + a few hours of one-off setup | Near-zero once running | Daily/2-daily resolution; consistent, personalisation-free; enables persistence analysis | Anyone with basic scripting help; what we run ourselves |
| Paid AEO tools (e.g. Otterly.AI, Peec AI, Profound) | Otterly from US$29/month for 15 prompts; Peec and Profound don’t publish pricing | Minimal — set prompts, read dashboards | Multi-engine coverage out of the box (ChatGPT, Perplexity, Gemini, AI Overviews and more, varying by tool), competitor benchmarking | Teams tracking many brands or engines, or with no one to run scripts |
Honest take: the tools are legitimate and the entry-level pricing is reasonable, but don’t buy one before you’ve done four manual weekly checks. The manual pass forces you to write good prompts and read real answers — which is most of the skill — and tells you whether your niche even shows up in AI answers yet.
What this method won’t tell you
Matching the honesty of the dataset post: this measures citations, not customers. A citation is not a click, and a small prompt set is a sample, not a census. Expect sampling noise — one-day flickers where an engine drops you and picks you back up are normal, which is exactly why you look at persistence over weeks instead of any single poll. And your results will be specific to your niche and domain; ours are published here as a reference point, not a benchmark.
FAQ
How many prompts do I need to track AI search visibility?
Fifteen well-chosen buyer questions is a workable minimum — enough to see per-engine trends without making weekly checks a chore. We run 68 and consider 70 a sensible ceiling for a single brand: beyond that, extra prompts mostly add polling cost, not insight. Prompt quality (buying intent, your verticals, comparisons) matters far more than count.
Is AI search visibility worth tracking if AI answers send so little traffic?
Yes — precisely because the traffic model is changing. Pew Research Center’s analysis of real browsing data from 900 U.S. adults found users clicked a traditional result in only 8% of Google visits when an AI summary appeared, versus 15% without one. When fewer people click through, being the name inside the answer is the visibility that’s left — and you can’t manage it without measuring it.
Why do my results change from day to day?
Because the engines genuinely vary their answers — model updates, retrieval changes, and plain non-determinism all cause flicker. In our 77-day dataset, ChatGPT re-cited an already-won prompt in only about 44% of later polls, while Gemini held about 70%. Day-to-day movement is noise; judge yourself on multi-week trends per engine, not on any single answer.
Want the instrumented version without building it?
Everything above is genuinely DIY — a spreadsheet and an hour a week gets you real data. If you’d rather your lead flow came already measured this way, that’s what we do: 50,769+ AI-booked sales appointments since 2017 and 1M+ leads generated, on a pay-per-result basis. Book a call and we’ll show you what your share of answer looks like today.
