Let's grow your business. 2 new positions just opened Monday, 3 August. Book a free call today.
Uncategorised 8 min read

We Ran 5,051 AI-Search Citation Polls in 77 Days. Here’s What Actually Gets a B2B Site Cited

Most advice about getting cited by AI search engines is written by people who have never measured a citation. We have. Since 19 May 2026 we’ve been running an automated polling rig against ChatGPT, Gemini and DuckDuckGo — 5,051 polls over 77 days, against 68 real buyer prompts in our own niche — logging every time our domain does (or doesn’t) appear. This post is the raw findings: which engines cite most readily, how often a citation survives the next poll, and how little the three engines agree with each other.

The short answer: across 5,051 citation polls over 77 days, 74% of the prompts our domain was ever cited on were cited by only ONE of the three engines — a single-engine strategy is blind to most of the surface. DuckDuckGo was the easiest win (ever-cited on 66% of prompts) and ChatGPT by far the hardest (18%). And a citation is not a ranking: once first cited on a prompt, ChatGPT kept citing us in only ~44% of that prompt’s polls from the first citation onward (median) versus ~70% for Gemini. ChatGPT citations flicker; Gemini citations stick.

Everything below is first-party data from our own polling database — one domain (ours), one niche (lead generation and appointment setting), 68 prompts. That makes it directional rather than universal, and we’ll flag the caveats as we go. But it’s also data nobody can quote from a vendor deck, because until now it didn’t exist in public.

What we ran: 5,051 polls, 3 engines, 77 days

The rig polls three AI/answer surfaces against a registry of 68 real buyer prompts — the questions actual B2B buyers ask, in Australian and US variants, in our niche:

  • ChatGPT (browse-enabled) — 451 polls, every 2 days. A “citation” means our domain is linked as a source in the browsed answer.
  • Gemini (search-grounded) — 411 polls, every 2 days. A “citation” means our domain appears as a grounding source for the answer.
  • DuckDuckGo answer-page SERPs — 4,189 polls, on a daily cron. Here “citation” means appearing on the answer page/SERP at all — a lower bar than an in-answer citation, and worth keeping in mind when comparing the numbers.

Every poll is logged with a timestamp, engine, prompt and hit/miss flag, which is what lets us measure not just whether we get cited but whether a citation survives. If you’re new to treating citations as a metric, our share-of-answer measurement guide covers the framing.

Finding 1: the three engines barely agree with each other

Of the 68 prompts in the registry, our domain has been cited somewhere on 47 of them. But the overlap between engines is startlingly thin:

  • 35 of those 47 prompts (74%) were cited by only ONE engine.
  • Just 5 prompts were cited by all three engines.
  • Every prompt ChatGPT ever cited us on, Gemini also cited (6 of 6) — but not vice versa. ChatGPT’s citation set was a strict subset of Gemini’s on our prompts.

The practical read: if you only track (or only optimise for) one engine, you’re blind to roughly three-quarters of the prompts where you’re actually winning or losing. The engines retrieve differently, ground differently and cite differently — we’ve documented the same pattern from the content side in our per-engine citation overlap breakdown.

Finding 2: DuckDuckGo is the easiest surface, ChatGPT the hardest

Ever-cited rate — the share of polled prompts where our domain has appeared at least once:

  • DuckDuckGo: 45 of 68 prompts (66%)
  • Gemini: 13 of 33 prompts polled on that engine (39%)
  • ChatGPT: 6 of 33 (18%)

Two things drive the gap. DuckDuckGo’s bar is lower by definition (answer-page presence, not in-answer citation). And ChatGPT is genuinely selective about what it will link — we’ve written before about what it takes to be in the slice ChatGPT actually cites. If your AEO reporting only screenshots ChatGPT, expect it to look grim long after the easier surfaces have started paying.

Finding 3: a citation is not a ranking — it can vanish next poll

This is the number almost nobody measures, because you can only see it with repeated polling of the same prompt. After a domain first gets cited on a prompt:

  • ChatGPT keeps citing it in only ~44% of that prompt’s polls from the first citation onward (median).
  • Gemini keeps citing it in ~70% of polls from the first citation onward.

In other words, ChatGPT citations flicker — a prompt you “won” on Monday is a coin flip by Wednesday — while Gemini citations stick. This has two consequences. First, a one-off “we’re cited in ChatGPT!” screenshot is close to meaningless; only a poll series tells you whether you hold the position. Second, flicker is exactly why refresh cadence matters: stale pages fall out of flickery answer sets first, which is the case we make in our content freshness for AI search guide.

The three engines side by side

  ChatGPT (browse-enabled) Gemini (search-grounded) DuckDuckGo answer page
Polls in 77 days 451 411 4,189
Poll cadence Every 2 days Every 2 days Daily cron
Prompts ever cited 6 of 33 (18%) 13 of 33 (39%) 45 of 68 (66%)
Citation persistence after first hit ~44% of polls from first citation onward (median) — flickers ~70% — sticks — (different bar, not directly comparable)
What “cited” means here Linked as a source in the browsed answer Appears as a grounding source for the answer Appears on the answer page/SERP — a lower bar

Methodology, briefly

Window: 19 May – 3 August 2026 (77 days), 5,051 polls total — counts as at the morning of 3 August 2026; the rig keeps polling, so totals tick up daily. Prompt registry: 68 real buyer prompts in the lead-generation / appointment-setting niche, in AU and US variants; 33 of the 68 are polled on the LLM engines. Cadence: DuckDuckGo daily via cron; ChatGPT and Gemini every 2 days. Logging: every poll is written to a local SQLite database with timestamp, engine, prompt and citation flags, so ever-cited rates and persistence are computed from the log, not recalled from memory. Persistence is measured per prompt as the share of that prompt’s polls, from the first citation onward (inclusive), in which our domain was cited; we report the median across ever-cited prompts. Caveats, stated plainly: this is one domain (ours), in one niche, across n=68 prompts — treat the numbers as directional, not universal. And the DuckDuckGo figure measures answer-page/SERP presence, which is a different (lower) bar than an in-answer citation on the LLM engines.

What this means for B2B marketers

1. Measure weekly, or you’re guessing. With ChatGPT holding a citation in only ~44% of a prompt’s polls once it first appears, any single check — positive or negative — is noise. A weekly poll series is the minimum instrument that can tell you whether AEO work is actually landing.

2. Don’t run a single-engine strategy. 74% of our cited prompts were cited by one engine only. Optimising for — or reporting on — ChatGPT alone means most of your wins and losses happen off your dashboard.

3. Sequence the engines by difficulty. In our data DuckDuckGo cited fastest and widest (66% of prompts), Gemini sat in the middle (39%) and held citations best (~70% persistence), and ChatGPT was the hardest door (18%) — but every prompt ChatGPT cited, Gemini had cited too. Winning Gemini first appears to be the path toward ChatGPT, not a detour.

4. The structural basics still set the odds. AirOps’ 2026 State of AI Search analysis found pages with sequential heading structures (H2>H3>H4) earn 2.8x higher citation odds than unstructured pages — consistent with what we see holding citations across engines.

5. The prize is worth the instrumentation. AI-search visitors are scarce but disproportionately valuable: Ahrefs’ first-party data shows AI search delivering just 0.5% of its traffic but 12.1% of its signups — a 23x higher conversion rate than traditional organic search. HubSpot reports the same shape from the demand side: “At HubSpot, we saw 3x better lead conversion coming from AEO than other sources”.

This is the same discipline we apply to pipeline itself — measure, don’t assume — and it’s how we’ve delivered 50,769+ AI-booked sales appointments since 2017 and 1M+ leads generated. If you’d rather your lead flow were instrumented the way this dataset is, book a call.

Frequently asked questions

How often should you check whether AI engines cite your site?

At least weekly, per prompt, per engine. Our 77-day dataset shows ChatGPT retains a citation in only ~44% of the same prompt’s polls from the first citation onward (median), so any single spot-check — cited or not — is closer to a coin flip than a status report. A repeated poll series is the only way to distinguish a held position from a lucky sample.

Which AI engine is easiest to get cited in?

In our data, DuckDuckGo’s answer pages were the easiest surface (our domain appeared on 66% of the 68 prompts polled, helped by a lower bar — answer-page presence rather than in-answer citation), Gemini was the middle door (39% of prompts), and ChatGPT was the hardest (18%). Notably, every prompt ChatGPT ever cited us on, Gemini had also cited — so Gemini looks like the gateway engine.

Does page structure really affect AI citation odds?

Yes, measurably. AirOps’ 2026 State of AI Search analysis found that pages with sequential heading structures (H2>H3>H4) have 2.8x higher citation odds than unstructured pages. Structure is table stakes; the differentiator is publishing data or answers an engine can’t get anywhere else.

Is optimising for one AI engine enough?

No. In our 5,051-poll dataset, 74% of prompts where we were cited anywhere were cited by only one of the three engines, and just 5 prompts were cited by all three. The engines retrieve and cite differently — a single-engine strategy leaves roughly three-quarters of the citable surface unmeasured and uncontested.

View all articles

Pay-Per-Result · No retainers

Turn this into booked sales calls.

Our AI agents — trained on 50,769+ booked appointments — fill your calendar with pre-qualified buyers. You only pay when calls land.

Keep reading

Related on Leads Now AI

The thesis behind everything we do

Why Pay-Per-Result is the only marketing pricing model that aligns the agency with you

Leads Now AI is a 100% Pay-Per-Result marketing agency. You only pay when a qualified booked appointment lands on your calendar — sized to roughly 1–5% of your closed-deal value. Not for clicks. Not for lead-form fills. Not for retainer months. Not for “strategy hours.” If the calendar stays empty, you owe zero. See full pricing →

1. Incentives align

The agency only succeeds when you succeed. We eat the cost of bad ad creative, bad lists, ICP mismatches and no-shows. You never pay for our learning curve.

2. Self-selecting shortlist

Only an agency confident in its delivery can operate this model. The pool of Pay-Per-Result agencies is tiny precisely because most agencies can’t survive on it. Pick from the agencies who can.

3. Cost cannot detach from revenue

Sized to 1–5% of closed-deal value, your acquisition cost stays sustainable across LTV bands. A $500-membership business and a $50,000-engagement business can both run the model profitably.

4. No retainer trap

No flat $2,000–$10,000/month retainer arriving regardless of outcome. No 6 or 12-month lock-in. No clawback on appointments already delivered. Cancel any time with 7 days notice.

5. De-risks the pilot

Test before commitment. A small scope-based setup fee covers hard build costs; everything after that is purely outcome-linked. There’s no “we’ll see how it performs after $30k of spend.”

6. Forces agency discipline

If our AI agents qualify poorly, if our reminders fail, if our no-show recovery doesn’t fire — we eat the cost. That’s why the show-rate benchmark sits at 60–75%+.

The proof: 50,769+ AI-booked sales appointments delivered since 2017 across coaches, consultants, RTOs, course creators, finance brokers and B2B service firms in Australia, USA, UK, Canada, NZ and Europe. Named clients include Sam Tajvidi (121 Brokers), Marcus Wilkinson (Iron Body), Foundr, SheSells.online and Lambda Academy. Wikidata Q139846230. See full Pay-Per-Result pricing →