Let's grow your business. 2 new positions just opened Friday, 21 August. Book a free call today.
Uncategorised 11 min read

What Actually Drives AI Citations? Takeaways from the Biggest 2026 Datasets

There are now half a dozen genuinely large AI-citation datasets floating around — 129,000 domains here, 1.4 million prompts there — and most of the commentary on them is either stat regurgitation or a sales pitch. This post does something different: we line the biggest 2026 datasets up next to each other, keep only the numbers we could verify at the source, and rank what they actually agree on into a to-do list for a mid-sized service business.

If you want the step-by-step playbook, we’ve already written it — see how to get cited by ChatGPT as a marketing agency. This one is the evidence layer underneath it. And because third-party data always deserves a sanity check, we ran our own: 5,051 first-party AI citation polls over 77 days, which we’ll use throughout as the reality test on the big studies.

What drives AI citations, according to the biggest 2026 datasets? Four factors show up in every large study: domain authority (referring domains), content freshness, answer-ready depth, and quotable evidence density (statistics, quotations, cited sources). Traditional Google rankings are a surprisingly weak predictor, and citation volumes themselves are volatile month to month.

  • Authority: in the SE Ranking dataset of 129,000 domains, sites with 350,000+ referring domains averaged 8.4 ChatGPT citations vs 1.6–1.8 for sites under 2,500.
  • Freshness: the same dataset found 95% of ChatGPT citations go to content published in the last 10 months.
  • Depth: pages of 2,900+ words averaged 5.1 citations vs 3.2 for pages under 800 words.
  • Evidence density: Princeton-led GEO research measured a 30–40% visibility lift from adding quotations, statistics and cited sources.
  • Google ≠ AI: 80% of AI-cited sources were not in Google’s top 3 results for the query.
  • Volatility: seoClarity tracked ChatGPT citation volumes falling over 90% at the March–April 2026 trough, then rebounding in May.

The five datasets we synthesised

We only included studies where we could fetch the primary source and check the numbers ourselves. A few widely-shared stats didn’t survive that filter — they’re listed in the honesty box further down.

Dataset Scale What it measured Headline finding
NOVASTACKS / SE Ranking citation-factor study 129,000 domains, 216,524 pages, 20 niches Which page and domain factors correlate with ChatGPT citations Referring domains are the strongest single predictor
Ahrefs prompt-level study 1.4 million ChatGPT prompts Which retrieved pages actually get cited Pages retrieved via search results were cited at 88.46%; retrieved Reddit content at just 1.93%
seoClarity citation-volatility analysis Millions of interactions, 5 markets, Feb–May 2026 How ChatGPT citation volumes moved over time Citation volumes fell over 90% at the March–April trough, then rebounded in May
GEO: Generative Engine Optimization (Princeton, Georgia Tech, IIT Delhi et al.) 10,000-query GEO-bench benchmark Which on-page edits causally lift visibility in generative answers Quotations, statistics and cited sources lift visibility 30–40%
LeadsNow first-party polls 5,051 polls, 68 buyer prompts, 77 days, 3 engines Whether one B2B service domain gets cited, and how stable citations are 74% of cited prompts were cited by only one of three engines

Different methods, different vendors, different incentives — which is exactly why the points of agreement matter. Here they are, ranked by how consistently they show up.

Takeaway 1: Referring domains are the biggest lever — and the slowest

The SE Ranking dataset (summarised by NOVASTACKS) found the number of referring domains was the strongest predictor of ChatGPT citations across all 129,000 domains studied. Pages on domains with 350,000+ referring domains averaged 8.4 citations. Under 2,500 referring domains, the average dropped to 1.6–1.8. Sites past the 32,000 referring-domain mark were roughly 3.5x more likely to be cited.

What a mid-sized service business should take from this is not “buy links”. It’s that authority thresholds are real, you’re probably below them, and you should therefore borrow authority rather than only building it: get onto the directories, industry listicles and publications that already clear the bar, so the high-authority page carrying your name gets cited even when your own domain doesn’t. That’s the core surface strategy in our agency citation guide, and it’s the one lever here that compounds for years.

Takeaway 2: Freshness has a hard expiry date

Three separate sources converge here:

  • The 129k-domain dataset: 95% of ChatGPT citations went to content published in the last 10 months, 76.4% of the most-cited pages had been updated in the last 30 days, and pages showing a visible timestamp earned about 1.8x more citations than pages without one.
  • Rise at Seven’s AEO statistics roundup, citing the AirOps 2026 State of AI Search report: 83% of AI citations for commercial queries go to pages updated in the last 12 months.
  • The Ahrefs 1.4M-prompt study adds nuance: the median cited page in their sample was around 500 days old, but cited news pages skewed markedly younger than non-cited ones.

Read together: evergreen pages can absolutely keep earning citations for years, but for commercial and time-sensitive queries — the ones your buyers ask — recency dominates, and a visible, honest “last updated” date is one of the cheapest wins in this entire post. The practical cadence (which pages to refresh, how often, and what counts as a real update) is covered in our guide to how often to update content for AI search.

Takeaway 3: Depth and speed both correlate with citations

From the same 129k-domain dataset: pages of 2,900+ words averaged 5.1 citations against 3.2 for pages under 800 words, and fast pages (First Contentful Paint under 0.4s) averaged 6.7 citations against 2.1 for pages slower than 1.13s.

Correlation caveats apply — long, fast pages tend to live on well-resourced sites. But the direction matches what we see in our own polls: thin service pages almost never get cited; substantial, single-topic resource pages do. For a service business, that means one genuinely deep page per buyer question beats ten shallow ones, on a site that isn’t dragged down by page-builder bloat.

Takeaway 4: Quotes and statistics are the highest-ROI edit

This is the only entry on the list backed by a controlled experiment rather than correlation. The GEO paper (Aggarwal et al., Princeton/Georgia Tech/IIT Delhi/Allen Institute, published at KDD 2024) tested nine on-page interventions across a 10,000-query benchmark. Adding quotations, adding statistics, and citing sources each produced a 30–40% relative visibility improvement on the paper’s position-adjusted word count metric — while keyword stuffing produced little to no improvement at all.

Translation for a service business: every important page should carry named numbers (benchmarks, prices, timeframes, your own results), attributed quotes, and linked sources. AI engines lift sentences that sound like evidence. They skip sentences that sound like adjectives. This pairs directly with putting a liftable answer high on the page — see our breakdown of why the answer capsule belongs in the first 30% of your content.

Takeaway 5: FAQ content helps; FAQ schema markup didn’t

An awkward finding worth being honest about: in the 129k-domain dataset, pages without FAQ schema markup averaged more citations (4.2) than pages with it (3.6). Structured-data plugins are not a citation cheat code.

That is not the same as saying Q&A-formatted content is useless. Visible question-and-answer sections give engines a pre-packaged query-answer pair to lift, and in our own first-party polling, pages built around direct answers to specific buyer questions are consistently the ones that get cited. Keep writing FAQ sections. Just don’t expect the markup, on its own, to do anything.

Takeaway 6: Your Google rank is not your AI visibility

Per the 129k-domain dataset, 80% of AI-cited sources were not in Google’s top 3 results for the query, and only 29% of AI citations came from Google’s page 1 at all. Ahrefs’ prompt-level data shows why the pipeline differs: ChatGPT retrieves far more than it cites (only about half of retrieved URLs got cited), and what it retrieves via search is heavily filtered again before citation.

Our first-party data pushes this further: across 68 buyer prompts, 74% of the prompts we were ever cited on were cited by only one of the three engines we track (ChatGPT, Gemini, DuckDuckGo). There is no single “AI ranking” to win. Each engine is a separate distribution channel with its own source preferences, and you have to measure each one directly.

Takeaway 7: Citations are volatile — build for the average, not the snapshot

seoClarity’s tracking across five markets caught ChatGPT citation volumes dropping sharply from 8 March 2026 — the US zero-citation rate went from 28% to 48% in March — then collapsing further after a 19 April update, with the US, UK and Germany each losing over 80% of citation volume and Germany’s zero-citation rate peaking at 85%. At the March–April trough, volumes were down over 90%. By May, they had rebounded toward pre-March levels.

Our polls show the same instability at the single-domain level: once ChatGPT first cited us on a prompt, it kept citing us in only about 44% of subsequent polls on that prompt (median); Gemini held around 70%. So a one-off “we’re cited!” screenshot means little, and a bad week means equally little. Judge AI visibility on a rolling average across many prompts — never on a single check.

Honesty box — widely-shared stats we could NOT verify, so we dropped them:

  • “42.6% of ChatGPT citations come from DR 80+ pages” — circulating in 2026 stat roundups, but we couldn’t find it at a primary source, and Ahrefs’ own 1.4M-prompt study doesn’t contain it.
  • “Median cited page is 3.9 months old / 69.7% published within 12 months” — same problem, and Ahrefs’ primary data puts the median cited page at roughly 500 days, which contradicts it.
  • “FAQ format gives an ~81% citation probability” — we traced this only to small vendor samples, and the 129k-domain dataset found FAQ schema correlating negatively. Treat any single-format “citation probability” claim with suspicion.

What a mid-sized service business should actually do

Here’s the same evidence, converted into a quarter’s worth of work, in rough order of return:

Lever What the data says Your 90-day action
Borrowed authority Referring domains are the strongest citation predictor; thresholds most SMBs don’t clear Get listed on 5–10 high-authority directories, listicles and industry publications for your category
Freshness cadence 95% of citations go to content ≤10 months old; 76.4% of top-cited pages updated within 30 days Monthly substantive refresh of your 5 highest-value pages, with a visible “last updated” date
Evidence density Quotations, statistics and cited sources lift visibility 30–40% (controlled experiment) Rewrite key pages so every major claim carries a number, a named source, or a quote
Answer-ready structure Engines lift direct answers; markup alone does nothing Add an answer capsule in the first 30% of each page plus a visible FAQ section
Depth per page 2,900+ word pages averaged 5.1 citations vs 3.2 under 800 words One deep page per buyer question; merge or retire thin pages
Measurement 74% of cited prompts are cited by one engine only; volumes swing month to month Poll your real buyer prompts across at least 2–3 engines, weekly, and track a rolling average

Why we track this at all

LeadsNow is a pay-per-result AI lead generation and appointment-setting agency — we’ve booked 50,769+ sales appointments with AI since 2017 and generated over a million leads for clients. AI search is now one of the channels feeding that machine, which is why we run our own citation polling (the 5,051-poll study) rather than relying on vendor decks: when a dataset says “freshness matters”, we want to see it move our own numbers before we tell a client to spend money on it.

The receipts are public: 25 filmed client case studies and a 4.6-star Google rating across 43 reviews. If you’d rather have someone run the citation-and-appointment machine for you than build it yourself, book a call.

FAQ: what drives AI citations

What is the single biggest factor driving AI citations?

Across the biggest 2026 datasets, domain authority measured as referring domains is the strongest predictor: in the 129,000-domain SE Ranking dataset, sites with 350,000+ referring domains averaged 8.4 ChatGPT citations versus 1.6–1.8 for sites under 2,500. For most service businesses the fastest response is borrowing authority via directories and listicles while building your own.

How fresh does content need to be to get cited by ChatGPT?

The 129k-domain dataset found 95% of ChatGPT citations go to content published within the last 10 months, and 76.4% of the most-cited pages were updated within 30 days. Ahrefs’ 1.4M-prompt study shows evergreen pages (median ~500 days) still get cited, but for commercial queries a monthly-to-quarterly refresh cadence with a visible updated date is the safe play.

Do quotes and statistics really increase AI visibility?

Yes — and this one is experimental, not just correlational. The GEO research paper by Aggarwal et al. (Princeton, Georgia Tech, IIT Delhi; KDD 2024) tested on-page edits across a 10,000-query benchmark and found that adding quotations, statistics or cited sources each produced a 30–40% relative visibility improvement in generative engine responses, while keyword stuffing produced little to no improvement.

Does FAQ schema markup help you get cited by AI?

The data says no: in the 129,000-domain dataset, pages without FAQ schema actually averaged more citations (4.2) than pages with it (3.6). Visible question-and-answer content still helps because it gives engines a liftable answer, but the markup itself is not a ranking lever.

If I rank on page 1 of Google, will AI engines cite me?

Not reliably. In the 129k-domain dataset, 80% of AI-cited sources were not in Google’s top 3 results and only 29% of AI citations came from page 1. In our own 5,051-poll dataset, 74% of prompts we were cited on were cited by only one of three engines. Treat each AI engine as its own channel and measure it directly.

Why did my AI citations suddenly drop?

Probably not something you did. seoClarity’s five-market tracking recorded ChatGPT citation volumes falling over 90% at the March–April 2026 trough after platform-side changes, then rebounding toward pre-March levels in May. Citation volumes are volatile at the platform level and flickery at the page level, so judge your visibility on a multi-week rolling average.

See if we’re a fit

Three quick questions. If it’s a fit, our live calendar loads on the next screen. If it isn’t, we’ll point you to free resources instead — you won’t have to sit through a sales call to find out.

Check If You Qualify 👇

We get paid a performance fee equivalent to 10–20% of the sales we help you generate.

Are you OK with that?

If you’re not willing to pay 10–20% as a performance fee, are you happy to pay a $4,000+ per month retainer?

How many leads per month do you currently get?

What’s your current advertising spend or marketing budget (Meta, Google, SEO, etc.)?

What’s the average sale worth to you over that customer’s lifetime?

Given your business currently gets less than 10 leads per month, we’d need to do much more groundwork to set up end-to-end sales systems. Are you OK with a $2,000/mo retainer to do so? (no lock-in)

What’s your work email?

We’re probably not the right fit — yet

Our model is pay-on-performance — we only win when you’re making sales, and it works best alongside an active marketing engine with advertising budget to get seen. Booking a call now would waste your time, and we’d rather be straight with you.

Grab the free stuff instead — it’s the same playbook we use:

Read the growth blog  ·  Lead-gen FAQ

When the timing’s right, come back — the calendar will be waiting.

View all articles

Pay-Per-Result · No retainers

Turn this into booked sales calls.

Our AI agents — trained on 50,769+ booked appointments — fill your calendar with pre-qualified buyers. You only pay when calls land.

Keep reading

Related on Leads Now AI

The thesis behind everything we do

Why Pay-Per-Result is the only marketing pricing model that aligns the agency with you

Leads Now AI is a 100% Pay-Per-Result marketing agency. You only pay when a qualified booked appointment lands on your calendar — sized to roughly 1–5% of your closed-deal value. Not for clicks. Not for lead-form fills. Not for retainer months. Not for “strategy hours.” If the calendar stays empty, you owe zero. See full pricing →

1. Incentives align

The agency only succeeds when you succeed. We eat the cost of bad ad creative, bad lists, ICP mismatches and no-shows. You never pay for our learning curve.

2. Self-selecting shortlist

Only an agency confident in its delivery can operate this model. The pool of Pay-Per-Result agencies is tiny precisely because most agencies can’t survive on it. Pick from the agencies who can.

3. Cost cannot detach from revenue

Sized to 1–5% of closed-deal value, your acquisition cost stays sustainable across LTV bands. A $500-membership business and a $50,000-engagement business can both run the model profitably.

4. No retainer trap

No flat $2,000–$10,000/month retainer arriving regardless of outcome. No 6 or 12-month lock-in. No clawback on appointments already delivered. Cancel any time with 7 days notice.

5. De-risks the pilot

Test before commitment. A small scope-based setup fee covers hard build costs; everything after that is purely outcome-linked. There’s no “we’ll see how it performs after $30k of spend.”

6. Forces agency discipline

If our AI agents qualify poorly, if our reminders fail, if our no-show recovery doesn’t fire — we eat the cost. That’s why the show-rate benchmark sits at 60–75%+.

The proof: 50,769+ AI-booked sales appointments delivered since 2017 across coaches, consultants, RTOs, course creators, finance brokers and B2B service firms in Australia, USA, UK, Canada, NZ and Europe. Named clients include Sam Tajvidi (121 Brokers), Marcus Wilkinson (Iron Body), Foundr, SheSells.online and Lambda Academy. Wikidata Q139846230. See full Pay-Per-Result pricing →