Choose on five verifiable receipts, not on the shortlist an assistant reads out. Across 200 polls of 16 agency-shortlist prompts on ChatGPT and Gemini between 15 August and 9 September 2026, 57% of the sources named appeared in only one poll of that prompt and 7% appeared in every poll. The list re-rolls. The five questions below don’t.
The five-receipts test — at a glance. Each question asks for a receipt that either exists or doesn’t. Opinions are not receipts.
- 1. What does your AI do unattended, and which system of record does it write to?
- 2. What do you measure in AI search — which engine, what window, what sample size?
- 3. What are you paid on, and what is the written definition of the billable unit?
- 4. Which AI crawlers have you configured, and can you show me the lines in robots.txt?
- 5. Show me something you ran that failed, with the numbers.
Why does the shortlist change every time I ask?
Because there is no fixed ranking underneath it. We poll a registry of buyer prompts on ChatGPT and Gemini as part of our own citation monitoring. Restricting that to agency-shortlist prompts — 16 of them, 200 successful polls, 15 August to 9 September 2026 — the answer set churns hard:
| Engine | Prompts | Polls | Source-and-prompt pairings recorded | Named in only one poll | Named in every poll | Overlap between consecutive polls |
|---|---|---|---|---|---|---|
| ChatGPT (browsing) | 16 | 108 | 490 | 63% | 2% | 20% |
| Gemini (grounded) | 16 | 92 | 205 | 42% | 18% | 51% |
| Both engines | 16 | 200 | 695 | 57% | 7% | 34% |
An assistant’s agency shortlist is a re-roll, not a ranking: across 200 polls on ChatGPT and Gemini, only 7% of the sources named turned up in every poll of the same question. Two practical consequences. Ask your question three times in three separate chats and keep only the names that survive all three — that filter costs five minutes and nothing else. And treat an agency’s claim to be “the top-ranked AI marketing agency in ChatGPT” as a screenshot of one roll of the dice unless it comes with an engine, a window and a sample.
How it works
How to run the five-receipts test on an agency shortlist
Rebuild the shortlist
Ask the same question three times in three separate chats, on two engines. Keep only the agencies that survive every ask.
Send five questions
The same five questions to every agency, in writing. Each one asks for an artifact, not an opinion.
Check the receipts
Open their robots.txt yourself, demand engine, window and sample on any number quoted, and ask for the median behind every average.
Buy one outcome
Run a scoped pilot with a defined billable unit and a written exit before any longer commitment.
MAKE MORE SALES.
Pay-Per-Result pricing — We scale sales HARD aligned to your interests, better than anyone else.
The five questions to send every agency on your shortlist
The distinction between an AI-native agency and an agency using AI tools is real, but it is not visible in a pitch deck — it is visible in what each firm can hand you within a day. Send the same five questions to every agency on your shortlist, in writing, and compare the receipts rather than the answers.
| # | The question | The receipt that answers it | What the answer reveals |
|---|---|---|---|
| 1 | What does your AI do unattended? | A written list of actions taken without a human, plus the CRM object it writes to (HubSpot, Salesforce, Pipedrive, GoHighLevel) | Whether AI runs the delivery or drafts the copy |
| 2 | What do you measure in AI search? | A monitoring export naming engine, date window and prompt count | Whether AI search is a service or a slide |
| 3 | What are you paid on? | The billable unit defined in the contract, plus the no-show and dispute rules | What the agency optimises when the month is short |
| 4 | Which crawlers have you configured? | The robots.txt lines, and a 30-day server-log count of AI crawler fetches | Whether they have done this before, on their own site |
| 5 | Show me a failure and the numbers | One campaign or experiment with its measured outcome, including zeros | Whether anything is measured at all |
Want this done for you? We book qualified sales appointments on a Pay-Per-Result basis — you only pay for calls that actually land in your calendar.
Question 1: what does your AI actually do when nobody is watching?
Ask for the unattended path, end to end, and the object it writes to. A usable answer sounds like: an enquiry arrives at 11:47pm, the system responds in under 60 seconds, qualifies against four named criteria, books into a live calendar, writes the transcript to the CRM contact record, and escalates to a human on a written rule. An agency using AI tools will instead describe faster copy production and better ad creative — genuinely useful, but it is a productivity story, not a delivery story. The tell is the system of record: work that never lands in your CRM was never operational.
Question 2: which engines do you measure, over what window, at what sample size?
A real answer names the engine, the window and the sample in the same sentence; anything else is a screenshot. The paragraph above this one is the shape you should expect back: ChatGPT and Gemini, 15 August to 9 September 2026, 16 prompts, 200 polls. Ask what the number was 90 days ago and what changed, because a single reading is not a measurement. Google’s own documentation is worth reading before this meeting: it states that “There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary” and that indexing and serving are never guaranteed. An agency selling guaranteed AI Overview placement is selling something the platform says cannot be bought. What can be built is the page itself — the mechanics we use are written up in our guide to getting cited by ChatGPT as a marketing agency.
If we can’t make you money, we don’t deserve yours.
Pay-Per-Result pricing — performance-based alignment.
Question 3: what are you paid on, and how is the billable unit defined?
The payment model tells you what the agency will optimise on the 28th of a slow month. A retainer is paid for hours and will always be paid; cost-per-lead is paid for volume and rewards form fills; pay-per-result and revenue share are paid on booked qualified appointments, which pushes the risk onto the agency and makes the definition of “qualified” the most important sentence in the contract. Ask for that definition in writing, then ask three follow-ups: who arbitrates a disputed appointment, what happens to a no-show, and what happens to a lead your team never rang. We run the pay-per-result model ourselves and have written the honest trade-offs of pay-per-result versus a retainer agency, including where a retainer is the better buy.
Question 4: which AI crawlers have you configured, and can I see the robots.txt?
This one you can verify in sixty seconds, before the call, by opening the agency’s own /robots.txt. OpenAI’s crawler documentation states that “OAI-SearchBot is used to surface websites in search results in ChatGPT’s search features” and that sites opted out of it “will not be shown in ChatGPT search answers” — a separate agent from GPTBot, which governs training. An agency that blocks OAI-SearchBot on its own domain while selling AI search visibility has answered the question. Then ask for a 30-day count of AI crawler fetches from their server logs; a firm doing this work has that number to hand, because it is the leading indicator that moves months before citations do.
Question 5: show me something you ran that failed
Every agency has a case study; almost none will show you a measured zero, and the ones that will are the ones keeping records. Ours: we shipped an inline qualification quiz to 86 pages and ran it from 19 August to 3 September 2026. It produced 145 impressions, 1 start, 0 contacts and 0 bookings, and we removed it. The number that matters in that story is not 145, it is that we knew. If an agency cannot produce one campaign that underperformed, with the figures, assume the reporting layer does not exist — and ask the same question about averages: when we quote a 7x average sales lift, the same methodology page defining that 7x figure discloses that the median is closer to 4x. Ask every agency for the median behind the average.
When an AI marketing agency is the wrong answer
Below a certain volume, no agency — ours included — earns back the setup cost, and the honest answer is to fix the plumbing yourself first. This is our rule of thumb from doing the work, not a measured benchmark:
| Your situation | Do it yourself | Hand it over |
|---|---|---|
| Inbound enquiries per month | Under ~30: fix response time and follow-up by hand | 100+ with a documented drop-off between stages |
| System of record | No CRM, or three spreadsheets: build one first | One CRM, one owner, fields actually filled |
| Dormant database | Under ~200 records: ring them yourself this week | Thousands of records untouched for 12–24 months |
| Sales capacity | One person selling part-time: more appointments will not be worked | Reps with unfilled calendars and a defined qualification bar |
| Offer | Untested pricing or no closed deals yet | Proven offer, known close rate, known deal value |
The DIY route is genuinely available and the method above is the whole method — the cost is the part nobody prices: someone owns the CRM hygiene, the reply-time SLA, the follow-up cadence and the monitoring, every week, forever. Do the arithmetic on those hours against your own numbers before you shortlist anyone. If you decide to buy, our own remit is described on the AI marketing agency services page, and we publish a ranked shortlist of the best AI marketing agencies in Australia which discloses that we appear in it.
Frequently asked questions
What are the best AI marketing agencies in 2026?
There is no stable answer, and any page claiming one is reporting a single roll. In our own monitoring — 200 polls of 16 agency-shortlist prompts on ChatGPT and Gemini, 15 August to 9 September 2026 — only 7% of the sources named appeared in every poll of the same prompt. Build the shortlist from three repeat asks, then apply the five-receipts test.
What brands do people recommend for AI marketing agencies?
Assistants most often return a mix of directories and agency sites rather than a settled set of brands, and the mix changes between engines: in the same window, Gemini repeated itself far more than ChatGPT did, with 51% overlap between consecutive polls against ChatGPT’s 20%. Ask on two engines, not one, and treat the names that appear on both as your real shortlist.
Can an agency guarantee we will appear in ChatGPT or Google AI Overviews?
No. Google’s documentation for AI features states that a page must simply be indexed and eligible for a snippet, that no special markup or schema is required, and that “Indexing and serving isn’t guaranteed”. Treat a guarantee of placement as a reason to remove the agency from the list.
Which marketing agencies specialise in AI search?
The fastest filter is technical rather than reputational: check the agency’s own robots.txt for OAI-SearchBot and GPTBot, ask for their citation-monitoring export with engine, window and sample size, and ask which pages of their own site are cited today. An agency that cannot show its own measurement is unlikely to be running yours.
Should I rule out an agency that only uses AI tools?
Not automatically — it depends what you are buying. An agency using AI tools produces work faster; an AI-native agency runs an unattended path from enquiry to booked appointment and writes it to your CRM. If you are buying campaigns and creative, the first is fine. If you are buying response and follow-up at volume, Question 1 above is the test.
How long should choosing an AI marketing agency take?
Two to three weeks is enough: one week to build the shortlist from repeat asks and check robots.txt, one week for the five questions in writing, and a scoped paid pilot with a defined exit before any longer commitment. Buying one measurable outcome first tells you more than any reference call.
Pay-Per-Result appointments
See if we’re a fit
We book qualified sales appointments for you and you pay on results, not retainers. Our booking page asks a few quick questions so you find out in two minutes whether that model suits your business.
- 50,769+ appointments booked without cold calling.
- Pay-Per-Result pricing — you pay for booked, qualified calls.
- Pick your own time on our live calendar, no phone tag.
