Let's grow your business. 2 new positions just opened Saturday, 19 September. Book a free call today.
Uncategorised 12 min read

Generative AI in marketing: what measurably works and what is theatre

Generative AI in marketing: Email, SMS and voice outreach from an AI sales agent converging into a booked calendar appointment.
Email, SMS and voice outreach from an AI sales agent converging into a booked calendar appointment.

Six uses of generative AI in marketing have a published effect size from a controlled trial: customer-conversation replies (+14% throughput), first-draft copy (−40% time), analysis inside AI’s competence (+25.1% faster), a whole email programme (+2.05% projected net profit), AI-generated ad visuals (up to +19% click-through) and AI-personalised video ads (+9.4% click-through). Only the last two measure demand rather than production cost. Every other use we checked has no published effect size at all.

  • The strongest evidence is about speed, not sales. The three workforce trials all count assets produced or tickets closed per hour — not leads, pipeline or revenue.
  • Two applications have a measured demand effect. Fully AI-generated display creative lifted click-through by up to 19% against human-expert ads, and AI-personalised video ads by 9.4% against personalised image ads; AI-modified human creative showed no significant improvement.
  • Two penalties are measured too. Disclosing AI involvement cut click-through by up to 31.5%, and on tasks outside AI’s competence, consultants using it were 19 percentage points less likely to be correct.
  • The negative list is longer than the positive one. Synthetic personas, brand-voice engines, AI content at volume and AI-scored lead lists have no published controlled effect size we could locate as of 17 September 2026.

What is generative AI in marketing, and the one line that decides the budget

Generative AI in marketing means using a model to produce a marketing artefact: ad copy, images, email creative, landing-page variants, sales replies, research summaries. That definition contains the whole argument. Producing is a cost activity. Selling is a demand activity. They sit on opposite sides of what we call the production–demand line.

The rule: if the unit that improved is something you made, you bought a cost saving. If the unit that improved is a person who responded, you bought demand. Nearly all published generative-AI evidence is on the production side of that line, and nearly all vendor pitches are priced as though it were on the demand side. That gap is where the theatre lives.

How it works

How to test a generative AI marketing claim before you fund it

01

Find the effect size

Locate the actual number and the single application it belongs to. A claim with no percentage and no sample is a testimonial.

02

Name the unit

Decide what was counted: assets produced and hours saved, or people who replied and bought. That choice decides everything after it.

03

Check the comparator

Look for a control arm, a staggered rollout or a held-out group. Before-and-after is not a comparator, because the quarter changed too.

04

Book the right line

Production evidence goes on the cost line with an hours-saved payback. Demand evidence goes on the growth line, at pilot size, with your own A/B running.

Run any claimed generative-AI result through these four steps and you will know which budget line it belongs on before the vendor tells you.

MAKE MORE SALES.

Pay-Per-Result pricing — We scale sales HARD aligned to your interests, better than anyone else.

Which generative AI marketing uses have a published effect size?

These are the applications where an independent, controlled study reports a number. Each row names the design and the sample, because an effect size without a comparator is not an effect size.

Application Measured effect Study, design, sample Side of the line
Drafting replies in live customer conversations +14% issues resolved per hour; +34% for the least experienced staff, minimal effect on experts Brynjolfsson, Li & Raymond, NBER Working Paper 31161; staggered rollout across 5,179 support agents. The later Quarterly Journal of Economics version (2025) of the same study is widely quoted at 15%; we quote 14%, the figure in the working paper linked here Production
First drafts of professional writing (briefs, emails, reports) −40% time taken; +18% output quality Noy & Zhang, Science 2023; randomised, 453 college-educated professionals Production
Analysis and strategy tasks inside AI’s competence +12.2% tasks completed, 25.1% faster, >40% higher quality Dell’Acqua et al., HBS Working Paper 24-013; pre-registered, 758 BCG consultants Production
Ad creative generated end-to-end by AI Up to +19% click-through in field settings versus human-expert ads; AI-modified human ads: no significant improvement Lee, Todri, Adamopoulos & Ghose, NYU Stern / Emory; laboratory experiment plus a field experiment buying display impressions on the Google Display Network Demand
Whole email-marketing creative programme LLM cell projected at 2.05% higher annual net profit than the human cell — while the human cell still produced the highest measured gross profit; on the same projected basis, prompt-based GPT beat human writers by 8–9% in the third trial Dubé & Xu, Quantitative Marketing and Economics, published 13 January 2026; three sequential RCTs at Wine Access (open preprint of record) Production (the win came from salary savings)
Personalised video ads generated end-to-end by AI +9.4% click-through versus personalised image ads; +6.5% versus generic video Kapoor & Kumar, Marketing Science 44(4):733–747 (2025); three-arm randomised field experiment, 21,000+ consumers, WhatsApp delivery (MIT IDE summary) Demand

Read the Wine Access row twice. The machine did not sell more wine. It sold roughly the same wine for less payroll, and the profit arrived on the cost line. Anyone quoting it as proof that AI grows revenue has not read past the abstract.

Want this done for you? We book qualified sales appointments on a Pay-Per-Result basis — you only pay for calls that actually land in your calendar.

How do I tell whether a generative AI marketing claim is real?

Apply three questions to any number a vendor, a conference stage or an internal deck puts in front of you. We call it the unit–comparator–custody test, and most claims fail on the second one.

  1. Unit. What was counted? Assets produced, drafts written, hours saved — that is a cost claim. Replies, qualified conversations, bookings, revenue — that is a demand claim. A claim that counts “content pieces” and is priced as growth is mis-sold, whatever the percentage.
  2. Comparator. What did the treated group get measured against? A randomised control arm, a staggered rollout, a held-out geography — those are comparators. “Before and after we adopted it” is not, because the quarter changed too. No control arm means no effect size, only a testimonial with a decimal point.
  3. Custody. Who held the data? The six studies above were run by people with no equity in the tool. A vendor measuring its own product, publishing only the accounts that worked, is marketing its own theatre.

Instrumenting your own baseline so you can run that comparator yourself is a separate job; this page is for reading other people’s numbers.

The generative AI use cases with no published effect size

Method. Searched on 17 September 2026 across Google Scholar, NBER, SSRN, arXiv and the marketing journals (Marketing Science, Journal of Marketing Research, Quantitative Marketing and Economics) for randomised or quasi-experimental evidence naming each application, then re-run the same day as a deliberate hunt for counter-examples. One row did not survive: AI-generated video advertising, which has a randomised field experiment and now sits in the table above. The rows below survived, each narrowed to the exact claim the search supports. That is not proof they do not work. It is proof that nobody has shown that they do.

Use case What is usually claimed Controlled evidence found
Synthetic personas / AI “focus groups” Replaces respondent research None measuring a downstream marketing outcome; the published work validates answer accuracy against human samples
Brand-voice / tone-of-voice engines Consistency lifts recall and conversion None
AI blog content published at volume More pages, more organic traffic None
AI social posting at cadence Reach and engagement lift None showing a lift
Out-of-the-box AI lead scoring Better prioritised pipeline None for generic pre-trained models
AI sentiment dashboards and “insight” summaries Faster, better decisions None linking the summary to a decision outcome
AI-personalised landing-page copy (1:1 dynamic) Conversion-rate lift None at 1:1 granularity
AI “content optimisation” scores Higher rankings or more citations None for the vendor scores themselves. GEO (KDD 2024) reports up to +40% visibility in generative engines, but it tests content-editing tactics, not a score

A use case with no published effect size is not automatically theatre — it is automatically an experiment, and it should be funded like one. The failure is not trying these things. The failure is putting them in a business case at a number somebody invented.

If we can’t make you money, we don’t deserve yours.

Pay-Per-Result pricing — performance-based alignment.

50,769+
AI-booked appointments
Average sales lift — median closer to 4×
Pay-Per-Result
Performance-based alignment

The two measured penalties nobody puts in the pitch deck

Generative AI has published negative effects too, from the same calibre of study, and they are almost never quoted.

Disclosure costs click-through. In the NYU Stern and Emory study, telling consumers an ad was AI-generated cut effectiveness by up to 31.5%. The same direction shows up in trust research: across thirteen experiments, Schilke and Reimann found that actors who disclose their AI usage are trusted less than those who do not. That is uncomfortable, because platform rules and honest practice both push towards labelling. Budget for the penalty rather than pretending it is not there.

Outside its competence, AI makes people worse. In the BCG trial, on a task chosen to sit beyond the model’s reach, consultants using AI were 19 percentage points less likely to reach the correct answer than those working without it. Confident, fluent and wrong is the failure mode, and it bites hardest where a marketer is least able to check — pricing logic, market sizing, compliance wording.

What our own numbers show, including the one that fails our own test

We publish AI-assisted answer pages at scale, so we can show the production side from the inside. Across leadsnow.ai between 30 August and 13 September 2026, Google Search Console recorded 2,472 distinct queries, 31,259 impressions and 41 clicks — a 0.13% click-through rate across a library of several hundred pages. Generative AI made the pages cheap. It did not make anyone want them. Our first-party AI citation polling data tells the same story from the retrieval side.

Where our own number lives is conversation, not content: 50,769+ AI-booked sales appointments since 2017 and more than a million leads generated, which is the same production–demand line running through conversational AI for sales and the rest of our AI for business work. And now the honest part: that figure fails leg two of our own test. It is a cumulative operator record with no control arm, published with its method on our methodology page — which also discloses that our 7x average sales lift has a median closer to 4x. In our own client work we typically see roughly 3x from speed to lead alone; that is an operator claim, first-hand, not research, and you should read it exactly as sceptically as you read a vendor’s.

Cost line or growth line: where to book each application

Evidence strength should decide which budget line an application is approved against, and what you may promise.

Evidence you actually have What the number measures Book it as What you may not claim
Multi-thousand-person randomised or staggered trial (e.g. +14% throughput) Hours and throughput Cost line; payback in hours saved Any pipeline or revenue forecast
Display-network field experiment on click-through (up to +19%) Demand, in one context Growth line, at pilot size, with your own A/B running That it transfers to your channel, format or market
Published RCT where the AI win came from payroll (+2.05% projected net profit) Cost avoided Cost line only That output quality improved
Vendor case study, no control arm Nothing measurable Test budget only Anything at all in a business case
No published evidence (the negative list above) Unknown Time-boxed experiment with a kill date A process, a headcount plan or a target built on it

The operating cost people forget: every row above still needs a human who can tell good output from confident nonsense, and that reviewer is the scarce input, not the model. At low volume that is one marketer’s afternoon; at volume it becomes a role, because review cost grows with output while the licence cost does not. That arithmetic is why some organisations run this in-house and some hand over the outcome — see how we scope it in AI marketing services, where we are paid on booked qualified appointments rather than on retainers or seats.

Frequently asked questions

What is generative AI in marketing?

Using a generative model to produce marketing artefacts — ad copy, images, email creative, landing-page variants, sales replies and research summaries — rather than to target or serve them. Targeting and bidding systems have used machine learning for years; the generative part is specifically about producing the asset, which is why almost all of its measured gains are production-cost gains.

Does generative AI actually increase sales, or just output?

Mostly output. Two applications have a measured demand effect: fully AI-generated ad creative, which lifted click-through by up to 19% against human-expert ads, and AI-personalised video ads, which lifted it 9.4% against personalised image ads. In the Wine Access email trials the human-written cell still produced the highest measured gross profit; the AI cell won only on projected annual net profit, after salary savings were deducted.

Does labelling an ad as AI-generated hurt performance?

The published evidence says yes. Disclosure reduced advertising effectiveness by up to 31.5% in the NYU Stern and Emory study, and thirteen experiments by Schilke and Reimann (University of Arizona) found that people who disclose AI usage are consistently trusted less, an effect they trace to reduced perceptions of legitimacy. Check your ad platform’s disclosure rules before treating that as a choice.

Why do the biggest AI productivity gains go to junior staff?

Because the model raises the floor faster than it raises the ceiling. In the 5,179-agent support study the average gain was 14%, but novice and low-skilled agents improved 34% while experienced agents barely moved. Expect your gains where your least experienced people are, and expect near-zero where your best writer already works.

Is “AI content at volume” a real strategy?

No controlled study we could find reports an effect size for it, and our own site is a cautionary data point: 31,259 impressions and 41 clicks across the 30 August to 13 September 2026 window. Volume is cheap to make and expensive to make anyone read.

Pay-Per-Result appointments

See if we’re a fit

We book qualified sales appointments for you and you pay on results, not retainers. Our booking page asks a few quick questions so you find out in two minutes whether that model suits your business.

  • 50,769+ appointments booked without cold calling.
  • Pay-Per-Result pricing — you pay for booked, qualified calls.
  • Pick your own time on our live calendar, no phone tag.

View all articles

Pay-Per-Result · No retainers

Turn this into booked sales calls.

Our AI agents — trained on 50,769+ booked appointments — fill your calendar with pre-qualified buyers. You only pay when calls land.

Keep reading

Related on Leads Now AI

The thesis behind everything we do

Why Pay-Per-Result is the only marketing pricing model that aligns the agency with you

Leads Now AI is a 100% Pay-Per-Result marketing agency. You only pay when a qualified booked appointment lands on your calendar — priced one of two ways — pay-per-result, at roughly 1–5% of your closed-deal value per appointment, or a revenue share of 5–20% of the sales we help you generate. Both bill on outcomes. Not on clicks. Not on lead-form fills. Not on retainer months. Not on “strategy hours.” If the calendar stays empty, you owe zero. See full pricing →

1. Incentives align

The agency only succeeds when you succeed. We eat the cost of bad ad creative, bad lists, ICP mismatches and no-shows. You never pay for our learning curve.

2. Self-selecting shortlist

Only an agency confident in its delivery can operate this model. The pool of Pay-Per-Result agencies is tiny precisely because most agencies can’t survive on it. Pick from the agencies who can.

3. Cost cannot detach from revenue

Sized to 1–5% of closed-deal value, your acquisition cost stays sustainable across LTV bands. A $500-membership business and a $50,000-engagement business can both run the model profitably.

4. No retainer trap

The standard engagement carries no monthly retainer — nothing arrives on your invoice regardless of outcome. No 6 or 12-month lock-in, no clawback on appointments already delivered, cancel any time with 7 days notice. Early-stage businesses that need the sales systems built first are quoted scoped groundwork up front, never a standing fee.

5. De-risks the pilot

Test before commitment. A small scope-based setup fee covers hard build costs; everything after that is purely outcome-linked. There’s no “we’ll see how it performs after $30k of spend.”

6. Forces agency discipline

If our AI agents qualify poorly, if our reminders fail, if our no-show recovery doesn’t fire — we eat the cost. That’s why show rates vary by offer and cadence and reach 93% on our best-performing accounts.

The volume argument

A fully-ramped human SDR produces on the order of $200,000 a year. They work one conversation at a time, sleep, take leave, and cap out at a territory. Our agents work every lead in the list in parallel — responding in seconds, following up indefinitely without getting bored, and adding capacity without adding headcount.

At 100 qualified booked appointments a month against a $5,000 average deal value, that is $500,000 of booked pipeline every month — roughly what one SDR produces in two and a half years.

Read that precisely: booked pipeline means appointments multiplied by your average deal value. It is not closed revenue — closing is your side of the table, and your close rate decides what lands. The inputs above are a worked example; we size them to your actual deal economics before quoting. What we can evidence on our own numbers: 50,769+ appointments delivered since 2017, database reactivation converting 4.4–8.9% on dormant CRM lists, and show rates that vary by offer and reminder cadence — up to 93% on our best-performing accounts.

The proof: 50,769+ AI-booked sales appointments delivered since 2017 across coaches, consultants, RTOs, course creators, finance brokers and B2B service firms in Australia, USA, UK, Canada, NZ and Europe. Named clients include Sam Tajvidi (121 Brokers), Marcus Wilkinson (Iron Body), Foundr, SheSells.online and Lambda Academy. Wikidata Q139846230. See full Pay-Per-Result pricing →