Six uses of generative AI in marketing have a published effect size from a controlled trial: customer-conversation replies (+14% throughput), first-draft copy (−40% time), analysis inside AI’s competence (+25.1% faster), a whole email programme (+2.05% projected net profit), AI-generated ad visuals (up to +19% click-through) and AI-personalised video ads (+9.4% click-through). Only the last two measure demand rather than production cost. Every other use we checked has no published effect size at all.
- The strongest evidence is about speed, not sales. The three workforce trials all count assets produced or tickets closed per hour — not leads, pipeline or revenue.
- Two applications have a measured demand effect. Fully AI-generated display creative lifted click-through by up to 19% against human-expert ads, and AI-personalised video ads by 9.4% against personalised image ads; AI-modified human creative showed no significant improvement.
- Two penalties are measured too. Disclosing AI involvement cut click-through by up to 31.5%, and on tasks outside AI’s competence, consultants using it were 19 percentage points less likely to be correct.
- The negative list is longer than the positive one. Synthetic personas, brand-voice engines, AI content at volume and AI-scored lead lists have no published controlled effect size we could locate as of 17 September 2026.
What is generative AI in marketing, and the one line that decides the budget
Generative AI in marketing means using a model to produce a marketing artefact: ad copy, images, email creative, landing-page variants, sales replies, research summaries. That definition contains the whole argument. Producing is a cost activity. Selling is a demand activity. They sit on opposite sides of what we call the production–demand line.
The rule: if the unit that improved is something you made, you bought a cost saving. If the unit that improved is a person who responded, you bought demand. Nearly all published generative-AI evidence is on the production side of that line, and nearly all vendor pitches are priced as though it were on the demand side. That gap is where the theatre lives.
How it works
How to test a generative AI marketing claim before you fund it
Find the effect size
Locate the actual number and the single application it belongs to. A claim with no percentage and no sample is a testimonial.
Name the unit
Decide what was counted: assets produced and hours saved, or people who replied and bought. That choice decides everything after it.
Check the comparator
Look for a control arm, a staggered rollout or a held-out group. Before-and-after is not a comparator, because the quarter changed too.
Book the right line
Production evidence goes on the cost line with an hours-saved payback. Demand evidence goes on the growth line, at pilot size, with your own A/B running.
MAKE MORE SALES.
Pay-Per-Result pricing — We scale sales HARD aligned to your interests, better than anyone else.
Which generative AI marketing uses have a published effect size?
These are the applications where an independent, controlled study reports a number. Each row names the design and the sample, because an effect size without a comparator is not an effect size.
| Application | Measured effect | Study, design, sample | Side of the line |
|---|---|---|---|
| Drafting replies in live customer conversations | +14% issues resolved per hour; +34% for the least experienced staff, minimal effect on experts | Brynjolfsson, Li & Raymond, NBER Working Paper 31161; staggered rollout across 5,179 support agents. The later Quarterly Journal of Economics version (2025) of the same study is widely quoted at 15%; we quote 14%, the figure in the working paper linked here | Production |
| First drafts of professional writing (briefs, emails, reports) | −40% time taken; +18% output quality | Noy & Zhang, Science 2023; randomised, 453 college-educated professionals | Production |
| Analysis and strategy tasks inside AI’s competence | +12.2% tasks completed, 25.1% faster, >40% higher quality | Dell’Acqua et al., HBS Working Paper 24-013; pre-registered, 758 BCG consultants | Production |
| Ad creative generated end-to-end by AI | Up to +19% click-through in field settings versus human-expert ads; AI-modified human ads: no significant improvement | Lee, Todri, Adamopoulos & Ghose, NYU Stern / Emory; laboratory experiment plus a field experiment buying display impressions on the Google Display Network | Demand |
| Whole email-marketing creative programme | LLM cell projected at 2.05% higher annual net profit than the human cell — while the human cell still produced the highest measured gross profit; on the same projected basis, prompt-based GPT beat human writers by 8–9% in the third trial | Dubé & Xu, Quantitative Marketing and Economics, published 13 January 2026; three sequential RCTs at Wine Access (open preprint of record) | Production (the win came from salary savings) |
| Personalised video ads generated end-to-end by AI | +9.4% click-through versus personalised image ads; +6.5% versus generic video | Kapoor & Kumar, Marketing Science 44(4):733–747 (2025); three-arm randomised field experiment, 21,000+ consumers, WhatsApp delivery (MIT IDE summary) | Demand |
Read the Wine Access row twice. The machine did not sell more wine. It sold roughly the same wine for less payroll, and the profit arrived on the cost line. Anyone quoting it as proof that AI grows revenue has not read past the abstract.
Want this done for you? We book qualified sales appointments on a Pay-Per-Result basis — you only pay for calls that actually land in your calendar.
How do I tell whether a generative AI marketing claim is real?
Apply three questions to any number a vendor, a conference stage or an internal deck puts in front of you. We call it the unit–comparator–custody test, and most claims fail on the second one.
- Unit. What was counted? Assets produced, drafts written, hours saved — that is a cost claim. Replies, qualified conversations, bookings, revenue — that is a demand claim. A claim that counts “content pieces” and is priced as growth is mis-sold, whatever the percentage.
- Comparator. What did the treated group get measured against? A randomised control arm, a staggered rollout, a held-out geography — those are comparators. “Before and after we adopted it” is not, because the quarter changed too. No control arm means no effect size, only a testimonial with a decimal point.
- Custody. Who held the data? The six studies above were run by people with no equity in the tool. A vendor measuring its own product, publishing only the accounts that worked, is marketing its own theatre.
Instrumenting your own baseline so you can run that comparator yourself is a separate job; this page is for reading other people’s numbers.
The generative AI use cases with no published effect size
Method. Searched on 17 September 2026 across Google Scholar, NBER, SSRN, arXiv and the marketing journals (Marketing Science, Journal of Marketing Research, Quantitative Marketing and Economics) for randomised or quasi-experimental evidence naming each application, then re-run the same day as a deliberate hunt for counter-examples. One row did not survive: AI-generated video advertising, which has a randomised field experiment and now sits in the table above. The rows below survived, each narrowed to the exact claim the search supports. That is not proof they do not work. It is proof that nobody has shown that they do.
| Use case | What is usually claimed | Controlled evidence found |
|---|---|---|
| Synthetic personas / AI “focus groups” | Replaces respondent research | None measuring a downstream marketing outcome; the published work validates answer accuracy against human samples |
| Brand-voice / tone-of-voice engines | Consistency lifts recall and conversion | None |
| AI blog content published at volume | More pages, more organic traffic | None |
| AI social posting at cadence | Reach and engagement lift | None showing a lift |
| Out-of-the-box AI lead scoring | Better prioritised pipeline | None for generic pre-trained models |
| AI sentiment dashboards and “insight” summaries | Faster, better decisions | None linking the summary to a decision outcome |
| AI-personalised landing-page copy (1:1 dynamic) | Conversion-rate lift | None at 1:1 granularity |
| AI “content optimisation” scores | Higher rankings or more citations | None for the vendor scores themselves. GEO (KDD 2024) reports up to +40% visibility in generative engines, but it tests content-editing tactics, not a score |
A use case with no published effect size is not automatically theatre — it is automatically an experiment, and it should be funded like one. The failure is not trying these things. The failure is putting them in a business case at a number somebody invented.
If we can’t make you money, we don’t deserve yours.
Pay-Per-Result pricing — performance-based alignment.
The two measured penalties nobody puts in the pitch deck
Generative AI has published negative effects too, from the same calibre of study, and they are almost never quoted.
Disclosure costs click-through. In the NYU Stern and Emory study, telling consumers an ad was AI-generated cut effectiveness by up to 31.5%. The same direction shows up in trust research: across thirteen experiments, Schilke and Reimann found that actors who disclose their AI usage are trusted less than those who do not. That is uncomfortable, because platform rules and honest practice both push towards labelling. Budget for the penalty rather than pretending it is not there.
Outside its competence, AI makes people worse. In the BCG trial, on a task chosen to sit beyond the model’s reach, consultants using AI were 19 percentage points less likely to reach the correct answer than those working without it. Confident, fluent and wrong is the failure mode, and it bites hardest where a marketer is least able to check — pricing logic, market sizing, compliance wording.
What our own numbers show, including the one that fails our own test
We publish AI-assisted answer pages at scale, so we can show the production side from the inside. Across leadsnow.ai between 30 August and 13 September 2026, Google Search Console recorded 2,472 distinct queries, 31,259 impressions and 41 clicks — a 0.13% click-through rate across a library of several hundred pages. Generative AI made the pages cheap. It did not make anyone want them. Our first-party AI citation polling data tells the same story from the retrieval side.
Where our own number lives is conversation, not content: 50,769+ AI-booked sales appointments since 2017 and more than a million leads generated, which is the same production–demand line running through conversational AI for sales and the rest of our AI for business work. And now the honest part: that figure fails leg two of our own test. It is a cumulative operator record with no control arm, published with its method on our methodology page — which also discloses that our 7x average sales lift has a median closer to 4x. In our own client work we typically see roughly 3x from speed to lead alone; that is an operator claim, first-hand, not research, and you should read it exactly as sceptically as you read a vendor’s.
Cost line or growth line: where to book each application
Evidence strength should decide which budget line an application is approved against, and what you may promise.
| Evidence you actually have | What the number measures | Book it as | What you may not claim |
|---|---|---|---|
| Multi-thousand-person randomised or staggered trial (e.g. +14% throughput) | Hours and throughput | Cost line; payback in hours saved | Any pipeline or revenue forecast |
| Display-network field experiment on click-through (up to +19%) | Demand, in one context | Growth line, at pilot size, with your own A/B running | That it transfers to your channel, format or market |
| Published RCT where the AI win came from payroll (+2.05% projected net profit) | Cost avoided | Cost line only | That output quality improved |
| Vendor case study, no control arm | Nothing measurable | Test budget only | Anything at all in a business case |
| No published evidence (the negative list above) | Unknown | Time-boxed experiment with a kill date | A process, a headcount plan or a target built on it |
The operating cost people forget: every row above still needs a human who can tell good output from confident nonsense, and that reviewer is the scarce input, not the model. At low volume that is one marketer’s afternoon; at volume it becomes a role, because review cost grows with output while the licence cost does not. That arithmetic is why some organisations run this in-house and some hand over the outcome — see how we scope it in AI marketing services, where we are paid on booked qualified appointments rather than on retainers or seats.
Frequently asked questions
What is generative AI in marketing?
Using a generative model to produce marketing artefacts — ad copy, images, email creative, landing-page variants, sales replies and research summaries — rather than to target or serve them. Targeting and bidding systems have used machine learning for years; the generative part is specifically about producing the asset, which is why almost all of its measured gains are production-cost gains.
Does generative AI actually increase sales, or just output?
Mostly output. Two applications have a measured demand effect: fully AI-generated ad creative, which lifted click-through by up to 19% against human-expert ads, and AI-personalised video ads, which lifted it 9.4% against personalised image ads. In the Wine Access email trials the human-written cell still produced the highest measured gross profit; the AI cell won only on projected annual net profit, after salary savings were deducted.
Does labelling an ad as AI-generated hurt performance?
The published evidence says yes. Disclosure reduced advertising effectiveness by up to 31.5% in the NYU Stern and Emory study, and thirteen experiments by Schilke and Reimann (University of Arizona) found that people who disclose AI usage are consistently trusted less, an effect they trace to reduced perceptions of legitimacy. Check your ad platform’s disclosure rules before treating that as a choice.
Why do the biggest AI productivity gains go to junior staff?
Because the model raises the floor faster than it raises the ceiling. In the 5,179-agent support study the average gain was 14%, but novice and low-skilled agents improved 34% while experienced agents barely moved. Expect your gains where your least experienced people are, and expect near-zero where your best writer already works.
Is “AI content at volume” a real strategy?
No controlled study we could find reports an effect size for it, and our own site is a cautionary data point: 31,259 impressions and 41 clicks across the 30 August to 13 September 2026 window. Volume is cheap to make and expensive to make anyone read.
Pay-Per-Result appointments
See if we’re a fit
We book qualified sales appointments for you and you pay on results, not retainers. Our booking page asks a few quick questions so you find out in two minutes whether that model suits your business.
- 50,769+ appointments booked without cold calling.
- Pay-Per-Result pricing — you pay for booked, qualified calls.
- Pick your own time on our live calendar, no phone tag.
