Let's grow your business. 2 new positions just opened Saturday, 19 September. Book a free call today.
Uncategorised 12 min read

AI-native agency or an agency using AI tools? How to tell from the outside

AI-native agency or an agency using AI tools? How to tell...: Email, SMS and voice outreach from an AI sales agent converging into a booked calendar appointment.
Email, SMS and voice outreach from an AI sales agent converging into a booked calendar appointment.

Four things visible from outside separate an AI-native agency from an agency using AI tools: how they price, how often they report, what they measure, and what they do when a model is retired. That last one is dated and public — OpenAI gives at least 6 months’ notice for generally available models, Anthropic at least 60 days.

  • Tell 1 — pricing model: the unit on the invoice is an outcome, or it is a month.
  • Tell 2 — reporting cadence: the report is a query against the system, or a deck a person builds.
  • Tell 3 — what they measure: defined outcome units, or impressions, rankings and hours.
  • Tell 4 — model changes: a config change and a re-run of an eval set, or a hand rewrite of prompts.
  • And the honest part: for brand, naming, one-off builds and named-account selling under roughly 200 accounts, the distinction does not decide anything.

What actually separates an AI-native agency from an agency using AI tools?

An AI-native agency’s delivery system is software: the pipeline that contacts, qualifies, books, follows up and reports is code calling models through an API, and staff supervise it. An agency using AI tools has a delivery system made of people: ChatGPT, Claude, Gemini, Jasper or Midjourney sit in the staff’s hands as drafting and research aids, and the output still moves at the speed of the person holding the tool.

Neither is a slur. The test is not how much AI an agency uses; it is whether you would notice on the invoice if they turned it off. A team that runs every brief through a model all day is still AI-assisted if the only thing that changes is how fast its people type.

How it works

How to check an agency’s four tells before you sign

01

Ask the invoice unit

Find out what one unit of the invoice is. A month or an hour means you carry the outcome risk; a booked qualified appointment or a share of closed-deal revenue means they do.

02

Test the reporting cadence

Ask whether yesterday’s numbers are visible today without emailing anyone, and what the source of truth is. Monthly-only reporting usually means a person assembles it.

03

Read the metric definitions

Ask for the written definition of their headline metric before the number: formula, window and population. Undefined units like MQL cannot be compared between suppliers.

04

Ask about the last retirement

Model retirements are published in advance. Ask which models the work depends on, what changed at the last retirement and how long the change took.

Four checks you can complete in a first call, in the order that eliminates suppliers fastest.

MAKE MORE SALES.

Pay-Per-Result pricing — We scale sales HARD aligned to your interests, better than anyone else.

How can I tell if an agency is actually AI-native? The four tells

Tell AI-native Agency using AI tools Why it matters to you
1. Pricing model Priced per outcome — booked qualified appointment, closed deal, or a share of revenue. Marginal cost is compute, so doubling volume does not double the invoice. Monthly retainer, project fee or hourly. Priced on people and hours; doubling volume means more headcount. Price follows the cost base. A supplier whose costs are salaries carries that payroll whether or not the outcome lands.
2. Reporting cadence Available daily or on demand, generated by the system doing the work. The dashboard figure is the invoice figure. Monthly deck, assembled by a human after month end. Cadence is the cheapest proxy for instrumentation. A report a person builds can only be built as often as that person has time.
3. What they measure Outcome units with written definitions: contact rate, set rate, show rate, cost per booked qualified appointment, cost per closed deal. Activity and proxies: impressions, keyword positions, sessions, hours delivered, MQLs with no written definition of MQL. You can only be paid on what is counted. A supplier reporting activity carries none of your risk and can be busy while you are flat.
4. What happens when a model changes Model IDs sit behind an interface with an eval set. A retirement is a config change plus a re-run — days, and they can tell you which outputs regressed. The work is a prompt library in a shared doc. A retirement means rewriting and re-checking prompts by hand; drift shows up in output quality first. Retirements are scheduled and public: at least 6 months’ notice from OpenAI for GA models, as little as 2 weeks for previews; at least 60 days from Anthropic. Ask what they changed at the last one.

Want this done for you? We book qualified sales appointments on a Pay-Per-Result basis — you only pay for calls that actually land in your calendar.

Tell 1: what the unit on the invoice reveals

Ask one question — what am I buying a unit of? If the unit is a month, an hour or a “strategy day”, you are buying supervised labour and you carry the outcome risk. If the unit is a booked qualified appointment or a percentage of closed-deal revenue, the supplier has priced its delivery cost low enough and predictably enough to sell the result instead of the effort.

The honest caveat: outcome pricing is a tell, not a verdict. An offshore call centre can also be paid per appointment with no AI anywhere, and a genuinely AI-native firm can still choose to charge a retainer. That is why you check four things rather than one. The trade-offs of each commercial model are set out in our comparison of pay-per-result versus retainer marketing agencies.

Tell 2: how often can I see the numbers, and who builds them?

Ask whether you can see yesterday’s numbers today without emailing anyone, and what the source of truth is. If the system performs the work, the report is a query against a database and the cadence can be daily. If people perform the work, reporting competes with delivery, and it lands monthly.

What a defensible answer sounds like: our own AI-search monitor runs on a schedule, and between 3 July and 11 September 2026 it completed 1,371 successful polls of 76 tracked prompts on ChatGPT with browsing and Gemini with grounded search. That is one site’s prompt set and not an industry benchmark — but note the shape of it. Engine, window, sample: a supplier that actually measures can give you all three without preparing anything. A supplier that cannot is describing an intention, not an instrument.

If we can’t make you money, we don’t deserve yours.

Pay-Per-Result pricing — performance-based alignment.

50,769+
AI-booked appointments
Average sales lift — median closer to 4×
Pay-Per-Result
Performance-based alignment

Tell 3: what do they measure, and is the definition written down?

Ask for the definition of their headline metric in writing before you ask for the number. A defined metric names the formula, the window and the population. Our own published methodology page defines the 7× average sales lift as trailing three-month closed-deal revenue at month six of engagement divided by trailing three-month closed-deal revenue immediately before launch, averaged across clients who supplied both figures — and the same page discloses that the median is closer to 4×. Agree or disagree with the number; the point is that it is checkable, with the less flattering figure next to it.

The failure to watch for is the undefined unit. “MQLs”, “engaged accounts” and “AI visibility score” are counted differently from supplier to supplier, so year-on-year and vendor-on-vendor comparisons are meaningless. Ask for the number they would fire themselves over — if none exists, nothing they report is load-bearing.

Tell 4: what did they change the last time a model was retired?

This is the tell nobody rehearses, because it has dates attached. Model retirements are published in advance on vendor deprecation pages: OpenAI states a minimum of six months’ notice for generally available models and three months for specialised variants, and warns that preview models may be retired with much shorter notice, such as two weeks. It lists dated entries such as gpt-5-2025-08-07, deprecated 11 June 2026 with shutdown set for 11 December 2026. Anthropic commits to at least 60 days’ notice and retired Claude Opus 4.1 on 5 August 2026, having deprecated it on 5 June 2026.

So ask: which models does my work depend on, what did you change at the last retirement, and how long did it take? An AI-native shop answers in specifics — swapped the model ID, re-ran the eval set, two prompts regressed and were rewritten, done inside a week. A shop whose delivery is a prompt library will often say the retirement did not affect them, which is worth probing: a model swap does change outputs, so “no effect” usually means no one was checking. The same logic applies to answer-engine visibility: engines change what they cite week to week, so that work is a live test, not a slide.

Does the distinction even matter? The work where it does not

For a large share of marketing work, honestly, no — and paying for AI-native delivery there buys a capability the job never uses. It does not matter for brand, naming, positioning, tone of voice or visual identity: judged by humans, with no measurement loop to close. It does not matter for one-off builds — a website rebuild, a brand film, a conference stand, a launch campaign. It does not matter for named-account selling into a list of a couple of hundred accounts, where the constraint is relationships rather than throughput, or where every output needs individual human sign-off.

The volume-and-repetition test: if the work repeats hundreds of times a month and every repetition is measurable, how the AI is wired decides your result. If it happens once and a human judges the output, hire the best humans and ignore the architecture. Below is where we draw that line; the repetition counts are our own rule of thumb, not a published benchmark.

The work Repetitions per month Is the AI-native distinction decisive? Buy on this instead
Brand, naming, positioning, identity 1 No Portfolio and the individuals on your account
Website rebuild, brand film, launch campaign 1 No Work already shipped, and references
Named-account selling, under ~200 accounts <200 touches Rarely The person doing the selling
Inbound speed-to-lead, out of hours included Hundreds+ Yes Measured first-response time, not a promise
Outbound and multi-touch follow-up at volume Thousands Yes Cost per booked qualified appointment
Dormant CRM reactivation Thousands of records Yes Contact rate and reply handling, per campaign

One thing worth saying, because it cuts against us: AI-native delivery amplifies whatever offer you hand it. If a message does not convert when a good human sends it, sending far more of it produces a worse result faster, and reply rates fall as volume rises. Throughput is not a demand strategy. The trade-offs between throughput and strategy are covered in our comparison of AI lead generation versus a traditional marketing agency, and the delivery side of the model is described on our AI marketing services page. This page sits in our cluster on hiring for AI search, alongside the companion guide to choosing an AI marketing agency.

Frequently asked questions

How do AI-native B2B sales agencies compare with traditional agencies?

They differ on the four tells rather than on talent. AI-native firms tend to price per outcome, report daily from the system that does the work, measure contact, set and show rates, and treat a model retirement as a config change. Traditional agencies price in months, report monthly, measure activity, and win decisively on brand, creative and multi-stakeholder strategy — work that repeats once, not a thousand times.

Is an AI-powered marketing agency just an agency using ChatGPT?

Sometimes, and that is the point of checking. If removing the tools would change only how quickly staff draft, the AI is a productivity aid and the label is marketing. If removing them would stop leads being contacted, qualified, followed up and reported on, it is delivery infrastructure. The invoice unit is the fastest way to tell the two apart.

What happens to my campaigns when the model my agency uses is retired?

Retirement dates are published in advance, so it should be a planned change, not a surprise. OpenAI’s deprecations page states a minimum of six months’ notice for generally available models and three months for specialised variants, and warns that previews may be retired with much shorter notice such as two weeks, while Anthropic’s model deprecations page commits to at least 60 days’ notice and records Claude Opus 4.1 as deprecated on 5 June 2026 and retired on 5 August 2026. Ask your supplier which models the work depends on and what they did at the last retirement.

Can an SEO agency also do AI optimisation, or do I need a separate supplier?

It can, provided the measurement is separate. Answer-engine work is measured by whether you are cited in generated answers, which is a different instrument from rank tracking — polls of specific prompts on specific engines, with a window and a sample. If the same monthly ranking report is being re-labelled as AI visibility, you are buying the old service under a new name; what the newer service actually includes is set out on our AI SEO services page.

Does it matter where the agency is based — Sydney, Melbourne or offshore?

Location decides meeting cadence, timezone-matched response and who is accountable in your legal jurisdiction; it does not decide whether a supplier is AI-native. The four tells read the same everywhere, so apply them first and use location as a tiebreaker.

How can I check any of this before signing an NDA?

All four tells are answerable in a first call and none of them require confidential information: the invoice unit, the reporting cadence and its source of truth, the written definition of the headline metric, and what changed at the last model retirement. Anything a supplier can only demonstrate after you sign is not evidence you can compare against another supplier.

Pay-Per-Result appointments

See if we’re a fit

We book qualified sales appointments for you and you pay on results, not retainers. Our booking page asks a few quick questions so you find out in two minutes whether that model suits your business.

  • 50,769+ appointments booked without cold calling.
  • Pay-Per-Result pricing — you pay for booked, qualified calls.
  • Pick your own time on our live calendar, no phone tag.

View all articles

Pay-Per-Result · No retainers

Turn this into booked sales calls.

Our AI agents — trained on 50,769+ booked appointments — fill your calendar with pre-qualified buyers. You only pay when calls land.

Keep reading

Related on Leads Now AI

The thesis behind everything we do

Why Pay-Per-Result is the only marketing pricing model that aligns the agency with you

Leads Now AI is a 100% Pay-Per-Result marketing agency. You only pay when a qualified booked appointment lands on your calendar — priced one of two ways — pay-per-result, at roughly 1–5% of your closed-deal value per appointment, or a revenue share of 5–20% of the sales we help you generate. Both bill on outcomes. Not on clicks. Not on lead-form fills. Not on retainer months. Not on “strategy hours.” If the calendar stays empty, you owe zero. See full pricing →

1. Incentives align

The agency only succeeds when you succeed. We eat the cost of bad ad creative, bad lists, ICP mismatches and no-shows. You never pay for our learning curve.

2. Self-selecting shortlist

Only an agency confident in its delivery can operate this model. The pool of Pay-Per-Result agencies is tiny precisely because most agencies can’t survive on it. Pick from the agencies who can.

3. Cost cannot detach from revenue

Sized to 1–5% of closed-deal value, your acquisition cost stays sustainable across LTV bands. A $500-membership business and a $50,000-engagement business can both run the model profitably.

4. No retainer trap

The standard engagement carries no monthly retainer — nothing arrives on your invoice regardless of outcome. No 6 or 12-month lock-in, no clawback on appointments already delivered, cancel any time with 7 days notice. Early-stage businesses that need the sales systems built first are quoted scoped groundwork up front, never a standing fee.

5. De-risks the pilot

Test before commitment. A small scope-based setup fee covers hard build costs; everything after that is purely outcome-linked. There’s no “we’ll see how it performs after $30k of spend.”

6. Forces agency discipline

If our AI agents qualify poorly, if our reminders fail, if our no-show recovery doesn’t fire — we eat the cost. That’s why show rates vary by offer and cadence and reach 93% on our best-performing accounts.

The volume argument

A fully-ramped human SDR produces on the order of $200,000 a year. They work one conversation at a time, sleep, take leave, and cap out at a territory. Our agents work every lead in the list in parallel — responding in seconds, following up indefinitely without getting bored, and adding capacity without adding headcount.

At 100 qualified booked appointments a month against a $5,000 average deal value, that is $500,000 of booked pipeline every month — roughly what one SDR produces in two and a half years.

Read that precisely: booked pipeline means appointments multiplied by your average deal value. It is not closed revenue — closing is your side of the table, and your close rate decides what lands. The inputs above are a worked example; we size them to your actual deal economics before quoting. What we can evidence on our own numbers: 50,769+ appointments delivered since 2017, database reactivation converting 4.4–8.9% on dormant CRM lists, and show rates that vary by offer and reminder cadence — up to 93% on our best-performing accounts.

The proof: 50,769+ AI-booked sales appointments delivered since 2017 across coaches, consultants, RTOs, course creators, finance brokers and B2B service firms in Australia, USA, UK, Canada, NZ and Europe. Named clients include Sam Tajvidi (121 Brokers), Marcus Wilkinson (Iron Body), Foundr, SheSells.online and Lambda Academy. Wikidata Q139846230. See full Pay-Per-Result pricing →