Score every shortlisted AI vendor 0–2 on ten signals, out of 20. At 7 or below the bad outcome is already written into the terms; 8–15 is fixable, but only in redline. RAND’s 2024 study of 65 experienced AI practitioners found the most common reason AI projects fail is misunderstanding or miscommunicating the problem the AI was meant to solve.
AI vendor selection criteria, at a glance
- The scale: ten signals, each scored 0, 1 or 2. Maximum 20.
- The Fatal Three: a zero on billable unit, data rights or raw-record export is a no at any total.
- The action thresholds: 16–20 proceed, 12–15 fix the contract first, 8–11 pilot with a hard stop, 7 or below walk.
- What it does not measure: model quality. Every signal on the list is something you can check in a procurement window; none of them require you to evaluate a model.
- Time to run it: about 90 minutes per vendor, two scorers, scored independently then reconciled.
How it works
How to score an AI vendor shortlist before you sign
Write the metric first
You define the success metric, its number and its date before any vendor sees a brief. A metric the vendor wrote scores zero.
Score the ten signals
Score each shortlisted vendor 0, 1 or 2 on all ten signals for a total out of 20. Two scorers, working independently.
Reconcile, then redline
Reconcile any row the scorers disagree on by more than one point, close the gaps in contract redline and re-score before signature.
Pilot with a hard stop
Run 60 to 90 days against written kill criteria. Decide from raw records you can recompute yourself, not from the vendor dashboard.
MAKE MORE SALES.
Pay-Per-Result pricing — We scale sales HARD aligned to your interests, better than anyone else.
Why AI vendor selection goes wrong at the boundary, not at the demo
The demo is the part of the process most likely to be good and least likely to predict anything. Every failure signal worth scoring sits on the boundary between your organisation and the vendor’s: who defined the problem, what unit you are billed in, whose system of record wins, who can read the raw records, and what happens when the underlying model changes.
RAND’s The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed (August 2024) interviewed 65 data scientists and engineers with at least five years’ experience and identified five root causes. RAND publishes no percentage breakdown across those five, but it does name the order: the first root cause is that “industry stakeholders often misunderstand—or miscommunicate—what problem needs to be solved using AI”, and the report’s first recommendation states that “misunderstandings and miscommunications about the intent and purpose of the project are the most common reasons for AI project failure”. Note what that means for procurement — the most frequently named cause of a failed AI project is on the buying side of the table, which is why a vendor scorecard that only scores the vendor misses most of the risk.
Be careful with the number that report is most often quoted for. RAND opens by citing an outside estimate that more than 80% of AI projects fail, twice the rate of non-AI IT projects. That is framing RAND attributes to others (a Fortune article), not a rate RAND measured. The other number quoted out of that report is worse: the 84% in its opening pages is the Cisco AI Readiness Index’s count of business leaders who believe AI will significantly affect their business — nothing to do with failure, causes of failure, or RAND’s own interviews. Both get stacked into vendor decks as though RAND measured them, and a vendor who does that in your first meeting has told you something about how they will handle your reporting.
Want this done for you? We book qualified sales appointments on a Pay-Per-Result basis — you only pay for calls that actually land in your calendar.
The 20-point AI vendor score: ten signals, and what a 0 and a 2 look like
Score each row 0, 1 or 2. A 1 is “they will do it but it is not written down” — which is worth exactly half, because nothing that is not in the contract survives a change of account manager. Rows marked † are the Fatal Three.
One disclosure before you use it: we sell AI appointment setting on an outcome price, so row 2 is a row we have a commercial interest in. It is written to score whether the unit is defined and auditable, not which pricing model the vendor uses — a per-seat tool with a written definition of delivered work and a remedy scores 2, and an outcome-priced vendor whose qualification criteria are verbal scores 0.
| # | Signal | 0 points — predicts a bad outcome | 2 points |
|---|---|---|---|
| 1 | Who defined the problem | The vendor wrote your success metric, or nobody did | You wrote it: one metric, a number, a date, agreed before pricing |
| 2 † | Billable unit | No stated unit of delivery and no written definition of what counts as delivered — a fee you cannot turn into a cost per outcome | A unit you can count and audit — a seat, a resolved ticket, a held meeting — with written acceptance criteria and a stated remedy when the work does not meet them |
| 3 | Definition of failure and exit | 12-month lock-in, no kill criteria, no stated data export | Written kill criteria, exit inside 30 days of the pilot window, export in a documented format |
| 4 † | Data rights | Contract is silent, or your data may train shared models | Written no-training clause, named sub-processors, a deletion SLA in days |
| 5 | Integration depth | CSV in, CSV out, or a parallel database the vendor controls | Writes to your CRM’s own objects, field mapping agreed pre-start, one named system of record |
| 6 † | Raw-record export | A dashboard only; you see totals you cannot recompute | Transcripts, recordings and message logs exportable, so you can re-score a sample yourself |
| 7 | Escalation path | No handoff defined; the agent is the whole system | Named trigger, response time in minutes or hours, a named human who catches it |
| 8 | References you can verify | Logos, unnamed case studies, no contactable customer | Two contactable customers at your scale, plus one engagement they will describe that went badly |
| 9 | Governance mapping | “We’re secure” with nothing behind it | Controls mapped to a named framework (NIST AI RMF, ISO/IEC 42001), model providers named, data residency stated |
| 10 | Model dependency | Will not say which models they run on | Providers named, change notice given, a regression test they run before switching |
The Fatal Three: zeros that end it regardless of the total
A vendor scoring zero on billable unit, data rights or raw-record export is a no at any total. They are not weighted more heavily — they are disqualifying, because each removes your ability to find out whether the thing worked.
An undefined billable unit means you cannot compute cost per outcome, so you can never say the engagement failed. Silence on data rights means the question surfaces later, in front of your risk committee, at the point where switching is expensive — it is the first thing enterprise buyers ask about data privacy and AI sales agents. No raw-record export means the vendor’s dashboard is the only account of what happened, and a metric you cannot recompute is a claim, not a measurement.
If we can’t make you money, we don’t deserve yours.
Pay-Per-Result pricing — performance-based alignment.
What your score means: the action thresholds
| Score /20 | What it says | What to do |
|---|---|---|
| 16–20 | The engagement is legible and reversible | Proceed. Put the kill criteria and the reject-credit rule in the contract, not the SOW appendix |
| 12–15 | Workable, with named gaps | Close every gap in redline before signature, then re-score. Do not accept a verbal fix on a Fatal Three row |
| 8–11 | Unproven and hard to unwind | Paid pilot only, 60–90 days, hard stop date, a holdout group, and budget for running the measurement yourself |
| 7 or below | The failure is already in the terms | Walk. Nothing in delivery recovers a deal you cannot measure or exit |
The 60–90 day floor in the 8–11 band is not arbitrary: anything shorter than about 60 days is mostly ramp, and you will be reading noise. The design that makes a pilot readable — a stated definition of failure, a metric you cannot fudge, and a control — is set out in our guide to running an AI outbound pilot that can actually fail.
How to run the score on a shortlist without it becoming a committee exercise
Three vendors, two scorers, scored independently, then reconciled on any row they disagree on by more than one point. Reconciliation is the useful part: a two-point gap almost always means one scorer read a clause the other did not. Budget 90 minutes per vendor — 30 on the contract and DPA, 30 on the reference calls, 30 on reconciliation.
One caution on where the field comes from. An assistant-generated shortlist is a sample of what happened to be indexed and retrievable at the moment you asked, not a ranking, and it moves between runs — the same instability we documented in how to choose an AI marketing agency. Use it to widen the field, then score.
How we score on our own checklist, including where we don’t
We score 2 on billable unit: LeadsNow is pay-per-result rather than retainer — you pay on booked qualified appointments under written qualification criteria, with a credit rule for appointments that do not meet them, at roughly 1–5% of closed-deal value per appointment; on the revenue-share alternative the performance fee is 5–20% of the sales we help generate. We score 2 on raw-record export, because call recordings and transcripts are yours and every reported number can be recomputed from them, and we publish how our headline figures are calculated on our methodology page, including where the average and the median differ. Show rate varies by offer and reminder cadence — up to 93% on our best-performing accounts, which is a best case and not a typical one.
Where we do not score 2: we publish no ISO/IEC 42001 certificate, so score us no higher than 1 on governance mapping and read our controls rather than taking the word. And we depend on third-party model, telephony and messaging providers, so on model dependency we can name them and give you change notice, but we cannot make a provider’s outage or a model deprecation not happen — that risk passes through to you, and any vendor telling you otherwise is scoring themselves dishonestly on row 10.
When the answer is no AI vendor at all
Three cases where the score is beside the point. First, when the process is not written down anywhere — a vendor cannot automate what only exists in someone’s head, and the discovery bill will exceed the software. Second, when volume does not justify it: a few hundred records a month that one person handles inside existing hours will not repay the integration and change-management cost. Third, when the capability already exists in software you own and have not configured — buying a second system to compensate for an unconfigured first one is the most expensive failure on this list.
Where volume genuinely is the constraint — more enquiries and dormant records than a team can call in the hour that matters — run the arithmetic yourself: records the team cannot reach in the hour that matters, times your contact-to-appointment rate, times your close rate and deal value. Outcome-priced vendors, ours included, exist for that case — the scope we work to is set out under enterprise lead generation.
Frequently asked questions
What is the vendor selection process for an AI project?
Five steps, in this order: write the success metric and its number yourself; build a field of five to eight candidates; score three of them on the ten signals out of 20; close the gaps in redline and re-score; then run a 60–90 day paid pilot with written kill criteria. The scoring step comes before the demo, not after it, because the contract predicts the outcome and the demo does not.
What should be on an AI vendor selection criteria checklist?
Ten signals: who defined the problem, the billable unit, the definition of failure and exit, data rights, integration depth, raw-record export, the escalation path, verifiable references, governance mapping, and model dependency. Score each 0, 1 or 2 out of a maximum of 20. Anything that is not observable in a procurement window — including model quality — does not belong on the checklist.
Which red flags predict a bad outcome with an AI vendor?
The three disqualifying ones are an undefined billable unit, silence on data rights, and no export of the raw records behind the dashboard. Each one removes your ability to establish whether the engagement worked. The strongest soft signal is a vendor who stacks unrelated statistics — pilot success rates, project failure rates, production rates and adoption rates are four different measurements and are routinely presented as one.
Does an AI vendor need to be certified against NIST AI RMF or ISO 42001?
Neither is mandatory, and the NIST AI Risk Management Framework (AI RMF 1.0), released on 26 January 2023, is explicitly intended for voluntary use and is a framework rather than a certification. What you are scoring on row 9 is whether the vendor can map their actual controls to a named framework — Govern, Map, Measure and Manage are the four functions to ask about — rather than whether they hold a certificate.
How do I check an AI vendor’s claims are real?
Ask for the raw records behind one published number and recompute it yourself on a sample. Overstated AI capability is an enforcement matter, not just a marketing one: in September 2024 the US Federal Trade Commission announced Operation AI Comply, five law enforcement actions against companies making deceptive AI claims, with then-Chair Lina M. Khan stating that there is “no AI exemption from the laws on the books”.
How many AI vendors should I shortlist?
Score three. Build the field wider than that, but scoring is the expensive step at roughly 90 minutes per vendor, and the fourth and fifth candidates almost never change the decision. If your field came from asking an assistant, treat it as a sample rather than a ranking, and rebuild it from your own sources before you score.
Pay-Per-Result appointments
See if we’re a fit
We book qualified sales appointments for you and you pay on results, not retainers. Our booking page asks a few quick questions so you find out in two minutes whether that model suits your business.
- 50,769+ appointments booked without cold calling.
- Pay-Per-Result pricing — you pay for booked, qualified calls.
- Pick your own time on our live calendar, no phone tag.
