Report four numbers to the board: realised value against plan, cost per outcome against a locked baseline, the share of AI spend in production with a named owner, and control exceptions. Keep self-reported hours saved off page one. In McKinsey’s 2026 survey, 80% of respondents said AI had improved their own productivity, but only 37% said it had contributed to their organisation’s EBIT.
Sources last checked: . Every external figure on this page links to the publisher that produced it, and was re-read at that source before publication.
- The template: the 4+1 board page. Four metrics on page one, and one metric that misleads, moved to the appendix with a label on it.
- The four: realised value ÷ plan, cost per outcome vs baseline, production share of AI spend, control exceptions.
- The one that misleads: hours saved and active users. They show use, not earnings.
- Cadence: quarterly. In Deloitte’s 2025 board survey (n=695), 17% of respondents said their board takes AI at every meeting, and 31% said it is not on the agenda.
- The slow line: realised value. On our own planning schedule, its first honest read comes in quarter 3 or 4, and it slips when nobody locked the baseline before go-live.
What do I actually report to the board about our AI programme?
Report outcomes the board can hold management to, not activity. Deloitte’s Governance of AI report (2nd edition, surveyed January–February 2025) puts the board’s performance question plainly: what metrics determine whether AI has contributed to strategic goals, and at what cadence are they measured and reported? The template below answers it on one page, and finance can re-derive every number in it.
| # | Metric | Formula | System of record | The board question it closes |
|---|---|---|---|---|
| 1 | Realised value vs plan | Finance-validated incremental benefit to date ÷ business-case benefit planned to date | General ledger: budget lines that fell, plus gross margin on holdout-tested incremental revenue | “Is the money arriving as promised?” |
| 2 | Cost per outcome vs baseline | (AI run cost + remaining human cost of the process) ÷ outcomes delivered, compared with the pre-go-live baseline | CRM or service desk for outcomes; invoices and payroll allocation for cost | “Is each unit of work cheaper, or is there more of it for the same cost?” |
| 3 | Production share of AI spend | AI run-rate spend on production use cases with a locked baseline and a named owner ÷ total AI run-rate spend | AI use register, reconciled to expense and licence records | “How much of what we spend can we actually measure?” |
| 4 | Control exceptions | Customer-affecting AI incidents per 10,000 AI-handled interactions, plus register coverage (registered uses ÷ uses found in reconciliation) | Output log, complaints register, AI use register | “Is it safe, and do we know where all of it is?” |
| +1 | Hours saved and active users (appendix only) | Self-reported hours × headcount; seats used ÷ seats bought | Staff surveys, licence telemetry | “Is it being used?” It is an input, not a result. |
The 4+1 board page puts one test to every figure: could it fall while the programme is working, or rise while it is failing? Hours saved can do both. If you are still building the case, metric 1’s plan line comes from the four-number AI business case your CFO will test. Metric 4 comes from the register and log in a four-artefact AI governance framework.
How it works
Building the quarterly AI board page
Lock baselines before go-live
Define numerator, denominator and cohort for each metric and date them. Keep a holdout group untouched.
Name one owner per metric
Finance owns realised value, the process owner owns cost per outcome, the programme lead owns spend share, risk owns exceptions.
Score against written thresholds
Mark each metric green, amber or red against the thresholds agreed in the business case. Two ambers in a row force a decision.
Report four, append one
Put the four outcome metrics on page one. Move hours saved and active users to the appendix, labelled capacity, not value.
MAKE MORE SALES.
Pay-Per-Result pricing — We scale sales HARD aligned to your interests, better than anyone else.
Why hours saved is the AI metric that misleads boards
Hours saved misleads because it measures how people feel about the tool, and people are poor judges of that. METR’s randomised trial (published 10 July 2025) had 16 experienced open-source developers work 246 real issues from repositories they knew well. They expected AI to speed them up by 24%. With AI they took 19% longer. Afterwards they still believed AI had made them about 20% faster. METR said in February 2026 that newer tools have probably shifted the result and it is redesigning the study, so this is no verdict on today’s tools. It does show that a self-report and a measurement can disagree in direction as well as size.
The same gap appears across whole companies. McKinsey’s The state of AI in 2026 (1,719 respondents, surveyed 4 May–8 June 2026) found eight in ten said AI had improved their own productivity. The share saying AI had contributed to their organisation’s EBIT was essentially unchanged from a year earlier, at 37%.
Worked example (illustrative inputs). 2,000 staff report saving 2 hours a week each. Over 46 working weeks that is 184,000 hours. At a $60 loaded hourly cost, the slide says $11.04 million. Now find the budget line that fell by $11.04 million. Usually there is none: headcount, contracts and spend are unchanged, so a finance function credits the benefit at zero. Hours saved become value only when a budget line falls or a measured output rises, and at that point they already appear in metric 1. Keep hours saved in the appendix, labelled “capacity released, not value realised”.
Want this done for you? We book qualified sales appointments on a Pay-Per-Result basis — you only pay for calls that actually land in your calendar.
When is an AI metric red, amber or green?
A threshold turns a board metric into a decision. The thresholds below are our own proposed defaults, to be replaced with the figures in your own business case. They are decision rules we suggest, not industry benchmarks or published standards.
| Metric | Green | Amber | Red | What the board decides at red |
|---|---|---|---|---|
| 1. Realised value vs plan | ≥ 90% of plan to date | 60–89% | < 60%, or amber two quarters running | Re-baseline the case in writing, or stop the use case |
| 2. Cost per outcome vs baseline | At or below baseline after one full sales or service cycle | Up to 10% above baseline, or more than 10% after only one cycle | More than 10% above baseline after two cycles | Management explains which cost line grew, or the use case is scaled back |
| 3. Production share of AI spend | ≥ 75% | 50–74% | < 50% | No new pilots are funded until existing spend has owners and baselines |
| 4. Control exceptions | No open customer-affecting incident; register coverage ≥ 95% | Any incident open over 30 days, or coverage 80–94% | Repeat incident of the same type, or coverage < 80% | The use case pauses until the named signer reports a fix |
Suppose the business case planned $450,000 of benefit by the end of quarter 2 and finance validates $310,000. That is 68.9% of plan, which is amber. If the next quarter is amber too, the programme lead has to re-baseline the case in writing or propose stopping it. An AI programme that has no red threshold on realised value will stay amber indefinitely.
What the board will see in each quarter of an AI programme
A board report on an AI programme fills from the bottom up. Metrics 3 and 4 can be read from the first quarter. Realised value is the last to become readable, and it is the stage that slips most. Give the board this schedule in advance, so a working deployment isn’t killed over a number that could not yet move. The schedule is our planning rule of thumb, not a published benchmark.
| Report | Readable | Shown as “not yet readable, from…” |
|---|---|---|
| Before go-live | Baselines for metrics 1 and 2, locked and dated; register coverage | All actuals |
| Quarter 1 | Metrics 3 and 4; the response-time parts of metric 2 | Metric 1; cost per outcome for anything with a buyer decision inside it |
| Quarter 2 | Metric 2, once one full sales or service cycle has closed | Metric 1 |
| Quarters 3–4 | First honest read of metric 1, against a holdout | EBIT contribution (in our view, rarely attributable within 12 months) |
Realised value slips for three reasons, and all three are set before go-live. No baseline was locked, so “before” gets redefined after the fact. No holdout group was kept, so incremental benefit cannot be separated from what would have happened anyway. And the benefit was capacity rather than cash, so nothing reached a budget line. The three-clock rule for AI efficiency metrics explains why a metric moves only as fast as the slowest human decision inside it. That is why quarter 1 cannot show metric 1.
Our own headline figure meets the disclosure half of that standard, not the holdout half. The 7x average sales lift LeadsNow publishes compares trailing three-month closed-deal revenue at month six with the three months before launch, averaged across clients who supplied both. It is a before-and-after comparison with no holdout, so it would not pass metric 1’s incrementality test, and it is an average, not a typical result: the measurement methodology page discloses that the median is closer to 4x. The board should get the same thing: a stated window, a stated denominator, and the median printed beside the average.
If we can’t make you money, we don’t deserve yours.
Pay-Per-Result pricing — performance-based alignment.
What boards are asking about AI in 2025 and 2026
Board oversight of AI is now disclosed far more often than a year earlier. The EY Center for Board Matters reviewed the 80 companies on the 2025 Fortune 100 list that had filed by 31 July 2025. It found 48% disclosed a focus on AI in the risk-oversight section of the proxy statement, up from 16% the year before, and 40% named at least one board committee responsible for AI, up from 11%. Knowledge has not kept pace. In Deloitte’s survey, 66% of respondents said their board has limited to no knowledge or experience of AI, down from 79%.
A board with limited AI knowledge depends on the format it is shown. A formula gives directors something to challenge, and a twenty-slide use-case gallery does not. McKinsey’s 2026 survey also found that AI high performers were twice as likely as other respondents to have defined processes for measuring the impact of their AI initiatives.
Who owns the AI board report, and what it costs to produce
One accountable executive signs the page, and each metric has one owner. Finance owns metric 1, because only finance can say a budget line fell. The process owner owns metric 2. The AI programme lead owns metric 3. Risk or compliance owns metric 4.
Most of the work is in the first cycle: defining numerator, denominator and cohort for each metric, pulling at least 90 days of pre-go-live history, setting aside a holdout, and reconciling the register against expense and single-sign-on records. As our own planning assumption, not a benchmark, allow roughly an analyst-week per use case before go-live, then about a day per use case each quarter. The skill that is hard to find is not analysis. It is someone senior enough to write “amber, second quarter” on a programme the CEO sponsored.
When an outside provider runs part of the process, require its reporting to use your denominator, not its own. Under a pay-per-result contract, where you pay per booked qualified appointment rather than per seat, the invoice already gives you most of metric 2. Metric 1 still has to show incrementality against a holdout. How that works in a large sales organisation, including the SLAs worth writing into the contract, is covered in AI appointment setting for corporate sales teams.
Frequently asked questions
Which board committee should oversee AI?
There is no single convention yet. Among the 80 Fortune 100 companies EY analysed in 2025, 40% disclosed that at least one board committee oversees AI. 21% disclosed oversight by the audit committee and 25% by a non-audit committee, and some named more than one (EY Center for Board Matters, October 2025). Pick the committee that already reviews the metric-4 controls, and send metric 1 to whoever reviews the business case.
How often should AI be on the board agenda?
In our view, quarterly is enough for most programmes. That matches how quickly realised value can change. In Deloitte’s 2025 Governance of AI survey of 695 board members and executives, 17% said their board covers AI at every meeting and 19% once a year, while 31% said AI is not yet on the board agenda. Add an out-of-cycle report only when a metric 4 threshold turns red.
What is the difference between AI governance reporting and AI performance reporting?
Governance reporting shows whether AI use is controlled: the register, the named signers, the log, the incidents. Performance reporting shows whether it is paying. The 4+1 board page puts both on one page. Metric 4 is the governance line, and metrics 1 and 2 are the performance lines.
Should we report hours saved to the board at all?
Yes, but only in the appendix, labelled as capacity released rather than value. Self-reports are unreliable: in METR’s 2025 randomised trial, developers took 19% longer with AI yet believed it made them about 20% faster. Hours saved belong on the front page only once a budget line falls or a measured output rises, and at that point they are already in metric 1.
What do we report if the AI programme has no results yet?
Report the schedule. Show the locked baselines, metrics 3 and 4, and the response-time parts of metric 2, and mark metric 1 as not yet readable, with the quarter it becomes readable. If you cannot name that quarter, the programme has no measurement plan. Fix that before the next board meeting.
Pay-Per-Result appointments
See if we’re a fit
We book qualified sales appointments for you and you pay on results, not retainers. Our booking page asks a few quick questions so you find out in two minutes whether that model suits your business.
- 50,769+ appointments booked without cold calling.
- Pay-Per-Result pricing — you pay for booked, qualified calls.
- Pick your own time on our live calendar, no phone tag.
