How AI Pilots Get Killed at the Budget Review
When you measure the wrong metric, even great looks bad.
It’s renewal season for the first wave of corporate AI, and the budget reviews have started doing what they’re famous for. At some time this quarter, a manager is going to walk into a renewal with a tool that works, and walk out having lost it. Not because it failed! But because the number it was created to move stopped moving, and no one realised that its value had quietly gone into moving another number.
The cleanest illustration is almost a decade old. In 2017 JPMorgan said a system it called COiN (for Contract Intelligence) was reading commercial-loan agreements and saving 360,000 hours of lawyer and loan-officer time a year, work that until then had been done by hand. It was the number that defined the era: AI would eat the back office, and the savings would compound!
Except hours saved is a zero sum stock - you drain it once, and that’s it. There’s only so much time that any tool can save. The contracts that existed got read; there was no second 360,000 hours waiting the following year, because the manual work it replaced was already gone.
On a scorecard that tracks hours saved per year, a system doing exactly what it was built to do begins, by year three, to look like a system that has stopped working.
And the real continuing benefit was somewhere else entirely. Now it’s in the loan-servicing mistakes that no longer happened because a person was no longer doing the reading by hand, the quiet money that shows up as the disputes you never have and the errors you never unwind. None of that sat on the hours line.
That gap is where working systems die. Picture the same tool in a smaller firm, one without the patience, cash reserves or the senior sponsor. The hours number is flat, the renewal review is on the calendar, the actual value it provides is sitting in a column nobody put on the form. It gets cancelled, and it goes into the file as one more failed pilot.
We are told most corporate AI moves nothing on the books, and most of it genuinely doesn’t. But a real share of what gets killed at renewal WAS working. It was just measured by the number that won it the budget, not the number where the benefit came out. The deployment was real. The metric just didn’t measure it.
This is not anyone being slow. A justification metric is a promise you make to get the money, and a promise has to be quick and clear, so it is always the simple, good solid number kind:
time saved,
tickets closed,
headcount avoided.
The benefits of a tool wired into real work are slower and sideways. A decision made better, a handoff that stops dropping things, a mistake that never gets the chance to compound. Those benefits are what typically gets forgotten to be put on the form.
So the move to make, is dull, and it works, and it has to happen before the review rather than during it.
Write down the number your tool was justified by.
Then write down the number where its value is actually showing up now
And if those are different numbers, change what the STORY the renewal is measured against while you are still the one holding the cards.
What I can’t tell you is whether you will have the standing to make the new number stick. A metric you swap in the week before a review looks exactly like a metric you swapped because the first one embarrassed you.
So the honest version is to choose the second number early, as early as possible. Ideally on the day you ask for the budget. You have to think in first, second and third order effects to get as close as possible to predicting what the real value of the tool will be that far down the line.
That’s how you beat this trap or being measured by the wrong metrics.





