A few months ago I wrote about why marketing transformation pilots stall between proof of concept and production: a well-drawn map that never quite becomes the territory. One idea in that piece got a sentence where it deserved an argument of its own: agreeing on graduation criteria before a pilot begins. I want to give it the room it was owed.
A Licence Earned in Stages
A learner driver doesn’t get to carry passengers or drive after dark the day someone decides they’ve gotten pretty good. They earn each stage against a specific, visible condition viz., hours logged, a supervised test passed, a fixed period without an incident, set before they ever got behind the wheel. Nobody in that system is relying on a feeling. The criteria exist precisely so that “ready” doesn’t have to be someone’s judgment call made under the pressure of the moment.
Most transformation pilots have no equivalent. What they have instead is a demo day: a moment when the pilot team shows its best result to an audience predisposed to be impressed, and the room’s enthusiasm substitutes for anything that was actually agreed in advance. That substitution is the real gap between pilot and production. It isn’t that organisations fail to notice a pilot has done well. It’s that “done well” was never defined narrowly enough, early enough, for anyone to say with confidence that the specific bar for moving to production had actually been cleared.
The first useful perspective is to separate proof of possibility (“This works under controlled conditions”) from
proof of readiness (“It works consistently, with ordinary users, errors detectable and correctable, and so on”).

The Demo Day We Mistake for a Graduation
A demo day answers one question: can this work, shown under the best available conditions, to people who want to see it succeed. A graduation criterion answers a different one: has this cleared a specific, pre-agreed bar for running without the conditions that made the demo possible. Those are not the same question, and a strong answer to the first tells you very little about the second.
The tell that an organisation is relying on demo day instead of a graduation criterion is simple: ask, before any pilot starts, what result would make everyone agree it’s ready to scale, precisely, not directionally, and watch how many people can answer without hedging. Most cannot, because the honest answer is “we’ll know it when we see it,” which is a demo-day standard wearing the language of a decision process.
A learner is not declared ready because they drove successfully once with an instructor on a familiar road.
They need to demonstrate competence under conditions that reveal whether the skill is transferable.
A pilot proves that something can work. Graduation proves that the organisation can live with it when it does not.
What a Real Criterion Actually Specifies
A genuine graduation criterion is specific enough to be checked by someone who wasn’t in the room when the pilot ran. Not “improved lead quality,” but a named metric, a named threshold, and a named comparison group: this segment’s conversion rate needs to clear this number against this baseline, sustained across this many weeks, including at least one week of degraded input data, because degraded input is what production actually looks like. Not “the team is comfortable using it,” but a defined handover: the receiving team runs the process unsupervised for a set period, on their own data, with the pilot team unavailable to intervene, before anyone calls it transferred.
The specificity is the entire point. A vague criterion can always be read as met by whoever wants it to have been met. A specific one either was or wasn’t, and that removes the argument at exactly the moment the argument tends to happen when momentum is high, the room is enthusiastic, and nobody wants to be the person raising a hand; after all it is tied to the KPI of important people in the room who likely are operating at a level of visibility that most people are not privy to.
A system is not production-ready when it produces good answers.
It is production-ready when the organisation knows what to do with the bad ones.
When the Criteria Themselves Are Wrong
None of this works if the criteria are treated as unchangeable once set, because some of them will be wrong. A threshold agreed before anyone fully understood the problem is a guess with the authority of a document, and guesses set early are sometimes set badly, too easy to look generous, too strict to be realistic, aimed at the wrong metric entirely because the right one wasn’t visible yet.
The answer isn’t to avoid setting criteria early for fear of getting them wrong. It’s building an explicit, visible process for revising them such as who can propose a change, what evidence justifies it, and a record of what changed and why, so a criterion never quietly loosens under pressure without anyone having to own that decision. A criterion that can be revised honestly is still a criterion. A criterion that gets renegotiated silently the week before it matters was never one to begin with.
Giving a real-life example, at an organisation I worked the ‘transformation leader with clout’ operated only with Miro boards and everything was negotiable and nothing firmed up. This bears no reflection on Miro itself, which is an incredibly useful tool for say design-collaboration at any level. But by the nature of it, it cannot be a ‘contract document’ while it is a wonderful drawing board.
Flexibility is legitimate when it responds to new evidence.
It becomes self-deception when it responds only to the desire to continue.
Who Signs Off, and Why Not Them
The team that ran the pilot should not be the ones deciding whether it has graduated. This isn’t a comment on their honesty. It’s that they have spent months building a case for their own work, they know its strengths better than anyone in the building, and they have every professional incentive to read an ambiguous result generously. That’s a completely ordinary human bias, and building a process that depends on them overcoming it is asking more of individual willpower than any process should.
The sign-off belongs with whoever will actually own the outcome after the pilot team moves on, specifically the function that will run it, the leadership that will answer for it if it fails at scale, checking the criterion as written against the result as measured, with no room for “it’s basically there.”
Production begins when responsibility moves, not when the software goes live.
The gap between pilot and production is not a technology gap. It is a decision gap.
Organisations run pilots to discover whether something works, but fail to define what evidence would justify making someone accountable for it.
A learner driver who has logged the hours and passed the supervised test doesn’t get to declare themselves graduated. Someone else checks the record against the standard and signs the licence. The standard was the entire point of making them wait.