Artifact critiqueAdvancedShipping & Model Lifecycle / Pilot design and POC-to-production / #9
Design the exit criteria that end a pilot in either direction.
The direct answer
Write two numbers before the pilot starts, not one: the score it has to clear to go live, and the score or date that kills it if it doesn't. Hand both numbers to someone outside the pilot team to hold, so the team that benefits from the pilot still looking promising can never quietly move either line.
Do this, in order
Write a pass number and a kill number before the pilot starts, not just a hopeful target.Why: without both written down, "we need more time" can be honest every single month and still be the wrong thing to keep saying by month fourteen.
Give ownership of those two numbers to the business side, not the team running the pilot.Why: the team running it benefits from the pilot staying promising forever, so they're the wrong people to hold the line that ends it.
Tie the kill date to something already on the calendar, like a contract renewal, instead of "when we feel ready."Why: a floating deadline slides quietly. A date sitting next to real money doesn't.
Once a quarter, ask the pilot team to state the pass number, the kill number, and the date, in one sentence.Why: if they can't, or the numbers have moved without anyone signing off, the pilot's already drifted past having real criteria.
Let small, low-stakes trials skip this entirely.Why: forcing the same two-number gate onto a two-week test with no budget line buries the one gate that actually matters.
How to answer this, stage by stage
Eight moves, from naming one real pilot to the line you'd close on.
1
Ground it in one real pilot before naming the framework
Say it like this
"Say a manufacturer pilots a tool that reads a vendor's invoice and checks it against the purchase order and the delivery receipt, so nobody has to match the three by hand. It was scoped to run one quarter, three months, on forty vendors. It's now in month fourteen, still called a pilot, and nobody's said yes or no yet. That's the exact shape of question I want to answer."
Why this works
Grounds "exit criteria" in one real number before any framework language shows up.
2
State your plan in one breath
Say it like this
"I'd run this through GUARD, because the real risk in a pilot with no ending isn't the model, it's who gets to decide when it's done. Who's running it versus who's waiting on it, where an open-ended pilot costs the most, who has no way to force an answer, the actual rule I'd write down, and how I'd catch it if that rule gets ignored."
Why this works
Two seconds naming the plan, not a recited acronym before the real thinking starts.
3
Reframe it as a decision-rights problem, not a scheduling one
Say it like this
"A team running a pilot is never neutral about when it ends. Every extra month it stays a pilot is a month nobody has to defend a number to leadership. So 'we need more time' can be completely honest every single month, and still be the wrong thing to keep saying by month fourteen."
Why this works
This is where a checklist answer and a real answer split apart.
4
Give the one decision: two named numbers, owned outside the pilot
Say it like this
"I'd put two numbers on the board before the pilot starts, not one. A pass number: ninety two percent of invoices matched clean, no person touching them, held for six straight weeks, and we go live. A kill number: if we're not past eighty five percent by month six, we kill it and go back to the vendor doing the matching by hand. Both numbers belong to the VP of finance, not the team running the pilot, because he's the one stuck waiting on it, not the one running it."
Why this works
A mechanism you could point to in the rollout plan, not a value everybody already agrees with.
5
Prove it with the compressed failure
Say it like this
"Here's what happens without those numbers. Match rate climbs the whole time, honestly, seventy one percent, then eighty one, then eighty seven. It's never bad enough to kill and never good enough to call finished, so every review ends the same way: give it one more quarter. Fourteen months in, the company's paid its old vendor's contract in full because nobody could say the pilot was done, and paid to keep the pilot's own engineering running the entire time too. Seven hundred fifty thousand dollars spent, and there's still no decision either way."
Why this works
The compressed version of the story below. Real numbers, a concrete cost, not a hypothetical one.
6
Say what you'd measure to catch it drifting
Say it like this
"Once a quarter, I'd ask the team running the pilot one question, cold: what's the number that ends this, and what's the date? If they can't say it back in one sentence, the pilot's already drifted past having real criteria, whether or not anyone upstream has noticed yet."
Why this works
Turns detection into a repeatable habit, not a one-time audit after the money's already gone.
7
Say what you'd leave alone
Say it like this
"A two-week test one engineer runs on their own laptop, no budget line, nobody's job waiting on the answer, doesn't need this. Write the two numbers down for anything with real spend attached or a real person stuck waiting. Skip the ceremony everywhere else."
Why this works
Shows judgment instead of forcing the same gate onto everything with an "AI" label on it.
8
Land the answer in one breath
Say it like this
"So: a pilot doesn't need a deadline for its own sake, it needs a verdict written down before anyone's attached to the answer. Skip that, and 'still promising' can run forever, because the only people watching it closely are the ones who benefit from it never quite finishing."
Why this works
Restates the decision in one breath, the line an interviewer remembers on the way out.
Let's learn
What does a pilot actually have to prove before anyone's allowed to keep it running?
Say a manufacturer builds a tool, call it LedgerMatch, that reads a vendor's invoice and checks it against the purchase order and the delivery receipt. If all three agree, the invoice pays itself. If they don't, it drops into a queue for a person to sort out.
Knowledge spark: what's a three-way match
Before a company pays a bill, someone checks that three documents agree: the purchase order (what was ordered), the delivery receipt (what showed up), and the invoice (what the vendor wants paid). If all three line up, it's clean. If one doesn't, a person has to work out why before the money moves.
Before LedgerMatch, Corvin Fasteners paid an outside firm, TriMatch Solutions, three hundred forty thousand dollars a year to do that matching by hand for its packaging vendors, forty of them. The pilot was scoped to replace that: one quarter, three months, the same forty vendors, running quietly beside TriMatch's own team so nobody was betting real money on it yet.
Now: month fourteen. The match rate has climbed the whole time. Seventy one percent clean matches in month one. Eighty one by month six. Eighty nine now. Nobody can say it's failing.
LedgerMatch's clean-match rate, month 1 to month 14
Climbing the whole time. The dashed line is the kill bar nobody ever wrote down.
At month six the real number was eighty one percent, four points under where a written kill line would have sat. Nobody had drawn that line, so nothing forced a call, and the climb just kept looking promising, month after month.
Here is the turn. Nobody can say the pilot is done, either, because nobody ever wrote down what "done" meant. So every quarterly review ends the same sentence: it's promising, give it a bit more time. And that sentence can be completely true, every single quarter, forever, because a slow honest climb from seventy one to eighty nine looks exactly the same at month three as it does at month fourteen, from inside the room.
The pilot was never failing. It just never had a way to stop looking promising and start being decided.
At its worst, that costs real money going nowhere. TriMatch's contract has already renewed once since the pilot began, in full, three hundred forty thousand dollars, because there was nothing written down that said stop. Add the fourteen months of engineering it's taken to keep LedgerMatch's own pilot running, the data connectors, the exception handling, the retraining, and Corvin has spent about seven hundred fifty thousand dollars finding out nothing either way.
What fourteen months with no verdict cost, against six with one
A contract renewal plus fourteen months of pilot engineering, versus six months of engineering and a renegotiated contract.
No decision made, month 14
$750,000
Kill line forces a call, month 6
$175,000
The pilot's engineering cost is the same either way, about twenty nine thousand dollars a month. The gap between the two bars is what one written contract renewal, made without a real decision behind it, costs on its own.
The decision I would take back
Nobody in the kickoff meeting wrote a pass number or a kill number on the board. Somebody said "low to mid nineties" felt about right, and it stayed a feeling instead of becoming a number, because guessing wrong out loud felt like the bigger risk. That made sense in the room. It just wasn't safe once the room stopped being the thing that had to pay for it.
What I would leave alone. Corvin also ran a two-week test of a tool that auto-tags expense receipts by category, one person, no budget request, nobody waiting on the answer. That one doesn't need a written pass number and kill number. Nobody's stuck if it just quietly stops.
The lesson. A pilot that never ends usually isn't a broken pilot. It's a pilot nobody made themselves brave enough to be wrong about, out loud, on a specific date, in front of the people paying for it. Guessing the number is the actual job. Refusing to guess doesn't protect you from being wrong. It just spreads being wrong out over a year instead of a quarter.
Now here is the same thing as a story
The short version sits above. Read this one for the fourteen months nobody put a number on the wall.
Conrad Iversen can build Corvin Fasteners' whole annual budget from memory by the third week of August, line by line, eleven years running.
He knows exactly what TriMatch's contract costs, exactly when it renews, exactly which vendors it covers. So when the LedgerMatch pilot kicked off, replacing that contract with something built in-house, Conrad was glad to see it. The first quarterly review even felt like his kind of meeting: a real number, seventy one percent, next to a number everyone in the room had agreed to call the target, low to mid nineties, and a plan to check back in three months.
One of them holds the finish line. The other one only holds a calendar reminder.
The habit faded in three passes he barely noticed at the time. Quarter one's review had a number and a target side by side on the slide. Quarter three's review had a number and the word "steady." Quarter five's review, month fifteen counting from the original kickoff, just said LedgerMatch: ongoing, no number at all, because by then everyone already knew it was climbing and nobody needed reminding.
There wasn't a Tuesday where it broke. That's the part Conrad keeps turning over. It built up slowly, the way weather does, until one afternoon in August, rebuilding next year's budget, he pasted last year's line items forward the way he always does, and there they were: TriMatch's contract, renewed in full a second time, and LedgerMatch's engineering line, funded straight through again, both sitting in the same spreadsheet, both fully expected, for a project that still didn't have an ending.
The step that was supposed to interrupt the renewal, and never got built
He wasn't hiding from it. Conrad emailed Zeynep Serrano, who runs the pilot, and asked, the way he'd asked every quarter for a year: where are we, really? The answer was the same honest answer it had been every time. Getting close. Give us one more quarter. It wasn't a dodge. Zeynep's team could see the number climbing too, and stopping something that's working felt like exactly the wrong call to make on a hunch.
We didn't build a bad tool. We built a tool that was always about to be good enough, and it never once had to prove that by a date.
All Conrad could do was sign both lines again. Third full year of double-funding, coming up, and still nobody on either side of that email had the standing to say: this ends now, one way or the other.
Back at that first kickoff meeting, eighteen months earlier, the room had thirty minutes and a whiteboard. Somebody floated ninety two percent as roughly where "done" should sit. Nobody wrote it down. Writing a number down felt like promising it, and nobody in that room wanted to be the one who'd guessed wrong in front of finance. Leaving it as a feeling felt like the safe choice. It made sense, for that half hour.
Run the same kickoff again, with one change. Before anyone leaves the room, two numbers go on the whiteboard and stay there: ninety two percent, held six weeks, and we go live. Under eighty five percent by month six, and we kill it, no vote required. The pilot runs exactly the same way, climbing exactly as honestly, seventy one percent, then eighty one by month six. Eighty one is under eighty five. The pilot ends that Tuesday, on schedule, by a rule written before anyone was attached to the answer. TriMatch's contract gets renegotiated down instead of renewed at full price. Total spend lands around a hundred seventy five thousand dollars, not seven hundred fifty, and everyone in that room knows exactly what's true by month six instead of guessing at it for another eight.
One design let the pilot decide for itself when it was finished. The other put a number on the wall before anyone had a reason to defend it.
What I'd tell myself, sitting in that first meeting: I thought refusing to guess a number was being careful. It just meant we'd be wrong for fourteen months instead of six.
GUARD, for a pilot with no expiry date
This reads like a scheduling question. The real test is who gets to decide when "promising" has to become a verdict.
G, groups. Zeynep's pilot team, who benefit from the pilot staying open and promising, versus Conrad and finance leadership, who are stuck funding both the old contract and the new tool with no way to make either decision themselves.
U, unequal. The cost lands hardest on whoever's paying for both systems at once with no verdict in sight, here Corvin's finance budget, and on the three people on Zeynep's own team who can't be reassigned to the next real project because their current one has no ending.
A, ability to contest. Conrad has no lever. He can ask for updates every quarter, forever, but he can't set the finish line, because the pilot's own definition of "ready" belongs entirely to the team running it.
R, reduce. Write the pass number and the kill number before the pilot starts, owned by the business side, not the pilot team, and attach the kill date to something already on the calendar, the BPO contract's renewal date, so there's no quiet way to slide past it.
D, detect. Once a quarter, ask the pilot team to state the pass number, the kill number, and the date, in one sentence. If they can't, or the numbers have moved since last quarter without anyone signing off on the move, the pilot's already directionless.
Where this answer would fail
If the fix here is "add another check-in meeting" or "ask the pilot team to report more often," it doesn't count. That's a bigger dial on the same broken design, more updates from the same team that benefits from staying open-ended. The only version that closes the gap is two numbers, owned outside the pilot team, attached to a real date, not a review cadence.
And if you want to be sure it really works, try it somewhere else
A regional grant-making foundation pilots a tool that reads a submitted grant application and flags whether it clears the basic eligibility rules for a given fund, before a program officer gives it a full read. Four months, one fund, community arts grants under ten thousand dollars, a little over a hundred applications a cycle. It's now run six cycles, close to two years, still called a pilot, and it's quietly become the thing three program officers lean on to decide which applications get read closely at all.
G, groups. Sten Kallio's small review team, who keep tuning the eligibility rules cycle over cycle, versus every future program officer and every applicant to a fund the pilot has never actually touched. U, unequal. Barely matters on the arts fund the tool was built and tuned on. It lands hardest on funds the pilot never ran against, like the youth-program fund, where the eligibility rules are genuinely different and nobody's checked whether the tool's rules apply there at all. A, ability to contest. An applicant to the youth-program fund has no way to know their application got a quick "likely ineligible" flag from a tool tuned on a completely different fund's rules, and no formal way to ask for a human first read instead. R, reduce. A named list of which funds the tool's rules have actually been checked against real applications, not just installed for. No fund gets the tool's flag treated as a first cut until it's on that list, with a fixed date by which the list either gets finished or the tool gets turned off for the funds still missing from it. D, detect. Track every "likely ineligible" flag by which fund it came from, and check monthly whether any fund not yet on the checked list is quietly having its applications screened anyway.
Swap the trigger and it still runs
Speed: leadership wants a board update before next quarter's meeting, so the two numbers get written down in a rush, copied from wherever the pilot already happens to be, instead of chosen honestly before anyone knew the result.
Cost: writing a real kill number means someone has to admit, in public, that the money spent so far might turn out to be wasted, so it keeps getting softened into "we'll know it when we see it."
The model gets better: the match rate keeps climbing, quarter over quarter, which makes it easier to keep saying "just a bit more time," not harder, because every review has a genuinely nicer number to point to.
Where people run it wrong
Treating a steadily improving number as proof a decision isn't needed yet, when the real question was never whether it's improving, it's whether it's improving fast enough by when it matters.
Writing the two numbers down after the fact, once the actual result is already known, which isn't a criterion, it's a caption.
Letting the team running the pilot own both numbers, so the same people who benefit from "still promising" get to decide what "still promising" means.
How to use it live
Ask "what's the number that ends this, and what's the date" before you ask anything about how the pilot's going. If nobody can answer both halves in one sentence, you've found the gap in about five seconds.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
What framework fits a question about designing a pilot's exit criteria, and why?
Tap to flip
ANSWER
GUARD, for risk and fairness. The real question isn't whether the model's good enough, it's who's exposed while nobody's forced to decide that.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Conrad Iversen, VP of finance at Corvin Fasteners, stuck funding both an outside vendor's contract and an AI pilot meant to replace it, with no way to force either decision.
3 · THE HABIT
What did Corvin's quarterly pilot reviews stop asking for, over time?
Tap to flip
ANSWER
They stopped naming a number and a target side by side. Quarter one's slide had both. By month fifteen, the slide just said "LedgerMatch: ongoing," no number at all.
4 · THE GAP
What's the gap this answer turns on?
Tap to flip
ANSWER
Nobody wrote a pass number or a kill number before the pilot started. A slow, honest, improving match rate looks exactly the same at month three and month fourteen, so nothing ever forced a real decision.
5 · THE OLD DECISION
What old decision does this answer take back, and why did it make sense at the time?
Tap to flip
ANSWER
The kickoff meeting where nobody put a pass number or a kill number on the board. It made sense because guessing a number out loud, and maybe being wrong in front of finance, felt riskier than leaving it open.
6 · THE NUMBER
Fill in: with an 85% kill line written down at the start, the actual month-six match rate of ______ percent would have ended the pilot eight months earlier.
Tap to flip
ANSWER
81 percent. Four points under the line. The pilot wasn't failing, it just never cleared a bar that was never written down, so nothing forced the call.
7 · THE REPLAY
Same pilot, two numbers written down this time. What changes?
Tap to flip
ANSWER
At month six, 81% is under the 85% kill line, so the pilot ends that Tuesday. TriMatch's contract gets renegotiated instead of renewed at full price. Total spend lands near $175,000 instead of climbing past $750,000.
8 · TRANSFER
Section four runs GUARD again on a different product. Which one, and what does the "unequal" gap become?
Tap to flip
ANSWER
A foundation's AI eligibility pre-screener for grant applications. The gap: it barely matters on the arts fund it was tuned on, and lands hardest on funds the pilot never ran against, like the youth-program fund, where nobody's checked whether its rules even apply.
Check yourself Score: 0 / 0
Short answer
1. What old decision does this answer take back, and why did it make sense when Corvin's team first scoped the LedgerMatch pilot?
Show hint
Think about what it would have felt like to say a specific number out loud in that first meeting, before anyone knew if it was right.
Show answer
Model answer: "The kickoff meeting where nobody wrote a pass number or a kill number on the board. It made sense at the time because guessing a number, and maybe being wrong about it in front of finance, felt like a real risk, while leaving it undefined felt safe, at least for that first meeting."
Multiple choice
2. The match rate climbed the whole time, seventy one percent to eighty nine percent over fourteen months. What does that climb actually prove?
A. LedgerMatch is now accurate enough to end the TriMatch contract.
B. Without a written pass or kill line, a steady improvement only proves it's happening, not that it's fast enough to matter.
C. Fourteen months is definitely too long for any pilot to run.
D. The climb proves TriMatch's manual process was worse than everyone assumed.
Show hint
Ask what a number with no bar attached can actually tell you, versus what it feels like it's telling you.
Show answer
B. A, C, and D all treat the climb as proof of something it was never built to prove. The real problem was never the direction of the number. It was that nobody had written down how fast it needed to move, or by when.
True or false
3. True or false: the two-week expense-receipt tagging test one engineer ran solo, with no budget request and nobody waiting on the outcome, needed the same written pass and kill numbers as the LedgerMatch pilot.
True
False
Show hint
Ask whether anyone is actually stuck waiting on that one, or any real money riding on it.
Show answer
False. Nobody's stuck waiting on that one, and no real money is riding on it. The heavy version of this gate is for pilots with real spend attached or a real person waiting on the verdict, not every small experiment in the building.
Fill in the blank
4. If Corvin had written down a kill line of eighty five percent by month six, the actual match rate at month six, ______ percent, would have ended the pilot right there instead of letting it run to month fourteen.
Show hint
It's the number that shows the pilot wasn't failing, it just hadn't cleared a bar nobody had written down.
Show answer
81 percent. Four points under the 85 percent line. That gap is the whole point: the pilot wasn't bad, it just never had a rule that could catch it not being good enough yet.
Short answer, apply it yourself
5. Think of a trial, a beta, or a "let's just pilot it for now" arrangement you've seen at work. What would the pass number and the kill number have been, if anyone had written them down on day one?
Show hint
Look for the one that's still "in pilot" well past when it was supposed to report back to anyone.
Show answer
Model answer: "A CRM add-on my team trialed for two quarters to auto-draft follow-up emails to leads. Nobody ever set a bar. Looking back, the pass number should've been something like eighty percent of drafts sent with no edits by week eight, and the kill number should've been under fifty percent by week four, revert to the old templates. Instead it just quietly kept running for five months until someone new joined and asked why we still had two systems doing the same job."
True or false
6. True or false: adding one more monthly check-in meeting where the pilot team gives leadership an update counts as fixing the missing exit criteria.
True
False
Show hint
Ask whether the fix changes who owns the number, or just how often the same team reports on itself.
Show answer
False. That's a bigger dial on the same broken design, more updates from the same team that benefits from staying open-ended. The fix is a written pass number and kill number owned by someone outside the pilot team, not a meeting where the pilot reports on itself.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.