ConceptIntermediateShipping & Model Lifecycle / Pilot design and POC-to-production / #2
What success criteria should be agreed before a pilot begins?
The direct answer
Before a pilot starts, write down two hard numbers: the smallest share of the real outcome the model has to catch, and the most false alarms the team can absorb without wasting its week. Name who signs off on go or no-go, and get both sides to sign the page. That's what turns the last day of the pilot into an honest decision instead of a fight over which number counts. It can tell you the feature works. It can't tell you the customer will actually buy.
Do this, in order
Write two hard numbers before kickoff: a catch-rate floor and a false-alarm ceiling.Why: without a number, a mixed result gets read as a win by whoever wants to keep the deal and a loss by whoever wants to kill it.
Name who signs off on go or no-go, and by what date.Why: two numbers on a page don't decide anything if nobody owns pulling the trigger.
Size the catch-rate target to what the team can actually act on.Why: a target the team has no capacity to work is a number that fails by design, not by the model's fault.
Write the false-alarm ceiling in hours or dollars, not just a percent.Why: a bare percent hides the real cost until someone counts the wasted calls.
Separate "the feature hit its numbers" from "the customer signs the contract."Why: budget, trust, and timing decide the purchase; the pilot's numbers only decide whether the feature works.
Skip pre-agreeing details nobody's actually going to argue about.Why: don't spend the kickoff meeting on email wording when the two numbers that matter haven't been pinned yet.
How to answer this, stage by stage
Eight moves. Name what the numbers can and can't settle before the story, or the answer sounds like project management instead of a decision.
1
Scope it to one pilot and one open question
Say it like this
"Let's make this concrete. Priyanka Nethercott runs product for Anchorlight, a tool that scores every subscriber's risk of canceling. Grovemill Crate, a spice and pantry subscription box with about thirty-eight thousand subscribers, is running an eight-week pilot of it. The question in front of Priyanka is what gets written down before that pilot even starts."
Why this works
A question about success criteria stays abstract until it's tied to one real pilot with a real end date coming.
2
Say the plan out loud
Say it like this
"I'll walk this through LEAD. L is the real outcome the criteria protect. E is the early signal, the actual numbers, written down before day one. A is how it gets gamed when nobody wrote them down. D is what a clean pass still can't tell you."
Why this works
Naming the four letters up front tells the interviewer you're about to make a call, not describe a project plan.
3
Say what the question is actually checking
Say it like this
"This sounds like a question about what to measure. It's really asking who gets to decide, once the results are in, what counts as a win, and whether that was settled before anyone had a reason to prefer one answer over another."
Why this works
This moves the answer from "measure the right things" to naming exactly what pre-agreement replaces: a fight that starts the moment the numbers land.
4
Give the decision straight
Say it like this
"Here's the answer. Before the pilot starts, write down two numbers: the smallest share of real cancellations the model has to catch in advance, and the most false alarms the team can absorb in a week without it costing them. Name who signs off on go or no-go. Get both sides to sign the page. That decides the pilot honestly. It does not decide whether Grovemill actually buys."
Why this works
This names exactly what the two numbers settle and what they leave open, instead of a vague "agree on metrics."
5
Prove it with the number that split the room
Say it like this
"Here's why that matters. Priyanka's pilot ran eight weeks with no number written down for what would count as a pass. It ended with the model catching fifty-one percent of subscribers who actually canceled, before they canceled, and sixty-nine percent of everyone it flagged turning out to be a false alarm. Anchorlight's side called fifty-one percent a clear win. Grovemill's retention lead called sixty-nine percent a clear reason to kill it. Both were reading the exact same eight weeks."
Why this works
One real result read two honest, opposite ways does more work than a paragraph about the importance of clear metrics.
6
Name the gaming path
Say it like this
"Nobody was acting in bad faith. Anchorlight's exec walked into the readout and said, 'we caught half the people who would have quit, before they quit, that's the whole pitch.' Grovemill's retention lead walked in and said, 'my team burned three days a week calling people who were never leaving.' Neither one was wrong. That's how a pilot gets gamed without anyone cheating: once the results exist, everyone reaches for whichever number was already true, and calls that number the point."
Why this works
Naming the exact mechanism, picking the flattering number after the fact, is what separates this from a generic "define your metrics" answer.
7
Name the hard limit
Say it like this
"And here's what even a clean number still couldn't tell Priyanka. Say she'd pre-agreed a forty percent catch-rate floor and the pilot cleared it easily. That proves the feature works. It doesn't tell you if Grovemill's CFO signs the annual contract, or whether this happens to be a quarter where nobody at Grovemill wants to add a new vendor, regardless of what the pilot showed. That's a separate decision, made by different people, on a different clock."
Why this works
Naming this boundary out loud is what keeps the answer from overclaiming what pre-agreed criteria can actually do.
8
Close on the artifact, one line
Say it like this
"So if I only get to change one thing about how Priyanka ran this: before day one, a one-page agreement, two numbers, one owner, one date, signed by both sides. Not because the pilot needs paperwork. Because the fight I just described happens the moment the numbers land, and by then it's too late to decide what they mean."
Why this works
Closing on the artifact instead of the abstract principle gives the interviewer something they could actually go build.
Let's learn
Every month, Grovemill Crate's retention team called the subscribers who had already paused or skipped that week's box. That was the whole method. Wait for someone to act like they were leaving, then try to talk them out of it.
Knowledge spark: what's churn?
The customers who cancel or stop renewing. If Grovemill loses nine out of every hundred subscribers a month, that's a nine percent churn rate. A tool that predicts churn is trying to name who's about to become one of the nine, before they do.
Over a normal eight-week stretch, that reactive method saved about nine subscribers. Not nine percent. Nine people, total, out of thirty-eight thousand, because by the time someone skips a box, they've usually already decided.
Anchorlight's pilot changed the shape of the problem. Every week, it scored all thirty-eight thousand subscribers for risk of canceling in the next thirty days, before anyone skipped anything. It flagged about fifteen hundred people a week as high risk. Grovemill's retention team, three people, could realistically call or email about five hundred. So the team worked the top five hundred scores each week and left the rest.
Here's the turn. Fifteen hundred flags against a five-hundred-person capacity was not really the problem either, though it should have been fixed. The real problem showed up eight weeks later, at the readout meeting, when the numbers came in and nobody in the room had agreed, before that morning, what a good number even looked like.
The model was never the problem. Nobody had written down, before day one, what a win would even look like.
At its worst, that kind of gap doesn't just waste a meeting. A pilot that genuinely worked gets shelved because the loudest objection in the room happened to land after the fact, with nothing written down earlier to weigh it against. Or a pilot that genuinely didn't work gets waved through, because the one flattering number got said first and nobody had a pre-agreed number to hold it against.
Weekly catch rate during Grovemill's eight-week pilot, against the two numbers nobody wrote down
By week four the model had already cleared the bar Anchorlight's side would privately have called a pass. It never got near the bar Grovemill's retention lead had in mind. Both bars were real. Neither was ever written down.
The false-alarm side of the same pilot told the opposite story. Sixty-nine out of every hundred people the model flagged would have stayed even if nobody called them. The retention team spent three of their five days a week on people who were never leaving.
Subscribers actually saved from canceling, before the pilot and during it
9
Eight weeks before the pilot, reactive outreach only, after someone had already paused
34
Eight weeks of the pilot, flagged before they canceled and offered a save
Thirty-four real saves is close to fourteen thousand dollars a year in subscriptions that would otherwise have walked. That's also true. It's just not the number the retention lead brought up first.
Same eight weeks. Same report. Two people circled two different numbers and both were right.
The choice I'd take back
At kickoff, Priyanka told both teams, "let's run it for eight weeks and figure out together what counts as good." It felt collaborative, not careless. I'd take it back and replace it with two numbers, a named decision-owner, and a decision date, written on one page both sides sign before day one.
What I'd leave alone. Nobody needed to pre-agree the exact wording of the retention offer email, or which color the dashboard uses for a high-risk score, or how many people get called first each morning. Those can flex all pilot long without changing whether the pilot's own question gets answered honestly.
The lesson. A pilot without a number agreed in advance doesn't fail to produce an answer. It produces two answers, both true, and lets whoever argues harder pick which one counts.
Now here is the same thing as a story
Use this version when you've got a few minutes. The short version is above. This is for when the decision actually needs to survive a room full of people who disagree.
Every Friday afternoon, before the call with Grovemill, Priyanka pulled up the week's model scores and read them the way other people read a weather report. Two years running product for Anchorlight had taught her to check the shape of the numbers before she checked whether they were good.
The first month of the pilot was quiet in the best way. The scores climbed steadily, week over week, nothing dramatic, no red flags in the pipeline. Grovemill's retention team worked their five hundred calls, reported back what worked, and the Friday call ran fifteen minutes. Priyanka liked those calls. Everyone seemed to be rowing the same direction.
By week six, the numbers were genuinely good. Forty-four percent of the people who actually canceled had been flagged in advance. Priyanka started drafting the readout deck early, the way you do when you're fairly sure how a thing is going to land.
Then week eight arrived, and with it the two final numbers. Fifty-one percent caught. Sixty-nine percent false alarm rate. Priyanka put both on the same slide, because both were true, and walked into the readout expecting a conversation about what to do next.
She got an argument instead.
Anchorlight's VP of sales spoke first. "We caught half the people who would have quit, before they quit. Nobody in retail loyalty gets that number. This is a clear yes." Grovemill's retention lead spoke right after him, and she wasn't wrong either. "My team spent three days a week for two months calling people who were never leaving. That's not a rounding error, that's most of their week." Neither one raised their voice. Neither one was lying. They were reading the same eight weeks and landing in different countries.
Nobody was gaming the pilot. The pilot had never been given a number to be gamed against, so both sides picked the true one that helped their case.
Priyanka sat through forty minutes of a meeting that was supposed to take ten. Nobody moved. The VP had the win-number memorized. The retention lead had the wasted-hours number memorized. Both of them had shown up ready to argue a case, because that's what an unwritten pass bar turns a readout into: a trial, with everybody as their own witness.
Two months earlier, at kickoff, someone from Grovemill had actually asked, "so what number are we looking for here?" Priyanka remembers answering, "let's run it for eight weeks and see, and we'll figure out together what counts as good," and the room nodded, because it sounded reasonable, even generous. Nobody circled back to pin it down. It just quietly became the plan.
I would take that back. I'd have said, in that same kickoff meeting, "before we start the clock, I want two numbers on paper: the lowest catch rate that makes this worth Grovemill's time, and the most false alarms your team can absorb without it costing them. Let's agree those today, and agree who says yes or no at the end, before either number exists to argue about."
Here's the replay. Same eight weeks, same fifty-one percent, same sixty-nine percent, but this time the kickoff page says: catch at least forty percent, keep false alarms under seventy-five percent of the flagged list, Grovemill's VP of ops signs off within five business days of the final number. The readout meeting still happens. It runs twelve minutes instead of forty, because the room isn't deciding what the numbers mean anymore, only reading them against a bar that was already agreed two months earlier.
One version hands the room two true facts and lets them fight. The other hands the room a yes or a no that was already decided, and just needs reading out loud.
And the thing I'd tell myself, back in that kickoff meeting: "we'll figure it out together once we see it" isn't collaboration. It's postponing the hardest conversation to the one moment everyone has the least reason to agree.
LEAD, for a pilot that has no number yet
This sounds like a question about what to measure. Underneath, it's still asking whether a team can act on a decision made in advance, instead of a fight made after the fact. That's LEAD, run on a pilot's own ending instead of a dashboard.
L, link. The real outcome pre-agreed criteria protect. Not the model's score in isolation. Whether the go or no-go decision, at the pilot's end, gets made honestly instead of becoming a retroactive negotiation over what counted as success. → Here, that's whether Grovemill's yes or no gets decided by a number agreed in March, not by who spoke first in May.
E, early signal. Specific, numeric pass thresholds written down before the pilot starts, not vague satisfaction. → A forty percent catch-rate floor and a false-alarm ceiling, agreed and signed before week one, not "let's see how it feels."
A, abuse. How success criteria get gamed without pre-agreement. Redefining "success" after seeing the results, to match whatever happened. → The VP led with the catch rate because it flattered the deal. The retention lead led with the false-alarm rate because it flattered her team's complaint. Both were true. Neither was pre-agreed.
D, decision. What pre-agreed criteria genuinely cannot settle. Whether the customer will actually buy at the end, only whether the feature performed as specified. → Even a clean forty-percent pass tells Grovemill the model works. It says nothing about their budget cycle, their trust in a new vendor, or whether this quarter was ever going to say yes to anything new.
The check that proves the agreement is real
Ask both sides, separately, before the pilot starts: "if the model lands at exactly this number, do we have a deal?" If either side hesitates or adds a new condition on the spot, the criteria aren't actually agreed yet, they're just written down.
And if you want to be sure it really works, try it somewhere else
Northfell Freight runs a hundred and twenty trucks out of three regional yards. Bertholt Aldercroft leads product for a predictive-maintenance tool that flags which trucks are likely to break down before they actually do, and Northfell agreed to a ninety-day pilot across the fleet.
L. Whether Northfell's fleet manager signs the annual license because the tool caught real breakdowns in advance, on agreed terms, not because of whichever anecdote from the ninety days got repeated loudest.
E. A pre-agreed number: catch at least seventy percent of major breakdowns with at least five days' lead time, so the shop can actually schedule the repair instead of scrambling.
A. Without that number, one bad miss outweighs eleven good catches in the room's memory. A truck the tool flagged, that the shop then cleared, broke down two days later on a highway outside Millbrook. Sales pointed at seventy-three percent caught. The fleet manager pointed at the one truck that stranded a driver.
D. Hitting the seventy-percent floor proves the tool can be trusted to flag real risk. It can't say whether Northfell signs the annual contract, that also turns on their budget cycle and whether freight volumes that quarter make anyone nervous about new spending at all.
Knowledge spark: what's lead time, here?
How many days of warning the shop actually gets before something breaks. A flag with no lead time is just a very fast confirmation of bad news. A flag with five days lets a mechanic swap the part on a truck's normal maintenance day instead of on the side of a highway.
Bertholt pulled the pilot's own logs the week after the highway incident: eleven of fifteen real breakdowns caught, with an average lead time of six days. He printed both numbers next to the seventy-percent floor Northfell had agreed to at kickoff, and the follow-up meeting, expected to run an hour, lasted nine minutes.
Waiting for the breakdown tells you last. The pre-agreed warning sign was already ringing days earlier.
Swap the trigger and it still runs
The model scores faster. Doesn't help by itself. A faster number nobody agreed on is still not a decision, it's just an earlier fight.
The pilot budget shrinks. Doesn't help either. Fewer weeks just means less time to write two numbers down, not less need for them.
The model genuinely gets more accurate. Still needs the same two numbers, pre-agreed. That's the one case where a pre-agreed bar should make saying yes easier, not harder, and only a written bar can do that.
Where people run it wrong
Writing a vague target like "high accuracy" instead of an actual catch-rate floor and a false-alarm ceiling.
Agreeing on the numbers but never naming who actually signs off on go or no-go.
Setting a catch-rate target the team has no realistic capacity to act on, so the pilot fails by design, whatever the model does.
How to use it live
Say the split out loud, first. "Before we talk about metrics, I want to separate two things: what number would count as a pass, and who decides once we have it. Let's pin both before the clock starts." That's not stalling. It names the two things a pilot actually needs settled, and it buys you a beat to structure the rest of the answer around a decision instead of a wish list.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits this question, and what does each letter stand for here?
Tap to flip
ANSWER
LEAD. L is the real outcome, an honest go or no-go decision, not a retroactive fight. E is the early signal, numeric pass thresholds written down before day one. A is how it gets gamed, redefining success after seeing the results. D is what it can't settle, whether the customer actually buys.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Priyanka Nethercott, who runs product for Anchorlight, a churn-prediction tool, and ran an eight-week pilot with Grovemill Crate, a subscription-box company, with no pass number agreed in advance.
3 · THE HABIT
What kept the readout meeting stuck for forty minutes?
Tap to flip
ANSWER
Both sides had a true number ready and no pre-agreed bar to check it against, so the meeting became an argument over which true number counted, instead of a read-out of a decision already made.
4 · THE EARLY SIGNAL
What's the E step here, in one line?
Tap to flip
ANSWER
Specific numeric pass thresholds, a catch-rate floor and a false-alarm ceiling, written down and signed by both sides before the pilot's first day, not decided once the results exist.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Telling both teams at kickoff, "let's run it for eight weeks and figure out together what counts as good." Replace it with two numbers, a named decision-owner, and a decision date, signed before day one.
6 · THE NUMBER
The model caught ______ percent of real cancellations in advance. Of everyone it flagged, ______ percent would have stayed anyway.
Tap to flip
ANSWER
51 percent caught. 69 percent false alarms. Both numbers are true. Nobody had agreed beforehand which one decided the pilot.
7 · THE REPLAY
Same eight weeks, same final numbers, criteria agreed on day one instead. What changes?
Tap to flip
ANSWER
The forty-minute readout becomes a twelve-minute one. The room isn't deciding what the numbers mean anymore, only reading them against a bar agreed two months earlier and signed by Grovemill's VP of ops.
8 · THE TRANSFER
Section four runs this same question again for a different product. Which one, and what's its early signal?
Tap to flip
ANSWER
Northfell Freight's predictive-maintenance pilot. Its early signal is a pre-agreed floor: catch at least 70 percent of major breakdowns with at least five days' lead time, agreed before the ninety-day pilot began.
Check yourself Score: 0 / 0
Fill in the blank
1. Priyanka's model caught ______ percent of subscribers who actually went on to cancel, before they canceled. Of everyone it flagged, ______ percent would have stayed anyway even without a call.
Show hint
These are the two numbers on the same slide that split the readout room.
Show answer
51 percent. 69 percent. One number reads as a clear win, the other as a clear cost. Both are true about the same eight weeks, which is exactly why a pre-agreed bar was needed to settle which one decided the pilot.
Multiple choice
2. Which of these is the decision Priyanka says she'd take back?
A. Adding more retention staff so the team could work all fifteen hundred weekly flags instead of five hundred.
B. Tightening the model's risk threshold in production so it flags fewer people.
C. Saying at kickoff, "let's run it for eight weeks and figure out together what counts as good," instead of pinning two numbers and an owner in writing first.
D. Watching the weekly scores more closely on the Friday calls.
Show hint
Look for a decision taken back, not a dial turned up or a habit of watching more closely.
Show answer
C. The other three are dials, more staff, a tighter setting, closer watching. None of them would have stopped the forty-minute argument, because none of them put a number on paper before the results existed.
True or false
3. True or false: once the pilot ended with a 51 percent catch rate, that number alone was enough to know Grovemill would sign the annual contract.
True
False
Show hint
Separate "the feature performed as specified" from "the customer buys."
Show answer
False. A catch rate, even a strong one, only proves the feature works. Whether Grovemill actually signs still depends on their budget cycle, their trust in a new vendor, and their timing, none of which a pilot's pass number can settle. That's the D step.
Multiple choice
4. Which of these genuinely didn't need to be pre-agreed before Grovemill's pilot began?
A. The catch-rate floor the model had to clear.
B. The most false alarms the retention team could absorb in a week.
C. Who signs off on go or no-go, and by what date.
D. The exact wording of the retention offer email sent to flagged subscribers.
Show hint
Look for the one detail that could flex all pilot long without changing whether the pilot's own question got answered honestly.
Show answer
D. Email wording, dashboard colors, and call order are the kind of thing you leave alone. The catch-rate floor, the false-alarm ceiling, and the decision owner are the three things that actually decide whether the pilot ends in a fight or a decision.
Short answer, apply it yourself
5. Think of a pilot, trial, or test period you've seen end in an argument about whether it "worked." What's one number that, if it had been written down before it started, would have settled the argument?
Show hint
Look for the number each side reached for after the fact, and ask what it would have taken to fix that number in place beforehand.
Show answer
Model answer: "A team trialed a new scheduling tool for six weeks. It ended with half the room saying it saved time and half saying it created double-bookings. Writing down beforehand, 'no more than two double-bookings a week, or we stop,' would have ended that argument in the first meeting instead of the last one."
Fill in the blank
6. If Grovemill and Anchorlight had written down a catch-rate floor of 40 percent before day one, the pilot would have cleared that bar by ______ points.
Show hint
The pilot's actual catch rate was 51 percent.
Show answer
11 points. 51 minus 40 is 11. A written 40 percent floor would have made the win undeniable on its own terms, whatever the retention lead thought about the false-alarm rate, because the two numbers would have been agreed as separate questions instead of competing for the same verdict.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.