CaseAdvancedShipping & Model Lifecycle / Pilot design and POC-to-production / #1

Design a four-week pilot for an AI feature with one enterprise customer.

The direct answer
Build the four weeks as four gates, not four scheduled milestones. Start narrow, on backtested numbers only, and let scope widen each week only when the model beats the customer's own accuracy and the review load still fits inside the hours their team actually has free. Close week four with a go or no-go built on the pilot's own numbers, not a demo timed to a date on a deck.
Do this, in order
  1. Gate every week to earned evidence, not a calendar date.Why: a schedule that expands scope by date instead of by proof is what lets an unreviewed pile of forecasts build up behind the scenes.
  2. Size each week's scope to the customer's real review hours, not to what the model can technically forecast.Why: the model can score four hundred parts overnight; two real planners can only check a fraction of that on top of their day job.
  3. Start narrow and backtested before the model makes one live ordering call.Why: week one's backtest is the cheapest place in the whole pilot to catch a bad forecast, before it touches a real order.
  4. Give the go/no-go a range, not one confident number, for how many parts will be live by week four.Why: forty parts if the integration ends up custom, two hundred forty if it's reused, and the room deserves both numbers before week one starts.
  5. Push hard in week zero to reuse the ERP integration wherever the customer's system allows it.Why: of the two big levers, this is the one that swings the final count the most, so it's worth negotiating before the pilot even starts.
  6. Leave the true one-off custom parts out of the pilot entirely.Why: a part that only ever gets ordered once has no history to learn from, so forecasting it isn't a shortcut, it's a guess wearing a spreadsheet.

How to answer this, stage by stage

Nobody is grading whether you land on exactly two hundred forty parts. They're grading whether you can defend what earns each week the right to the next one, whether the hours add up, and whether you close on a number the room can check. Eight moves get you there.

1
Scope it to one real product and one real customer
Say it like this
"Let's ground this in one thing. Say we've built a tool that reads a manufacturer's order history, current stock, and supplier lead times, and drafts a demand forecast for every part they carry. One enterprise customer means one real company, say a precision-parts plant with about four hundred parts they forecast by hand every week, not a hypothetical logo on a slide."
Why this works
Stops the answer from staying abstract before a single week gets planned.
2
Say the structure out loud before naming any numbers
Say it like this
"The number that matters here isn't four weeks. It's four gates. Each week only earns the next one if two things are true: the model's forecast is actually good enough, and the customer's own team has the hours to check it. A calendar with no gates behind it isn't a pilot plan. It's a hope with dates on it."
Why this works
Shows the equation before the arithmetic, so the numbers that follow read as a plan, not a guess.
3
Reframe what a four-week pilot is actually testing
Say it like this
"This isn't really 'can we get an AI feature live with an enterprise customer in four weeks.' It's 'how much of this customer's real work can earn its way onto the model in four weeks, without their team drowning trying to check it.' Those sound like the same question. They aren't, and only one of them survives contact with a real planning team."
Why this works
Separates a real pilot from a demo that just happens to run for four weeks.
4
Break down the week-by-week structure and what has to be true to advance
Say it like this
"Here's the shape. Week one: connect the data and backtest it, no live decisions yet. Week two: run the model alongside the planners' own forecast for one real ordering cycle, still no live decisions. Week three: let the model make live calls on whatever it already proved in week two, while anything new stays in shadow. Week four: everything proven goes live for one real cycle, then we sit down with the actual numbers and decide what's next. Each week only opens if the week before it cleared its bar."
Why this works
This is the B step of BOUND: the equation, the milestones, and their gates, stated before a single number gets attached.
5
Own the real numbers for each week
Say it like this
"For a plant with two hundred forty fast-moving parts, that's: week one, backtest against the last two months of real demand on those two hundred forty. Week two, shadow-run them for one real ordering cycle, capped at eight hours of planner review. Week three, expand the remaining slower movers into shadow while the original two hundred forty go live, capped at ten hours. Week four, all two hundred forty live for one real cycle, then the go, no-go conversation."
Why this works
This is the O step: a real proposed structure with real numbers, not "we'll test it and see how it goes."
6
Give the range, not one number
Say it like this
"If the customer's ordering system already has a connector we've built before, we can realistically get all two hundred forty parts live by week four. If their system is old enough that we're mapping every field by hand, week four might only have about forty parts live, and the rest is still sitting in backtest. Same four weeks, a six-times difference, and we'd know which one we're in by day two."
Why this works
A single number here would claim a confidence the pilot doesn't have yet.
7
Sanity check the plan against the customer's real hours
Say it like this
"Twenty-nine hours of planner review across four weeks sounds fine against a team with roughly twelve hours a week free. But nine of those hours land in week three and eleven in week four. That's twenty hours in two weeks, against a twelve-hour-a-week ceiling. It survives, but only if nothing else goes wrong that week, and on a factory floor something always does."
Why this works
This is the step most pilot plans skip, and it's what stops week three from quietly becoming the week nobody actually reads the forecasts.
8
Name what moves it most, then close on the one line
Say it like this
"If I had to bet on what changes this number most, it's not the four-week length. It's whether we can reuse an existing integration, and whether the customer's team actually has the hours they think they have. So here's the line I'd actually say: I'd run four weeks, but I'd spend day one finding out how much of the integration is already built and how many real hours the planning team has free, because those two answers decide whether we're forecasting forty parts or two hundred forty by week four. Not the calendar."
Why this works
Closes on the literal ask, a claim someone could actually check, not a vibe about "running a successful pilot."
If you remember one thing Let scope widen only when the model's accuracy and the customer's real review hours both clear a bar, never on a date alone. That's what turns four weeks into four earned steps instead of one long guess with a deadline.

Let's learn

The product is a tool that reads a manufacturer's order history, current stock, and supplier lead times, and drafts a demand forecast for every part they carry, instead of a planner rebuilding the same spreadsheet from scratch every Monday.

Before a pilot like this gets planned, a team told to "get it live with the customer in four weeks" usually writes the plan the way any project plan gets written: build in week one and two, test in week three, review in week four. On paper, that covers every part the customer makes from day one. It reads like a serious pilot.

Knowledge spark: what's a backtest? Running the model on demand that already happened, last month's real orders, say, and checking its guess against what actually got ordered. It's a practice run where you already have the answer key in hand.

Here's the turn. Getting the model connected and forecasting is not the hard part. Earning the right to let it make a live ordering decision is. A plan that puts every part live on a fixed date, whether or not the model and the customer's team are actually ready, isn't really a plan. It's a hope with a calendar taped to it.

A pilot with a date on every milestone and no gate behind any of them isn't a plan. It's a hope with a calendar taped to it.

At its worst, that gap shows up in week two, in the size of the pile a customer's own planners have to check by hand. If the schedule, not the evidence, decides how much of the plant runs on the model's numbers by a set date, the planners absorb the difference. And a pile that's too big to check gets skimmed instead of checked, which is exactly the moment a bad forecast can slip through unnoticed.

The decision that mattered Let scope expand only when the model's accuracy and the planners' real review hours both clear a bar, never on a date alone. Variety in hours checked, not a fixed calendar, is what a four-week pilot is actually buying you.

The choice I would take back. On the first pass, I planned the pilot by working backward from the demo date: all two hundred forty of the customer's fast-moving parts, live by week three, so the week-four meeting could open with a finished result instead of a partial one. That felt decisive. It wasn't a real gate. It was a due date wearing a pilot's clothes.

What I would leave alone. The true one-off parts, the ones made once to a single customer's exact spec and never reordered, don't belong in this pilot at all, and they never will. A forecast for a part that will only ever be ordered once isn't a forecast. It's a guess with extra steps. Leave those on the planners' own spreadsheet, forever, and don't spend a single pilot week on them.

The lesson. Four weeks is a fine amount of time for one enterprise pilot. The problem was never the calendar. It was letting the calendar decide how many parts went live, instead of letting the hours the planners actually had left decide it.

Now here is the same thing as a story

The short version is above. Read on if you want to feel why the gates are the whole decision, not a scheduling detail before the real work starts.

Katarzyna Nowak has designed pilots like this for three years, twelve of them by the time this one landed on her desk. She's good at it. She can tell within the first data pull whether a customer's export is going to be clean or a mess, just from how the column headers are named.

Thoresen Metalworks was her twelfth. Thoresen runs a precision-parts plant that forecasts demand for about four hundred repeat-order parts by hand every week, two hundred forty of which move fast enough to carry most of the plant's order value. Piotr Kaczmarek, who runs supply-chain planning there, had spent months pushing his own leadership to approve the pilot, and he wanted the four weeks to look decisive: real scope, real speed, a real answer by the end.

Katarzyna wrote the plan the way she'd written the last few. Week one and two, connect the data and build the model. Week three, run it live across all two hundred forty of Thoresen's fast movers. Week four, review the results with Piotr's boss. It read well in the kickoff deck. Everyone signed off in twenty minutes.

Week one went fine. The connector took three days instead of five, and the backtest looked solid enough on all two hundred forty parts to keep the calendar as written.

Then week two arrived, and with it, the actual size of the pile Piotr's three planners had to check by hand: two hundred forty new forecasts, on top of a supplier delay that was already eating most of their week. Only two of the three planners knew these parts well enough to judge whether a forecast looked right, and between them they had about twelve hours free once the firefighting was accounted for. Two hundred forty forecasts needed close to twelve hours to check properly. By Wednesday, they weren't checking properly. They were opening a forecast, glancing at the total, and moving to the next one, because glancing at two hundred forty lines and doing the job they were actually hired for both had to fit inside the same day.

Hand-sketched comparison of two desks. Left, labeled Scheduled by date, a document icon in red-orange with the caption 240 forecasts due, 12 hours needed, 12 hours free, nothing left. Right, labeled Earned by gate, a document icon in green with the caption 120 forecasts due, 6 hours needed, 12 hours free, six left over.
Same two planners, same twelve hours free. One schedule spends every hour they have. The other leaves half of it.

On Thursday, one of those glances waved through a forecast that had under-ordered a bearing housing with a six-week lead time from its only qualified supplier. Nobody caught it. Not Thursday, not Friday. It sat there until Monday morning, when a planner doing an unrelated stock check noticed the part was down to two days of supply against a forty-two day lead time to replace it.

They caught it with two days to spare, by paying for an emergency shipment that cost more than the part itself was worth. No line stopped. No customer of Thoresen's ever knew. But Piotr knew, because he called Katarzyna that Monday afternoon.

We didn't nearly stop a line. We nearly proved to the person who'd fought to get us in the building that his own team couldn't be trusted to catch our mistakes.

Katarzyna paused the pilot for two days and rebuilt the plan. Same four weeks, same customer, same two hundred forty parts eventually, but scope now had to be earned week by week against two things: the model's own backtested accuracy, and how many hours Piotr's two senior planners actually had free once the firefighting was taken out. Week one, backtest on one hundred twenty parts, not two hundred forty. Week two, shadow only, same one hundred twenty, capped at eight review hours. Week three, expand into shadow for the remaining one hundred twenty while the first batch goes live, capped at ten hours. Week four, all two hundred forty live, one real ordering cycle, then the actual go, no-go conversation, backed by numbers instead of a date on a slide.

In the original kickoff, when Katarzyna proposed "all two hundred forty live by week three," nobody in the room pushed back, because it sounded like exactly the kind of pilot a customer wants to see: fast, comprehensive, decisive. It made sense in that room, in that hour. It stopped making sense the moment the actual review hours had to come out of two real people's real week.

The second version of the pilot reached the same place, two hundred forty parts live, by the end of the same four weeks. The difference was ten fewer hours of planner review in the two crunch weeks, and zero forecasts that went out unread.

The thing I'd tell myself, in that first kickoff meeting: a pilot that looks decisive on a slide and a pilot that's actually been earned are not the same four weeks, and the only way to tell them apart in week zero is to ask whose hours the schedule is actually spending.

BOUND, gated before the kickoff deck goes out

This is a sizing question about how much of a customer's real work can go live in four weeks, and what that number depends on, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.

B, break it down. Four weeks isn't the plan. Four gates are. Week one earns the right to run in shadow: the pipeline has to work without hand-patched data, and the backtest has to beat or match the planners' own historical accuracy on most of the parts tested. Week two earns the right to go live: the shadow forecast has to land inside an agreed error band, and the review time has to fit inside the hours the planners actually have. Week three earns the right to expand: no stockout or overstock event traceable to the model among the live parts, and the review load still has to fit. Week four is the live run and the real decision, not a demo.
O, own the numbers. For Thoresen's two hundred forty fast-moving parts: week one, backtest against eight weeks of real demand on the first one hundred twenty. Week two, shadow-run those same one hundred twenty for one real ordering cycle, capped at eight review hours, coming in around six. Week three, add the remaining one hundred twenty into shadow while the first batch goes live, capped at ten hours, coming in around nine. Week four, all two hundred forty live for one real cycle, review and monitoring around eleven hours.
U, use a range. Before running anything, the honest range for how many parts are live by week four is wide. If Thoresen's ordering system already has a connector built for a past customer, all two hundred forty can realistically be live. If it's old enough that every field gets mapped by hand, week four might only have about forty parts live, with the rest still sitting in backtest. Same four weeks, a six-times difference, depending only on how much of the integration gets reused.
N, nail the sanity check. Twenty-nine hours of planner review across four weeks looks fine against a team with roughly twelve hours a week free, on average. But nine of those hours land in week three and eleven in week four, twenty hours across two weeks against a twenty-four hour ceiling for that pair of weeks. That's four hours of slack across the two busiest weeks of the pilot, which is exactly tight enough to survive and exactly tight enough to explain why the first draft, which needed nearly twelve hours in week two alone, didn't.
D, direction. Two assumptions could change this number a lot, and two barely move it. Whether the ERP integration gets reused instead of built custom swings the parts live at week four by about two hundred. Whether the two senior planners' real slack is six hours a week instead of twelve swings it by about one hundred fifty. Freeing up a fourth planner just for pilot review moves it by about forty. Extending the backtest window moves it by about ten. Integration reuse is the bigger number on paper, but bandwidth is the one that quietly kills a pilot, because you can throw two extra engineers at a connector overnight. You can't hand a planner more hours in their week.

The build-up: planner review hours across the four gated weeks
Week 1, backtest sanity check3 hrs
+ Week 2, shadow run, 120 parts9 hrs
+ Week 3, expand + first live calls18 hrs
+ Week 4, full live, 240 parts29 hrs
Doubling the part count roughly doubles the review hours. That's why the redesigned plan grows the pile one gate at a time instead of putting all of it on the planners in week two at once.
What moves the parts live by week four (swing from a baseline of 240)
ERP integration built custom, not reused−200
Senior planners' real slack is 6 hrs/week, not 12−150
A fourth planner freed up for pilot review+40
Backtest window extended from 8 to 16 weeks+10
Whether the integration gets reused swings the count the most, but the planners' real hours is the assumption nobody put on the kickoff deck, and it's the one that actually caused the near miss.

And if you want to be sure it really works, try it somewhere else

A bank pilots an AI tool that scores every card transaction for fraud risk, with one regional bank as the enterprise customer, over the same four weeks.

B, break it down. Same shape, different work. Week one earns shadow scoring: the model has to backtest against a quarter of confirmed fraud cases without missing the obvious ones. Week two earns a hold list: the model flags transactions but nothing gets blocked yet, analysts just review what it flagged. Week three earns an auto-hold: the highest-confidence flags get held automatically, everything else stays in review. Week four is full volume, live, then the go, no-go.
O, own the numbers. Week one, backtest on a fifty-thousand-transaction daily sample. Week two, shadow-score the full daily volume, analysts review the top two hundred flags a day. Week three, auto-hold the top twenty highest-confidence flags a day, analysts review those plus a spot-check of one hundred more. Week four, full volume, analysts review only what the model holds.
U, use a range. If the model plugs straight into the bank's existing case queue, the full auto-hold list can run by week four. If the bank's fraud rules engine needs custom mapping, week four might only have the newest card product covered, with the rest still in shadow.
N, nail the sanity check. Four analysts, each with about ninety minutes a day of slack beyond their existing caseload, gives the team roughly six hours a day. Two hundred flags a day at about two minutes each is close to seven hours, already over the ceiling, which is why the review count in week two has to be capped at what the team can do, not at what the model flags.
D, direction. The same tension, a different domain. Whether the model plugs into the existing case queue swings the timeline the most, because bank fraud rules engines are almost always bespoke. Analyst bandwidth is the second lever, and just as easy to leave off a kickoff deck.

Same shape, different lever At Thoresen, integration reuse and planner bandwidth were both real levers, with integration swinging the number slightly more. At the bank, integration is nearly always the dominant one, because a bespoke fraud rules engine is harder to reuse than a manufacturer's ERP export. The gate structure doesn't change. Which assumption to protect first does.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the line: gate every week to accuracy and to the customer's real hours, don't schedule scope by the calendar.
Cost: the customer will only pay for a two-week pilot, not four. Cut the part count, not the gates. Backtest and shadow in week one, live and go/no-go in week two, same structure, half the scope.
The model got better: a newer version rarely misses on demand spikes anymore. That doesn't remove the need for gates, it just moves which number you watch, from raw accuracy to how well it handles parts it's never seen before.

Where people run it wrong.
They schedule the part count and the go-live date before they know the model's backtest accuracy.
They size the review workload to what the model can produce, not to what the customer's actual team has hours for.
They treat week four as a demo instead of a go/no-go built on the pilot's own numbers.

How to use it live. Say the equation before naming a single week's plan: "the four weeks aren't the pilot, the four gates are, so before I lay out the calendar, I'd want to know how many hours a week your team actually has free to look at this." That buys the room to ask a real question instead of guessing a schedule that only sounds thorough.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a "design a pilot" question, and why not FLIPS or SPARK?
Tap to flip
ANSWER
BOUND. This is a sizing question: how much of a customer's real work can go live in four weeks, and what that depends on. It's not a person's trust switching between two settings, and it's not a single anchor design decision.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Katarzyna Nowak, who designs enterprise pilots for a demand-forecasting AI tool. Three years in, twelve pilots run, Thoresen Metalworks was her twelfth.
3 · WHAT THE FIRST PLAN GOT WRONG
What did Katarzyna's first draft schedule that limited what it could actually deliver?
Tap to flip
ANSWER
She scheduled all 240 fast-moving parts live by week three on the calendar, without checking whether the model had earned that scope or whether the planners had the hours to review it.
4 · THE STRUCTURE IN THIS STORY
What's the two-setting difference between the first plan and the second?
Tap to flip
ANSWER
Schedule-earned scope, where the part count is set by the calendar date, versus evidence-earned scope, where the part count is set by backtest accuracy and the planners' real review hours.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at first?
Tap to flip
ANSWER
Scheduling all 240 parts live by week three so the week-four meeting could open with a finished result. It made sense in the kickoff room because it sounded fast and decisive to a customer who wanted to look decisive too.
6 · THE NUMBER
Fill in the blank: the under-forecast sat unreviewed for ___ days, until a planner found the part down to ___ days of supply against a ___-day lead time to replace it.
Tap to flip
ANSWER
Four days. Two days of supply. Forty-two days of lead time.
7 · THE REPLAY
Same 240 parts, same four weeks, second design. What changes?
Tap to flip
ANSWER
Same end state, 240 parts live by week four, but ten fewer hours of planner review in the two crunch weeks, and zero forecasts that went out unread.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and which lever swings it most?
Tap to flip
ANSWER
A bank piloting an AI fraud-scoring tool with one regional bank. The lever is the same tension, but integration reuse dominates even harder there, because fraud rules engines are almost always bespoke.

Check yourself Score: 0 / 0

Multiple choice
1. In Katarzyna's redesigned plan, why does week three let planners act live only on the 120 parts already proven in week two, while the newly added 120 stay in shadow?
  • A. Acting live on unproven parts would break the gate that scope only advances once the model has already shown it works.
  • B. The model has a hard technical limit of 120 forecasts at a time.
  • C. Thoresen's ERP system can only import 120 part numbers a week.
  • D. The planners refused to review more than 120 forecasts under any circumstance.
Show hint
Look at what week two was actually for, and what "earned" scope means in this answer.
Show answer
A. The whole point of gating is that scope only widens once accuracy and review hours have both cleared their bar. Going live on the untested 120 would skip straight past that check.
True or false
2. True or false: the near miss with the bearing housing happened only because the model's forecast was wrong.
  • True
  • False
Show hint
Ask what would have happened to that same wrong forecast if the review pile hadn't already been too big to check properly.
Show answer
False. The model did under-forecast that one part. The real danger was that scope had already outgrown the planners' review hours, so a wrong forecast that would normally get caught sailed through unread for four days.
Fill in the blank
3. The redesigned pilot capped week two's review time at ___ hours and week three's at ___ hours, both under the two senior planners' combined ceiling of roughly ___ hours a week.
Show hint
Check the O step numbers for each week, and the N step's stated ceiling.
Show answer
8 hours; 10 hours; 12 hours. Actual hours came in a bit under each cap, six and nine, which is what left any slack at all in the crunch weeks.
Short answer, the number question
4. If Thoresen's real planner ceiling had been 8 hours a week combined instead of 12, would the week-three target of 120 live plus 120 in shadow still have fit inside the gate? Why or why not?
Show hint
Compare week three's actual review load against the smaller ceiling, not the one the plan assumed.
Show answer
No, it wouldn't fit. Week three's actual load was about 9 hours, already past an 8-hour ceiling. The plan would have had to hold the shadow expansion to week four instead, or shrink the live count, exactly the D-step point that planner bandwidth is the assumption that bites first when it's wrong.
Short answer, apply it yourself
5. Think of a pilot or trial you could design at your own job. What would you gate each stage on, and whose real hours would you check before scheduling the last stage's scope?
Show hint
Name a real accuracy or quality bar and a real person's free hours, not "we'll see how it goes."
Show answer
Model answer: "A hospital pharmacy piloting an AI tool that checks prescriptions for drug interactions. I'd gate week one on backtested accuracy against last month's confirmed interaction flags, and check how many spare minutes the two pharmacists on shift actually have between filling orders, not how many the schedule says they have on paper, before promising the tool covers every prescription by week four."
Short answer
6. Weeks three and four of the redesigned pilot need about 20 hours of review between them, against a ceiling of about 24 hours for those two weeks combined. Why does that leftover 4 hours matter, and what happens if Piotr's boss pushes to compress the pilot to three weeks instead of four?
Show hint
Think about what a compressed week would do to the hours needed in the busiest week, against the same ceiling.
Show answer
Model answer: The 4 hours of slack is what lets the plan absorb one bad week, a supplier fire or a sick day, without the review pile backing up the way it did the first time. Compressing to three weeks would force roughly 20 hours of review into a single week against a 12-hour ceiling, which is close to the exact math that caused the first near miss, and would likely mean cutting the part count in half again or accepting that some forecasts won't get properly checked.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more