Design a four-week pilot for an AI feature with one enterprise customer.
- Gate every week to earned evidence, not a calendar date.Why: a schedule that expands scope by date instead of by proof is what lets an unreviewed pile of forecasts build up behind the scenes.
- Size each week's scope to the customer's real review hours, not to what the model can technically forecast.Why: the model can score four hundred parts overnight; two real planners can only check a fraction of that on top of their day job.
- Start narrow and backtested before the model makes one live ordering call.Why: week one's backtest is the cheapest place in the whole pilot to catch a bad forecast, before it touches a real order.
- Give the go/no-go a range, not one confident number, for how many parts will be live by week four.Why: forty parts if the integration ends up custom, two hundred forty if it's reused, and the room deserves both numbers before week one starts.
- Push hard in week zero to reuse the ERP integration wherever the customer's system allows it.Why: of the two big levers, this is the one that swings the final count the most, so it's worth negotiating before the pilot even starts.
- Leave the true one-off custom parts out of the pilot entirely.Why: a part that only ever gets ordered once has no history to learn from, so forecasting it isn't a shortcut, it's a guess wearing a spreadsheet.
How to answer this, stage by stage
Nobody is grading whether you land on exactly two hundred forty parts. They're grading whether you can defend what earns each week the right to the next one, whether the hours add up, and whether you close on a number the room can check. Eight moves get you there.
Let's learn
The product is a tool that reads a manufacturer's order history, current stock, and supplier lead times, and drafts a demand forecast for every part they carry, instead of a planner rebuilding the same spreadsheet from scratch every Monday.
Before a pilot like this gets planned, a team told to "get it live with the customer in four weeks" usually writes the plan the way any project plan gets written: build in week one and two, test in week three, review in week four. On paper, that covers every part the customer makes from day one. It reads like a serious pilot.
Here's the turn. Getting the model connected and forecasting is not the hard part. Earning the right to let it make a live ordering decision is. A plan that puts every part live on a fixed date, whether or not the model and the customer's team are actually ready, isn't really a plan. It's a hope with a calendar taped to it.
At its worst, that gap shows up in week two, in the size of the pile a customer's own planners have to check by hand. If the schedule, not the evidence, decides how much of the plant runs on the model's numbers by a set date, the planners absorb the difference. And a pile that's too big to check gets skimmed instead of checked, which is exactly the moment a bad forecast can slip through unnoticed.
The choice I would take back. On the first pass, I planned the pilot by working backward from the demo date: all two hundred forty of the customer's fast-moving parts, live by week three, so the week-four meeting could open with a finished result instead of a partial one. That felt decisive. It wasn't a real gate. It was a due date wearing a pilot's clothes.
What I would leave alone. The true one-off parts, the ones made once to a single customer's exact spec and never reordered, don't belong in this pilot at all, and they never will. A forecast for a part that will only ever be ordered once isn't a forecast. It's a guess with extra steps. Leave those on the planners' own spreadsheet, forever, and don't spend a single pilot week on them.
The lesson. Four weeks is a fine amount of time for one enterprise pilot. The problem was never the calendar. It was letting the calendar decide how many parts went live, instead of letting the hours the planners actually had left decide it.
Now here is the same thing as a story
The short version is above. Read on if you want to feel why the gates are the whole decision, not a scheduling detail before the real work starts.
Katarzyna Nowak has designed pilots like this for three years, twelve of them by the time this one landed on her desk. She's good at it. She can tell within the first data pull whether a customer's export is going to be clean or a mess, just from how the column headers are named.
Thoresen Metalworks was her twelfth. Thoresen runs a precision-parts plant that forecasts demand for about four hundred repeat-order parts by hand every week, two hundred forty of which move fast enough to carry most of the plant's order value. Piotr Kaczmarek, who runs supply-chain planning there, had spent months pushing his own leadership to approve the pilot, and he wanted the four weeks to look decisive: real scope, real speed, a real answer by the end.
Katarzyna wrote the plan the way she'd written the last few. Week one and two, connect the data and build the model. Week three, run it live across all two hundred forty of Thoresen's fast movers. Week four, review the results with Piotr's boss. It read well in the kickoff deck. Everyone signed off in twenty minutes.
Week one went fine. The connector took three days instead of five, and the backtest looked solid enough on all two hundred forty parts to keep the calendar as written.
Then week two arrived, and with it, the actual size of the pile Piotr's three planners had to check by hand: two hundred forty new forecasts, on top of a supplier delay that was already eating most of their week. Only two of the three planners knew these parts well enough to judge whether a forecast looked right, and between them they had about twelve hours free once the firefighting was accounted for. Two hundred forty forecasts needed close to twelve hours to check properly. By Wednesday, they weren't checking properly. They were opening a forecast, glancing at the total, and moving to the next one, because glancing at two hundred forty lines and doing the job they were actually hired for both had to fit inside the same day.
On Thursday, one of those glances waved through a forecast that had under-ordered a bearing housing with a six-week lead time from its only qualified supplier. Nobody caught it. Not Thursday, not Friday. It sat there until Monday morning, when a planner doing an unrelated stock check noticed the part was down to two days of supply against a forty-two day lead time to replace it.
They caught it with two days to spare, by paying for an emergency shipment that cost more than the part itself was worth. No line stopped. No customer of Thoresen's ever knew. But Piotr knew, because he called Katarzyna that Monday afternoon.
Katarzyna paused the pilot for two days and rebuilt the plan. Same four weeks, same customer, same two hundred forty parts eventually, but scope now had to be earned week by week against two things: the model's own backtested accuracy, and how many hours Piotr's two senior planners actually had free once the firefighting was taken out. Week one, backtest on one hundred twenty parts, not two hundred forty. Week two, shadow only, same one hundred twenty, capped at eight review hours. Week three, expand into shadow for the remaining one hundred twenty while the first batch goes live, capped at ten hours. Week four, all two hundred forty live, one real ordering cycle, then the actual go, no-go conversation, backed by numbers instead of a date on a slide.
In the original kickoff, when Katarzyna proposed "all two hundred forty live by week three," nobody in the room pushed back, because it sounded like exactly the kind of pilot a customer wants to see: fast, comprehensive, decisive. It made sense in that room, in that hour. It stopped making sense the moment the actual review hours had to come out of two real people's real week.
The second version of the pilot reached the same place, two hundred forty parts live, by the end of the same four weeks. The difference was ten fewer hours of planner review in the two crunch weeks, and zero forecasts that went out unread.
The thing I'd tell myself, in that first kickoff meeting: a pilot that looks decisive on a slide and a pilot that's actually been earned are not the same four weeks, and the only way to tell them apart in week zero is to ask whose hours the schedule is actually spending.
BOUND, gated before the kickoff deck goes out
This is a sizing question about how much of a customer's real work can go live in four weeks, and what that number depends on, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.
B, break it down. Four weeks isn't the plan. Four gates are. Week one earns the right to run in shadow: the pipeline has to work without hand-patched data, and the backtest has to beat or match the planners' own historical accuracy on most of the parts tested. Week two earns the right to go live: the shadow forecast has to land inside an agreed error band, and the review time has to fit inside the hours the planners actually have. Week three earns the right to expand: no stockout or overstock event traceable to the model among the live parts, and the review load still has to fit. Week four is the live run and the real decision, not a demo.
O, own the numbers. For Thoresen's two hundred forty fast-moving parts: week one, backtest against eight weeks of real demand on the first one hundred twenty. Week two, shadow-run those same one hundred twenty for one real ordering cycle, capped at eight review hours, coming in around six. Week three, add the remaining one hundred twenty into shadow while the first batch goes live, capped at ten hours, coming in around nine. Week four, all two hundred forty live for one real cycle, review and monitoring around eleven hours.
U, use a range. Before running anything, the honest range for how many parts are live by week four is wide. If Thoresen's ordering system already has a connector built for a past customer, all two hundred forty can realistically be live. If it's old enough that every field gets mapped by hand, week four might only have about forty parts live, with the rest still sitting in backtest. Same four weeks, a six-times difference, depending only on how much of the integration gets reused.
N, nail the sanity check. Twenty-nine hours of planner review across four weeks looks fine against a team with roughly twelve hours a week free, on average. But nine of those hours land in week three and eleven in week four, twenty hours across two weeks against a twenty-four hour ceiling for that pair of weeks. That's four hours of slack across the two busiest weeks of the pilot, which is exactly tight enough to survive and exactly tight enough to explain why the first draft, which needed nearly twelve hours in week two alone, didn't.
D, direction. Two assumptions could change this number a lot, and two barely move it. Whether the ERP integration gets reused instead of built custom swings the parts live at week four by about two hundred. Whether the two senior planners' real slack is six hours a week instead of twelve swings it by about one hundred fifty. Freeing up a fourth planner just for pilot review moves it by about forty. Extending the backtest window moves it by about ten. Integration reuse is the bigger number on paper, but bandwidth is the one that quietly kills a pilot, because you can throw two extra engineers at a connector overnight. You can't hand a planner more hours in their week.
And if you want to be sure it really works, try it somewhere else
A bank pilots an AI tool that scores every card transaction for fraud risk, with one regional bank as the enterprise customer, over the same four weeks.
B, break it down. Same shape, different work. Week one earns shadow scoring: the model has to backtest against a quarter of confirmed fraud cases without missing the obvious ones. Week two earns a hold list: the model flags transactions but nothing gets blocked yet, analysts just review what it flagged. Week three earns an auto-hold: the highest-confidence flags get held automatically, everything else stays in review. Week four is full volume, live, then the go, no-go.
O, own the numbers. Week one, backtest on a fifty-thousand-transaction daily sample. Week two, shadow-score the full daily volume, analysts review the top two hundred flags a day. Week three, auto-hold the top twenty highest-confidence flags a day, analysts review those plus a spot-check of one hundred more. Week four, full volume, analysts review only what the model holds.
U, use a range. If the model plugs straight into the bank's existing case queue, the full auto-hold list can run by week four. If the bank's fraud rules engine needs custom mapping, week four might only have the newest card product covered, with the rest still in shadow.
N, nail the sanity check. Four analysts, each with about ninety minutes a day of slack beyond their existing caseload, gives the team roughly six hours a day. Two hundred flags a day at about two minutes each is close to seven hours, already over the ceiling, which is why the review count in week two has to be capped at what the team can do, not at what the model flags.
D, direction. The same tension, a different domain. Whether the model plugs into the existing case queue swings the timeline the most, because bank fraud rules engines are almost always bespoke. Analyst bandwidth is the second lever, and just as easy to leave off a kickoff deck.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the line: gate every week to accuracy and to the customer's real hours, don't schedule scope by the calendar.
Cost: the customer will only pay for a two-week pilot, not four. Cut the part count, not the gates. Backtest and shadow in week one, live and go/no-go in week two, same structure, half the scope.
The model got better: a newer version rarely misses on demand spikes anymore. That doesn't remove the need for gates, it just moves which number you watch, from raw accuracy to how well it handles parts it's never seen before.
Where people run it wrong.
They schedule the part count and the go-live date before they know the model's backtest accuracy.
They size the review workload to what the model can produce, not to what the customer's actual team has hours for.
They treat week four as a demo instead of a go/no-go built on the pilot's own numbers.
How to use it live. Say the equation before naming a single week's plan: "the four weeks aren't the pilot, the four gates are, so before I lay out the calendar, I'd want to know how many hours a week your team actually has free to look at this." That buys the room to ask a real question instead of guessing a schedule that only sounds thorough.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Pilot design and POC-to-production
- #2 What success criteria should be agreed before a pilot begins?
- #3 Explain the difference between a pilot and a beta.
- #4 How do you choose pilot customers, and what makes a bad one?
- #5 Describe the pilot-to-production gap and the work that lives in it.
- #6 Why do most AI POCs fail to reach production? Give four reasons.
- #7 What data do you need to collect during a pilot that you would not otherwise?