What data do you need to collect during a pilot that you would not otherwise?
- Collect three things production logging never captures: a yes-or-no on every alert, manual inspections on the riskiest machines, and a note behind every override.Why: a plant's normal logs record what the machine did, never whether the model's warning about it was actually right.
- Size the plan by how many people you can actually free up to confirm data, not by how many alerts the model raises.Why: alerts are free and automatic; a person saying "yes, that one was real" is the part that costs time, and that's the real ceiling.
- Put a floor of manual inspections on the machines you can least afford to misjudge, not every machine equally.Why: a wrong guess on a quiet conveyor motor costs a five-minute walk, a wrong guess on a machine that's already failed once costs a day of downtime.
- Capture the reason behind every override, not just the fact that one happened.Why: a raw disagreement count tells you nothing about whether the technician or the model had it right.
- Check the total against how many hours a small review team can actually read it back.Why: data nobody reads by the end of the pilot might as well never have been collected.
- Know whether headcount or review capacity is the assumption that would move the total most.Why: cut the wrong one and either the confirmations dry up or the review pile goes stale before anyone reads it.
How to answer this, stage by stage
Nobody is grading whether you land on exactly 368. They're grading whether you name what's actually new about pilot data before touching a number, whether the plan survives how many people are really free to do the confirming, and whether the total gets checked against real review hours. Seven moves get you there.
Let's learn
The product is a box of sensors bolted onto factory machines. It listens to vibration and heat and tells a maintenance team which machine is about to break, before it actually breaks.
Before Ferrix installed anything, Corvane Industrial's maintenance crew caught most problems the old way, walking the floor, listening, feeling a bearing get warm under a gloved hand. They were good at it. It just meant fixing things after they'd already started to fail, sometimes with a machine down for a day while a part got ordered.
Ten weeks into the pilot, the sensors were doing what they were built to do. Forty machines, watched around the clock, throwing out an alert here and there. On paper, that looked like exactly what Corvane paid for.
Here is the turn. The extra alerts were not the problem. The problem is nobody had decided, on day one, what to do with each alert once it fired. A technician would glance at it, note it, and keep working, the same way anyone glances past a warning light that's cried wolf before. Nobody wrote down whether the alert turned out to be right. Nobody wrote down why they ignored it when they did.
At its worst, that costs the whole pilot. Ten weeks in, a PM can have forty machines of sensor data and a folder of alerts, and not one confirmed answer to the only question a plant director actually asked: which of these turned out to be real?
The choice I would take back. The first data plan assumed the sensor readings, plus whatever showed up on Corvane's maintenance tickets, counted as the pilot's data. That's what the tool logs on its own, running or not. Nobody built a specific step where a technician looks at each alert and says, right then, real or false, and why they did what they did about it.
What I would leave alone. Not every machine needs the extra layer. Corvane's newer conveyor motors barely ever throw an alert, and a wrong guess there just means a five-minute walk to check nothing. Wiring in a weekly manual inspection for those would be effort with nothing to show for it. Save the extra checking for the machines where getting it wrong actually costs someone a day of downtime.
The lesson. A pilot's job isn't to prove the model can run somewhere real. It's to build, from the first week, the exact record you'll need afterward to say whether it worked. Normal running of the product will never write that record for you.
Now here is the same thing as a story
Skip this if you already believe a pilot needs its own data plan, separate from whatever the product logs on its own. Read on if you want to feel why.
For three weeks, Lindiwe Mokoena's status update to her VP was the same sentence: the sensors are up, the alerts are flowing, everything's on track.
Lindiwe runs product for Ferrix Analytics' predictive-maintenance tool, and Corvane Industrial's pump-manufacturing plant was the company's first real pilot outside a lab. Forty machines. Ten weeks. Corvane's maintenance crew had run that floor for years without any of this, catching a failing bearing by ear and a gloved hand on the housing, well enough that shutdowns were rare and mostly short.
The sensors went in on a Monday. By the second week they were throwing the odd alert, a compressor running hot, a lathe's motor showing a vibration pattern nobody liked the look of. Corvane's technicians would glance at the tablet mounted by the breaker room, note the alert, and go check the machine if they had a minute. Some mornings they did. Some mornings the floor was too busy and the alert just sat there, unopened, until it aged off the screen.
Nobody thought much of it. That's what an alert dashboard is for, isn't it, something you check when you have time.
Around week four, Corvane's plant director sat in on Lindiwe's weekly check-in and asked one question. Not about uptime. Not about the dashboard. He asked: "Of everything it's flagged so far, how many turned out to actually be something?"
Lindiwe didn't have the number. Nobody at Ferrix or Corvane had ever written down, alert by alert, whether it turned out to be real.
Here's the part that actually cost something. It wasn't the one awkward silence in that meeting. It was that four weeks of alerts, close to a hundred and fifty of them by then, had already come and gone with nobody confirming a single one against what the machine actually did next. Getting that confirmation later meant asking technicians to remember a compressor from three weeks back. Most of them couldn't.
So here's what Lindiwe took back. She'd assumed the readings Ferrix's sensors already logged, plus whatever showed up on Corvane's maintenance tickets, counted as the pilot's data. That's what the tool logs on its own, running or not. She'd never built a specific step where a technician looks at each alert and says, right then, yes or no, and why they did what they did about it.
She rebuilt it for the remaining six weeks. Every alert got a one-tap yes-or-no from whichever technician saw it first, real or false. Eight of the oldest, highest-risk machines got a manual inspection every week regardless of whether anything had fired, a second, independent check the model's own record could be measured against. Any time a technician did something other than what the model suggested, a fifteen-second voice note into the tablet, why.
Six weeks, 176 alerts, 56 manual inspections, 34 override notes. Two hundred and sixty-six confirmed data points where before there had been zero. When Lindiwe sat down with Corvane's director in week ten, she had an answer: the model's alerts turned out to be right seventy-one percent of the time, and every miss that mattered clustered on the same two machines, ones with a sensor mounted somewhere it kept picking up vibration from the machine next door.
The thing I'd tell myself, back in that week-four meeting where I had no number: a pilot doesn't need you to build a bigger dataset. It needs you to build the one small record that lets you answer the only question anyone's actually going to ask at the end.
BOUND, sized for forty machines and one plant floor
This is a sizing question about how much extra, human-confirmed data a pilot needs on top of what a plant already logs, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.
B, break it down. The number of pilot-only data points equals three things added together: a human label on every alert the model raises, a fixed floor of manual inspections on the riskiest machines, and a short note every time someone overrides the model.
O, own the numbers. Forty machines, a ten-week pilot, roughly six alerts per machine, gives 240 alerts, each needs a yes-or-no label. Eight of those machines are the ones Corvane can least afford to misjudge, so they get a manual inspection every week regardless of whether an alert fired, 8 times 10 weeks is 80. Technicians override the model on about one alert in five, 240 divided by 5 is 48 short notes on why. Total: 240 plus 80 plus 48 is 368.
U, use a range. That 368 assumes all four of Corvane's maintenance technicians are actually free to confirm alerts in real time. If only two are, because the rest are still doing their normal jobs, an alert older than a couple of days is worthless to label, since nobody remembers the machine well enough to say. Realistically that drops the total to about 210. Start at 368, the full-headcount case, and check which one the plant can actually staff before betting on it.
N, nail the sanity check. 368 entries at about five minutes each to read and code by hand is a little over thirty hours, split across a two-person review team over ten weeks, about an hour and a half a week each. Sane. If the plan had instead tried to label every raw sensor reading instead of just the alerts, that number would cross a hundred hours, more than a small team can actually read back before the pilot ends, and the extra data would go stale on a shelf.
D, direction. Technician headcount moves this more than alert-rate assumptions do. Going from four technicians free to confirm alerts down to two swings the total from 368 to 210, a drop of 158. Doubling the machine count or the alert rate would move it by less than that. If this number needs to shrink, headcount is the first place to look, not the machine count, since cutting machines means losing an entire part of the plant floor from the pilot's evidence.
And if you want to be sure it really works, try it somewhere else
A regional food bank runs the same idea on donated pallets instead of factory machines. A model looks at how long a pallet has sat, what's inside it, and the temperature log on the truck that delivered it, and flags which pallets are likely to spoil before they get distributed.
B, break it down. The number of pilot-only data points equals three things: a spoilage-outcome label on every flagged pallet, a floor of manual checks on pallets the model didn't flag, and a note every time a volunteer overrides a flag and distributes the pallet anyway.
O, own the numbers. An eight-week pilot, about 25 pallets flagged a week, gives 200 flagged pallets, each needs an outcome label, did it actually spoil before use. A floor of 5 non-flagged pallets checked every week as a control, 5 times 8 is 40. Volunteers override roughly one flag in four, 200 divided by 4 is 50 notes on why. Total: 200 plus 40 plus 50 is 290.
U, use a range. With one shift coordinator logging outcomes, weekends mostly get missed, closer to 200. With three coordinators rotating through every shift, closer to the full 290.
N, nail the sanity check. 290 entries at about three minutes each on a clipboard is roughly fourteen and a half hours over eight weeks, under two hours a week for a small volunteer team already on site. Sane for people who aren't being paid to do this.
D, direction. Same shape as Corvane's. Coordinator headcount swings the total from about 200 to 290, a difference of 90, more than any assumption about how many pallets get flagged in the first place.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at 90 seconds. Skip straight to the split: three categories, 368 at full headcount, 210 at half, about thirty hours to review either way.
Cost: instead of a data-point count, a manager caps review time at twenty hours. Work backward: at five minutes an entry, that's 240 entries, cover alert labels first and let manual inspections and override notes fill whatever's left.
The model got better: a newer version rarely misses an obvious failure anymore. The label count doesn't drop because of that, a quiet model still needs proof it's quiet for the right reason, not just guessing safe.
Where people run it wrong.
They assume whatever the tool logs automatically, sensor readings, tickets closed, already is the pilot's data.
They size the plan by how many alerts the model happens to raise instead of by how many people can actually confirm them.
They let the collection plan grow until reading it back costs more hours than a small team actually has, so half the data never gets looked at.
How to use it live. Say the equation before any number: "pilot-only data is a human label on every alert, a floor of manual checks on the riskiest cases, and a note behind every override, sized to who's free to confirm it, not to how much the model produces." That buys the time to name real categories instead of reaching for "collect everything" and calling it thorough.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Pilot design and POC-to-production
- #1 Design a four-week pilot for an AI feature with one enterprise customer.
- #2 What success criteria should be agreed before a pilot begins?
- #3 Explain the difference between a pilot and a beta.
- #4 How do you choose pilot customers, and what makes a bad one?
- #5 Describe the pilot-to-production gap and the work that lives in it.
- #6 Why do most AI POCs fail to reach production? Give four reasons.