ConceptIntermediateShipping & Model Lifecycle / Pilot design and POC-to-production / #7

What data do you need to collect during a pilot that you would not otherwise?

The direct answer
Collect three things ordinary production logging never captures: a human yes-or-no on every alert the model raises, a fixed floor of manual inspections on your riskiest machines, and a short note every time someone overrides the model. Size all three by how many people you can actually free up to confirm them, not by how many alerts the model happens to raise, then check the total against how many hours a small review team can read back by hand. A pilot that only logs what production already logs proves the sensors work, not that the model does.
Do this, in order
  1. Collect three things production logging never captures: a yes-or-no on every alert, manual inspections on the riskiest machines, and a note behind every override.Why: a plant's normal logs record what the machine did, never whether the model's warning about it was actually right.
  2. Size the plan by how many people you can actually free up to confirm data, not by how many alerts the model raises.Why: alerts are free and automatic; a person saying "yes, that one was real" is the part that costs time, and that's the real ceiling.
  3. Put a floor of manual inspections on the machines you can least afford to misjudge, not every machine equally.Why: a wrong guess on a quiet conveyor motor costs a five-minute walk, a wrong guess on a machine that's already failed once costs a day of downtime.
  4. Capture the reason behind every override, not just the fact that one happened.Why: a raw disagreement count tells you nothing about whether the technician or the model had it right.
  5. Check the total against how many hours a small review team can actually read it back.Why: data nobody reads by the end of the pilot might as well never have been collected.
  6. Know whether headcount or review capacity is the assumption that would move the total most.Why: cut the wrong one and either the confirmations dry up or the review pile goes stale before anyone reads it.

How to answer this, stage by stage

Nobody is grading whether you land on exactly 368. They're grading whether you name what's actually new about pilot data before touching a number, whether the plan survives how many people are really free to do the confirming, and whether the total gets checked against real review hours. Seven moves get you there.

1
Scope it to one real pilot, not a category of pilots
Say it like this
"Let's ground this. Say I'm ten weeks into a pilot of a predictive-maintenance tool at Corvane Industrial, a hydraulic pump manufacturer. We've put sensors on forty machines on their plant floor, and the model's starting to raise alerts."
Why this works
A named plant and a named machine count give the interviewer something to push on, instead of a shrug about "pilots in general."
2
Reframe the question before naming a single category
Say it like this
"This isn't asking me to log everything the sensors can technically capture. It's asking what a pilot lets me get that a normal, live run of this same tool never would, because once it's live, nobody's standing there confirming each alert."
Why this works
Rules out "collect everything" before it becomes the answer. That instinct sounds thorough and produces nothing anyone can act on.
3
State the build-up before touching a number
Say it like this
"The number of pilot-only data points equals three things added together: a human label on every alert, a fixed floor of manual inspections on the riskiest machines, and a short note every time someone overrides the model."
Why this works
States the equation before a single figure lands, so what follows reads as arithmetic, not a guess dressed up as one.
4
Own the numbers
Say it like this
"Forty machines, ten weeks, the model raises about six alerts a machine, that's 240 alerts, each needs a yes-or-no label. Eight of those machines are the ones Corvane can least afford to misjudge, those get a manual check every week regardless of whether the model says anything, 80 more. Technicians override the model on about one alert in five, that's 48 short notes on why. 240 plus 80 plus 48 is 368."
Why this works
Turns "as much data as we can get" into a number a reviewer can check against a real, named list of categories.
5
Give the range, tied to who's actually available to observe
Say it like this
"That 368 assumes all four of Corvane's maintenance technicians are free to confirm alerts. If only two really are, because the other two are still doing their normal jobs, an alert that sits unconfirmed for a couple of days is worthless, nobody remembers what the machine sounded like. Realistically that's closer to 210."
Why this works
Ties the range to real headcount, not a made-up confidence interval that sounds rigorous but explains nothing.
6
Test the total against a small team's real review hours
Say it like this
"368 entries at about five minutes each to read and code is just over thirty hours, split across two people on our side over ten weeks, that's an hour and a half a week each. If we tried to label every sensor reading instead of just the alerts, we'd be past a hundred hours, and nobody's reading that back before the pilot ends."
Why this works
Forces the plan to survive real review capacity, not just collection capacity, which is the check most estimation answers skip.
7
Name the lever, then close in one breath
Say it like this
"If I had to bet on what moves this number, it's technician headcount, not alert volume. Going from two available technicians to four swings the total from 210 to 368, more than doubling the alert rate would. So: three kinds of data, 368 at full headcount, 210 if we only get two technicians, about thirty hours to review either way, and headcount is the first thing I'd check before assuming the pilot needs more weeks or more machines."
Why this works
Answers the hardest follow-up directly and closes in one breath, the way a strong answer actually sounds.
If you remember one thing A pilot that only logs what the tool would log anyway, sensor readings, tickets closed, will always look busy. Size the extra layer as a floor per category, human-confirmed, checked against how many hours a small team can actually read it back.

Let's learn

The product is a box of sensors bolted onto factory machines. It listens to vibration and heat and tells a maintenance team which machine is about to break, before it actually breaks.

Knowledge spark: what does a plant already log, pilot or not? Two things, automatically. What the sensors read, and whether a maintenance ticket got closed. Neither one says whether an alert turned out to be right. That answer only exists if a person writes it down, and normal running of the tool never asks anyone to.

Before Ferrix installed anything, Corvane Industrial's maintenance crew caught most problems the old way, walking the floor, listening, feeling a bearing get warm under a gloved hand. They were good at it. It just meant fixing things after they'd already started to fail, sometimes with a machine down for a day while a part got ordered.

Ten weeks into the pilot, the sensors were doing what they were built to do. Forty machines, watched around the clock, throwing out an alert here and there. On paper, that looked like exactly what Corvane paid for.

Here is the turn. The extra alerts were not the problem. The problem is nobody had decided, on day one, what to do with each alert once it fired. A technician would glance at it, note it, and keep working, the same way anyone glances past a warning light that's cried wolf before. Nobody wrote down whether the alert turned out to be right. Nobody wrote down why they ignored it when they did.

We did not fail to build the model. We failed to build a way to know if it worked.

At its worst, that costs the whole pilot. Ten weeks in, a PM can have forty machines of sensor data and a folder of alerts, and not one confirmed answer to the only question a plant director actually asked: which of these turned out to be real?

The decision that mattered Size the pilot's data by what a normal, live run of the tool would never bother to ask for, a human confirmation on every alert, not by whatever the sensors were already going to log anyway.

The choice I would take back. The first data plan assumed the sensor readings, plus whatever showed up on Corvane's maintenance tickets, counted as the pilot's data. That's what the tool logs on its own, running or not. Nobody built a specific step where a technician looks at each alert and says, right then, real or false, and why they did what they did about it.

What I would leave alone. Not every machine needs the extra layer. Corvane's newer conveyor motors barely ever throw an alert, and a wrong guess there just means a five-minute walk to check nothing. Wiring in a weekly manual inspection for those would be effort with nothing to show for it. Save the extra checking for the machines where getting it wrong actually costs someone a day of downtime.

The lesson. A pilot's job isn't to prove the model can run somewhere real. It's to build, from the first week, the exact record you'll need afterward to say whether it worked. Normal running of the product will never write that record for you.

Now here is the same thing as a story

Skip this if you already believe a pilot needs its own data plan, separate from whatever the product logs on its own. Read on if you want to feel why.

For three weeks, Lindiwe Mokoena's status update to her VP was the same sentence: the sensors are up, the alerts are flowing, everything's on track.

Lindiwe runs product for Ferrix Analytics' predictive-maintenance tool, and Corvane Industrial's pump-manufacturing plant was the company's first real pilot outside a lab. Forty machines. Ten weeks. Corvane's maintenance crew had run that floor for years without any of this, catching a failing bearing by ear and a gloved hand on the housing, well enough that shutdowns were rare and mostly short.

The sensors went in on a Monday. By the second week they were throwing the odd alert, a compressor running hot, a lathe's motor showing a vibration pattern nobody liked the look of. Corvane's technicians would glance at the tablet mounted by the breaker room, note the alert, and go check the machine if they had a minute. Some mornings they did. Some mornings the floor was too busy and the alert just sat there, unopened, until it aged off the screen.

Nobody thought much of it. That's what an alert dashboard is for, isn't it, something you check when you have time.

Around week four, Corvane's plant director sat in on Lindiwe's weekly check-in and asked one question. Not about uptime. Not about the dashboard. He asked: "Of everything it's flagged so far, how many turned out to actually be something?"

Lindiwe didn't have the number. Nobody at Ferrix or Corvane had ever written down, alert by alert, whether it turned out to be real.

Hand-sketched comparison. Left panel, a gauge icon labeled what running it logs, captioned sensor readings, a ticket closed. Right panel, a question-mark box labeled what a pilot needs, captioned a yes or no, an inspection, a reason why.
The same forty machines, logged two different ways. One only ever recorded what the sensors would have recorded with nobody watching. The other recorded whether the warning was actually right.
We did not fail to build the model. We failed to build a way to know if it worked.

Here's the part that actually cost something. It wasn't the one awkward silence in that meeting. It was that four weeks of alerts, close to a hundred and fifty of them by then, had already come and gone with nobody confirming a single one against what the machine actually did next. Getting that confirmation later meant asking technicians to remember a compressor from three weeks back. Most of them couldn't.

So here's what Lindiwe took back. She'd assumed the readings Ferrix's sensors already logged, plus whatever showed up on Corvane's maintenance tickets, counted as the pilot's data. That's what the tool logs on its own, running or not. She'd never built a specific step where a technician looks at each alert and says, right then, yes or no, and why they did what they did about it.

She rebuilt it for the remaining six weeks. Every alert got a one-tap yes-or-no from whichever technician saw it first, real or false. Eight of the oldest, highest-risk machines got a manual inspection every week regardless of whether anything had fired, a second, independent check the model's own record could be measured against. Any time a technician did something other than what the model suggested, a fifteen-second voice note into the tablet, why.

Six weeks, 176 alerts, 56 manual inspections, 34 override notes. Two hundred and sixty-six confirmed data points where before there had been zero. When Lindiwe sat down with Corvane's director in week ten, she had an answer: the model's alerts turned out to be right seventy-one percent of the time, and every miss that mattered clustered on the same two machines, ones with a sensor mounted somewhere it kept picking up vibration from the machine next door.

The thing I'd tell myself, back in that week-four meeting where I had no number: a pilot doesn't need you to build a bigger dataset. It needs you to build the one small record that lets you answer the only question anyone's actually going to ask at the end.

BOUND, sized for forty machines and one plant floor

This is a sizing question about how much extra, human-confirmed data a pilot needs on top of what a plant already logs, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.

B, break it down. The number of pilot-only data points equals three things added together: a human label on every alert the model raises, a fixed floor of manual inspections on the riskiest machines, and a short note every time someone overrides the model.
O, own the numbers. Forty machines, a ten-week pilot, roughly six alerts per machine, gives 240 alerts, each needs a yes-or-no label. Eight of those machines are the ones Corvane can least afford to misjudge, so they get a manual inspection every week regardless of whether an alert fired, 8 times 10 weeks is 80. Technicians override the model on about one alert in five, 240 divided by 5 is 48 short notes on why. Total: 240 plus 80 plus 48 is 368.
U, use a range. That 368 assumes all four of Corvane's maintenance technicians are actually free to confirm alerts in real time. If only two are, because the rest are still doing their normal jobs, an alert older than a couple of days is worthless to label, since nobody remembers the machine well enough to say. Realistically that drops the total to about 210. Start at 368, the full-headcount case, and check which one the plant can actually staff before betting on it.
N, nail the sanity check. 368 entries at about five minutes each to read and code by hand is a little over thirty hours, split across a two-person review team over ten weeks, about an hour and a half a week each. Sane. If the plan had instead tried to label every raw sensor reading instead of just the alerts, that number would cross a hundred hours, more than a small team can actually read back before the pilot ends, and the extra data would go stale on a shelf.
D, direction. Technician headcount moves this more than alert-rate assumptions do. Going from four technicians free to confirm alerts down to two swings the total from 368 to 210, a drop of 158. Doubling the machine count or the alert rate would move it by less than that. If this number needs to shrink, headcount is the first place to look, not the machine count, since cutting machines means losing an entire part of the plant floor from the pilot's evidence.

The build-up: three categories, one plant floor
240 alert labels (40 machines, 6 alerts each)240
+ 80 manual inspections, 8 riskiest machines weekly320
+ 48 override notes, 1 in 5 alerts368
The alert labels alone are two-thirds of the total. The two smaller categories add just 128 more, but they're the 128 that would have told Lindiwe which machines to actually trust the model on.
What moves the total most
Technicians free to confirm alerts drops from 4 to 2−158
Alert rate per machine rises from 6 to 8 over the pilot+96
Override rate falls from 1-in-5 alerts to 1-in-8−18
One more machine added to the weekly-inspection floor+10
Technician headcount swings the total nearly twice as hard as anything else on this list, in either direction. That's why headcount, not alert volume, is the first thing worth re-checking if this number needs to move.

And if you want to be sure it really works, try it somewhere else

A regional food bank runs the same idea on donated pallets instead of factory machines. A model looks at how long a pallet has sat, what's inside it, and the temperature log on the truck that delivered it, and flags which pallets are likely to spoil before they get distributed.

B, break it down. The number of pilot-only data points equals three things: a spoilage-outcome label on every flagged pallet, a floor of manual checks on pallets the model didn't flag, and a note every time a volunteer overrides a flag and distributes the pallet anyway.
O, own the numbers. An eight-week pilot, about 25 pallets flagged a week, gives 200 flagged pallets, each needs an outcome label, did it actually spoil before use. A floor of 5 non-flagged pallets checked every week as a control, 5 times 8 is 40. Volunteers override roughly one flag in four, 200 divided by 4 is 50 notes on why. Total: 200 plus 40 plus 50 is 290.
U, use a range. With one shift coordinator logging outcomes, weekends mostly get missed, closer to 200. With three coordinators rotating through every shift, closer to the full 290.
N, nail the sanity check. 290 entries at about three minutes each on a clipboard is roughly fourteen and a half hours over eight weeks, under two hours a week for a small volunteer team already on site. Sane for people who aren't being paid to do this.
D, direction. Same shape as Corvane's. Coordinator headcount swings the total from about 200 to 290, a difference of 90, more than any assumption about how many pallets get flagged in the first place.

Same shape, different lever At Corvane, technician headcount decided how many alerts got a real confirmation. At the food bank, coordinator headcount decides how many flagged pallets get a real outcome label. Different plant floor, same rule: who's actually free to confirm the data moves the total more than how much of it the model produces.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at 90 seconds. Skip straight to the split: three categories, 368 at full headcount, 210 at half, about thirty hours to review either way.
Cost: instead of a data-point count, a manager caps review time at twenty hours. Work backward: at five minutes an entry, that's 240 entries, cover alert labels first and let manual inspections and override notes fill whatever's left.
The model got better: a newer version rarely misses an obvious failure anymore. The label count doesn't drop because of that, a quiet model still needs proof it's quiet for the right reason, not just guessing safe.

Where people run it wrong.
They assume whatever the tool logs automatically, sensor readings, tickets closed, already is the pilot's data.
They size the plan by how many alerts the model happens to raise instead of by how many people can actually confirm them.
They let the collection plan grow until reading it back costs more hours than a small team actually has, so half the data never gets looked at.

How to use it live. Say the equation before any number: "pilot-only data is a human label on every alert, a floor of manual checks on the riskiest cases, and a note behind every override, sized to who's free to confirm it, not to how much the model produces." That buys the time to name real categories instead of reaching for "collect everything" and calling it thorough.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a question about pilot-specific data collection, and why not FLIPS?
Tap to flip
ANSWER
BOUND. This is a sizing question, how much extra, human-confirmed data a pilot needs on top of what the plant already logs, not a person's trust flipping between two settings.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Lindiwe Mokoena, product manager for Ferrix Analytics' predictive-maintenance tool, piloted at Corvane Industrial, a hydraulic pump manufacturer.
3 · THE HABIT
What did Lindiwe's team wrongly assume covered "the pilot's data" for the first four weeks?
Tap to flip
ANSWER
That the sensor readings Ferrix's tool already logged, plus whatever showed up on Corvane's maintenance tickets, were enough. That's what the tool logs on its own, running or not.
4 · THE BUILD-UP, IN THIS STORY
What's the sizing build-up this answer turns on?
Tap to flip
ANSWER
240 alert labels (40 machines times 6 alerts each), plus 80 manual inspections (8 riskiest machines, weekly, for 10 weeks), plus 48 override notes (1 in 5 alerts), for 368 pilot-only data points.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at first?
Tap to flip
ANSWER
Assuming the sensor readings and maintenance tickets Corvane already generates counted as the pilot's data, instead of building a step where a technician confirms each alert. It made sense because that data was free and already flowing, until the plant director asked which alerts were actually real.
6 · THE NUMBER
Fill in the blank: 40 machines at 6 alerts each, plus 80 inspections and 48 override notes, comes to ___ pilot-only data points, about ___ hours to review by hand.
Tap to flip
ANSWER
368 data points, a little over 30 hours.
7 · THE REPLAY
Same pilot, rebuilt data plan for the last six weeks. What changes?
Tap to flip
ANSWER
176 alert labels, 56 manual inspections, 34 override notes, 266 confirmed data points where before there had been zero. By week ten, Lindiwe can tell Corvane's director the model's alerts were right 71 percent of the time, and that the misses clustered on two machines with a misplaced sensor.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question for a different product. Which product, and what plays the role technician headcount played at Corvane?
Tap to flip
ANSWER
A food bank's pallet-spoilage prediction tool. Volunteer shift-coordinator headcount plays the same role there, swinging the total from about 200 to 290.

Check yourself Score: 0 / 0

True or false, with why
1. True or false: most of the 368 pilot data points come from letting the model run and log its own alerts automatically.
  • True
  • False
Show hint
Ask what the model does on its own versus what a person has to add on top.
Show answer
False. The model raising an alert costs nothing extra, it happens whether or not anyone's running a pilot. The 368 counts only the human-added layer on top: a technician's yes-or-no label, a manual inspection, or an override note. That layer is the actual pilot-only data, and it's the part a normal production run never bothers to collect.
Multiple choice
2. Why does technician headcount move the total more than the alert-rate assumption does?
  • A. Technicians are paid overtime once alert volume rises.
  • B. An alert the model raises only counts toward the pilot's real data once a person confirms it, and that confirming capacity is capped by headcount, not by how many alerts fire.
  • C. More alerts always mean the model has gotten less accurate.
  • D. Corvane's plant director only trusts numbers that come from technicians directly.
Show hint
Ask what actually turns a raw alert into a usable pilot data point.
Show answer
B. The model can raise as many alerts as it wants for free. What's scarce is a person confirming each one before it goes stale. That confirming capacity is set by how many technicians are actually free, which is why headcount, not alert volume, is the lever that swings the total.
Fill in the blank
3. Forty machines at six alerts each over ten weeks is ___ alert labels. Add 80 manual inspections and 48 override notes and the total is ___ pilot-only data points.
Show hint
Check the O step's own arithmetic.
Show answer
240 alert labels; 368 total. 40 times 6 is 240. 240 plus 80 plus 48 is 368.
Short answer, apply it yourself
4. Pick an AI product you use that runs quietly in the background. What's one thing about whether it's actually right that nobody currently writes down?
Show hint
Look for the moment the product makes a call and nobody ever circles back to check it.
Show answer
Model answer: "A spam filter quietly moves emails into a folder I rarely open. Nobody, including me, ever confirms whether those emails were actually spam. The product only knows I complained when I manually pull one back out, so it never learns from the times it was wrong and I simply never noticed."
Multiple choice
5. What old decision does this answer actually take back?
  • A. Installing sensors on forty machines instead of twenty.
  • B. Assuming the sensor readings and maintenance tickets Corvane already generates counted as the pilot's data, without building a step where a technician confirms each alert.
  • C. Running the pilot for ten weeks instead of twenty.
  • D. Mounting the alert tablet by the breaker room instead of on the shop floor.
Show hint
Look for the actual data-sourcing decision made before the pilot started, not a hardware or scheduling choice.
Show answer
B. The first data plan only ever counted what the tool logged on its own. It had no way to say whether an alert was right, because nobody had built the step that asks a person to confirm it.
Short answer, the number question
6. If Corvane could only free up two technicians for the pilot instead of four, roughly how many pilot data points would you actually collect, and why is that lower than 368?
Show hint
Check the U step's range, then think about what happens to an alert that sits too long unconfirmed.
Show answer
About 210. With only two technicians free, alerts pile up faster than they can be confirmed, and an alert older than a couple of days is worthless to label since nobody remembers the machine well enough to say yes or no. Some alerts, manual inspections, and override notes simply never get logged, dropping the realistic total from 368 to about 210.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more