ConceptAdvancedAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #7
How many examples do you need before a feasibility judgement is trustworthy?
BOUNDforty photos said it was ready
Kestrion Precision Works machines turbine and pump parts, and RivetEye is the photo-based tool at the inspection bench that checks each part for six kinds of defect before it ships. Marisol Iyengar is the AI PM who had to decide how many example photos were enough to trust a feasibility call.
The direct answer
Do not pick one round number like 40 or 100. Work backward from the rarest thing you need to catch. If a defect shows up in 3 percent of parts, you need enough examples of that one defect, not just a big pile of easy ones, before its accuracy number means anything. In practice that is usually several hundred examples, spread across every real condition the tool will see, not a few dozen clean demo shots. A sample is trustworthy when the rarest class you care about has shown up enough times to move the needle, not when the total count feels big enough to relax.
Do this, in order
Size the sample off your rarest defect, not your total photo count.Why: a big pile of common examples can hide a rare class you have barely tested.
State the sample-size equation out loud before trusting any accuracy number.Why: a percentage with no equation behind it is a guess wearing a lab coat.
Pull examples from every real production line and lighting condition, not one staged batch.Why: a demo set proves the model can see. It does not prove it can see your shop floor.
Give a range for the required sample size, not one exact figure.Why: false precision invites a fight over the wrong decimal.
Keep growing the sample until the rarest class stops moving the accuracy number.Why: that is the real signal you have seen enough, not a fixed week on a calendar.
Say plainly when a small sample really is enough, like one common defect with no rare cousins.Why: shows judgment about where the extra examples earn their cost, not fear applied everywhere equally.
How to answer this, stage by stage
Nobody is scoring whether you can name a big enough number off the top of your head. They are scoring whether you can show your work for why that number, and not a smaller one, is the one worth trusting.
Stage 1
Scope it to one real feasibility call
Say it like this
"Let's ground this in RivetEye, the defect scanner at Kestrion Precision Works. That's the call where '40 photos and 96 percent' turned out to mean a lot less than it sounded like."
Why this works
Keeps the answer from floating into an abstract debate about sample sizes with nothing real behind it.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as BOUND. Break it down, the actual equation for how big a sample needs to be. Own the numbers. Use a range. Nail the sanity check. Direction, which assumption moves it most."
Why this works
Signals a repeatable way to size a sample, instead of a gut-feel number picked to sound thorough.
Stage 3
Reframe: it isn't "how many in total," it's "how many times have you actually seen your rarest case"
Say it like this
"This isn't really a question about hitting some total count. It's a question of whether the rarest defect you care about has shown up often enough that a good result on it actually means something."
Why this works
This is where a strong answer separates from someone who just says "a few hundred" and stops there.
Stage 4
Break the equation down, out loud
Say it like this
"The sample needs three things stacked together: enough of each defect type to see it at all, enough of the rarest one to trust its rate, and enough spread across production lines to catch real variation. Add those up and that's your number, not a single round guess."
Why this works
Stating the arithmetic before touching a single figure is what makes the rest of the answer checkable, not just asserted.
Stage 5
Own the numbers
Say it like this
"I'm assuming six known defect types, the rarest at about 3 percent of parts, based on Kestrion's own year of scrap records. To trust a rate on a 3 percent class within about 10 points, you need something like 300 examples of that class alone, not 300 photos total."
Why this works
Naming where each number came from turns "trust me" into "here's exactly where you can check my work."
Stage 6
Give a range, then run the sanity check
Say it like this
"All in, across six defect types and three production lines, that lands somewhere between about 500 and 650 photos, not one exact figure. And yes, that's more than ten times the 40 we started with, which sounds like a lot until you remember 40 photos gave us maybe one or two chances to see the rarest defect at all."
Why this works
A range keeps the number honest, and naming the "that's a lot" reaction yourself gets ahead of the objection.
Stage 7
Name the direction and the AI-specific reasoning, then close
Say it like this
"The assumption that moves this most is how rare the rarest defect actually is, not the total defect count. A model's accuracy on paper is only as trustworthy as the number of real chances it had to get the rare stuff wrong. I'm not defending a bigger number for its own sake. I'm defending the one defect type that forty photos never gave a fair test."
Why this works
Restates the direct answer and lands on the AI-specific judgment: an aggregate accuracy number can hide a class with almost no real evidence behind it.
Let's learn
For six years, an inspector at Kestrion checked every machined part by hand under a loupe, about six minutes a part, roughly 140 parts a shift.
RivetEye photographs each part instead and sorts it in under four seconds, so an inspector reviews the flagged pile instead of checking every part from scratch.
The five letters, held up as one page. Own the numbers is the step a forty-photo sample skipped.
Here's the turn: Marisol judged RivetEye feasible off 40 example photos pulled from a vendor demo set, staged under good light, hitting 96 percent accuracy against that set. Six of Kestrion's defect types were in that set, but the rarest one, a casting void that shows up in about 3 percent of parts, appeared only twice. Two correct calls out of two tells you almost nothing. It felt like proof anyway.
What a trustworthy sample actually adds up to
Forty photos barely covered "see each defect once." The rare-class trust and the line-to-line spread were never built at all.
Forty photos never even reached the first branch of this tree.
At its worst, a turbine housing with an undetected casting void ships to a customer, and the flaw surfaces during their own incoming inspection instead of Kestrion's, turning a private catch into a public one and a batch recall.
Ninety-six percent was never a lie. It was a number with almost no evidence behind the one defect that mattered most.
The choice I would take back
Kestrion's own go, no-go rule for any inspection tool was "40 clean examples, 90 percent or better, ship it," set back when RivetEye only had to catch one obvious defect type. It made sense then. It stopped making sense once RivetEye covered six defect types with wildly different base rates, because the rule never asked how many times the rarest one had actually been seen.
What I would leave alone: for hairline cracks, showing up in nearly a third of scrap parts, the original 40-photo sample was genuinely plenty. That accuracy held up in production without a single surprise.
The lesson: a sample is not trustworthy because the total looks big. It is trustworthy because the thing you are most worried about missing has actually shown up enough times to be tested, not just hoped for.
Now here is the same thing as a story
The short version above is what you'd say defending a sample size in a planning review. Read this one for what it felt like the week a routine audit turned a comfortable number into an uncomfortable one.
Marisol could tell a defect classifier that was actually ready from one that just looked ready, most of the time. She had shipped two smaller vision tools before RivetEye, both fine, both boring in the best way.
Same word, sample, describing two very different collections of photos.
RivetEye's early weeks were good ones. The demo set's 96 percent held up on the shop floor for the common defects, hairline cracks and chatter marks, week after week. Inspectors started trusting the flagged pile and skimming past the rest, since checking every single one kept confirming it was fine.
Four things the 40-photo demo set never had a chance to prove.
Then Kestrion's quality director ran a surprise audit, pulling ten recently shipped parts at random. One had a casting void RivetEye had scored as clean.
Feasible got decided in week 3. Nobody asked again until week 26.
The real question was never whether RivetEye was 96 percent accurate. It was whether that number had ever been tested on a defect rare enough that two lucky catches could pass for proof.
What moves the required sample size most
The rarest defect's own rate moves the required sample more than everything else combined.
When the 40-photo sample was first approved, someone said, "it's well above our bar, let's ship it," and it sounded reasonable, since the bar itself had never been rewritten since the day it covered one defect, not six.
Rerun that same audit with a 600-photo sample built to cover all six defect types across every line: the casting void class alone had been tested against 40 real examples, not two, and its true catch rate, 71 percent, was known and flagged for a second check before RivetEye ever reached the shop floor.
What I'd tell myself, hearing about that one shipped part: 96 percent was never the number that mattered. The number that mattered was two, and nobody had asked what two examples could actually prove.
BOUND, the arithmetic behind a number that felt safe enoughNot a script against ever trusting a small sample. BOUND is what tells you exactly which class never got a fair test.
B
Break it down. State the equation out loud.
Required sample equals enough to see each defect type once, plus enough of the rarest type to trust its rate, plus enough spread across production lines.
This is the hardest step, and the one a "40 photos, 96 percent" answer always skips.
O
Own the numbers. State each one, and where it came from.
Six defect types and a 3 percent rarest-defect rate, both from a year of Kestrion's own scrap records. About 300 examples needed of that rarest class alone to trust its rate within 10 points.
Naming the source of each number is what makes the sample size verifiable, not just confident.
U
Use a range. A low and a high, not false precision.
All in, across six defect types and three lines, the real requirement lands between about 500 and 650 photos, not one exact figure.
A range survives the first "that seems like a lot" objection before it is even raised.
N
Nail the sanity check. Does it survive a smell test?
Needing more than ten times the original sample sounds excessive, until you remember the rarest defect had only ever been tested on two examples, which is not a rate, it is a coin flip with a nice number attached.
Naming the "that can't be right" reaction yourself gets there before an audit does.
D
Direction. Which assumption would change the answer most?
How rare the rarest defect actually is moves the required sample by 380 photos, more than every other factor combined. If that defect turned out to be twice as common, the whole requirement would shrink.
Naming what would change your mind is what separates a real estimate from a number defended out of habit.
The recap, one line per letter: break it down is coverage plus rare-rate trust plus line spread, own the numbers is stating where the 3 percent and 300 came from, use a range is 500 to 650 instead of one figure, nail the sanity check is explaining the ten-times jump before anyone questions it, and direction is naming the rarest defect's own rate as the number worth watching.
And if you want to be sure it really works, try it somewhere elseSame five letters, a municipal recycling line instead of a machine shop. The rare class that never got tested changes. The missing arithmetic doesn't.
Solveig Marrapodi runs product at Thistlewood Waste Services, where BinScope photographs items on a sorting belt and flags what should be pulled before it reaches the baler. A pilot judged BinScope feasible off 60 photos, mostly plastic bottles and cardboard, hitting 94 percent. Mapped onto BOUND: break it down is coverage of every material type plus enough of the rarest contaminant, a coffee cup lined with foil that looks like paper but jams the baler, plus spread across the two collection routes that feed the belt. Own the numbers means pulling from Thistlewood's own six months of jam reports, where the foil-lined cup shows up in about 2 percent of loads. Use a range covers the 250 to 400 photos of that one item needed to trust its catch rate. Nail the sanity check means explaining why a 94 percent pilot number still let three jams through in the first month. Direction names the one thing that would flip the plan: if the foil-lined cup turned out to be rarer than 2 percent, fewer examples would already be enough.
Checking the math before the go, no-go call is the step a 60-photo pilot skipped the first time.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "size the sample off your rarest case, not your total count," and stop.
Cost: there is no time to gather hundreds of examples before a decision is due. Say so honestly, and name which rare class is still untested rather than pretending the small sample covers it.
The model actually gets better, for real: if a new version of RivetEye claims a higher accuracy, that claim still needs the same rare-class test before anyone trusts it, since a better model can still be tested on too few hard examples.
Where people run it wrong.
They treat a big total photo count as proof, without checking how many of those covered the rare case that actually matters.
They trust a percentage with no visible equation behind it, as if a number alone were the same thing as evidence.
They set a sample-size rule once, for one easy case, and never revisit it as the tool takes on harder, rarer cases.
How to use it live. The moment someone asks how many examples are enough, ask yourself: what is the rarest thing this tool has to catch, and how many times has it actually been seen. Name that number, and the rest of the sample size follows on its own.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust flip: the team checked whether the sample needed to keep growing every week, then stopped once the number held steady for three weeks running, and treated 40 photos as final.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marisol Iyengar, the AI PM at Kestrion Precision Works, who judged RivetEye's defect scanner feasible off a 40-photo demo sample.
3 · THE HABIT
What did the team stop doing once the number held steady for three weeks?
Tap to flip
ANSWER
They stopped adding new examples to the sample and re-checking whether the accuracy number would move, treating a steady number as proof instead of a coincidence of small numbers.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Keep growing the sample and re-checking the number, versus trust the number and stop growing the sample. No in-between setting once the number looked steady.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Kestrion's standing rule of "40 clean examples, 90 percent or better, ship it," written when RivetEye only had one defect type to catch, never rewritten once it covered six.
6 · THE NUMBER
Fill in the blank: the rarest defect, casting voids, appeared only ___ times in the original 40-photo sample, out of a true rate of about ___ percent of parts.
Tap to flip
ANSWER
Twice, out of about 3 percent.
7 · THE REPLAY
Same audit, a 600-photo sample already built. What changes?
Tap to flip
ANSWER
The casting void class gets tested against 40 real examples instead of two, its true 71 percent catch rate is known in advance, and it gets flagged for a second check before RivetEye ever reaches the shop floor.
8 · CROSS PRODUCT TRANSFER
Section 4 runs this again for a different product. Which one, and what stays the same?
Tap to flip
ANSWER
Thistlewood Waste Services' BinScope. The same missing piece stays: a rare item, the foil-lined cup, that a small pilot sample never tested enough times.
Check yourself Score: 0 / 0
Multiple choice
1. Why did a 96 percent accuracy number fail to predict the shipped casting void?
A. RivetEye's camera hardware broke down on the shop floor.
B. The 96 percent was measured on a sample where the casting void defect had appeared only twice, too few times to trust a rate on it.
C. Inspectors stopped using RivetEye entirely after the first week.
D. The quality director changed the definition of a casting void.
Show hint
Look at "here's the turn" in the first section.
Show answer
B. Two correct calls out of two examples tells you almost nothing about the true catch rate for that defect type.
True or false
2. True or false: this answer argues that every feasibility spike needs at least 600 examples, no matter what it's checking for.
True
False
Show hint
Look at "what I would leave alone."
Show answer
False. For a common defect like hairline cracks, the original 40-photo sample was genuinely enough. The number needed depends on how rare the case is, not a fixed floor.
Fill in the blank
3. Fill in the blank: varying how ___ the rarest defect is swings the required sample size the most, by about 380 photos.
Show hint
Look at the horizontal bar chart, "what moves the required sample size most."
Show answer
Rare. The rarer the defect, the more total examples you need before its own rate is trustworthy, more than lines covered, confidence margin, or total defect types.
Short answer, where it wouldn't matter
4. Name a defect type in this same tool where a 40-photo sample really was enough, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Hairline cracks, which show up in nearly a third of scrap parts. A small sample already contains plenty of real examples of a defect that common.
Short answer, apply it yourself
5. Think of a claim you've seen backed by "we tested it and it works." What rare case might that test have barely covered?
Show hint
Think about what the test's total sample size actually contained, not just its headline result.
Show answer
Model answer: A spam filter tested on a week of email might see very few examples of a brand new scam pattern, so a strong overall accuracy number can still hide a near-total miss on that one rare pattern.
Short answer, work the number
6. If the casting void defect were actually twice as common, at 6 percent instead of 3 percent, would you need roughly twice as many total photos?
Show hint
Think about how many total photos it takes to see a class a given number of times, as its rate changes.
Show answer
Model answer: No, roughly half as many. If the defect shows up twice as often, you reach the same number of real examples of it from about half the total photos, so the requirement would shrink, not double.
Before you close the answer
Why this works
Tests whether you'll treat a headline accuracy number as proof on its own, or ask how many real chances it actually had to be wrong on the case that matters most.
Follow-up traps
"Isn't 650 photos just an arbitrary bigger number?" Response: no, it is built from three named parts, coverage, rare-class trust, and line spread, each with its own source, not picked to sound thorough.
"What if you can't afford to collect that many examples?" Response: then say so honestly and name exactly which class is still untested, rather than letting a small sample quietly stand in for a rare one it never covered.
If pressed
The 10-point confidence margin used here comes from a standard proportion estimate at a 95 percent confidence level, tightened to 5 points for the casting void class specifically once the audit made its cost visible, which roughly doubles the examples needed for that one class alone.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.