InterviewAdvancedShipping & Model Lifecycle / Pilot design and POC-to-production / #22
Design a pilot live for a scenario I describe.
The direct answer
Run the tool against 500 real incoming applications in shadow mode, not a curated test set, with an underwriter reading and confirming every flag before it can touch a quote. Set the pass bar by segment, not by one topline number, especially for shelter and mixed-breed dogs, because that's exactly where a wrong breed call attaches an exclusion that shouldn't apply. And keep the pilot away from actual pricing entirely; that's a separate decision for later.
Do this, in order
Run the tool against 500 real incoming applications in shadow mode, with an underwriter checking every flag before it touches a quote.Why: a pilot that never gets checked against real, messy submissions hasn't proven anything the live queue actually needs.
Set a separate pass bar for shelter and mixed-breed applications from day one, not folded into one overall average.Why: that's the exact segment where a wrong breed call attaches an exclusion that never should apply.
Never let a flag change a quote or a policy term on its own; require a human confirmation on every single one for the whole pilot.Why: a wrongful exclusion is expensive and slow to undo once it's sitting inside thousands of live policies.
Watch what underwriters do right after a segment-level miss surfaces, not just how many misses there were.Why: a bad run of misreads on one kind of dog can make a team stop trusting flags they'd have trusted before, which costs more review time than the tool was built to save.
Leave any change to actual premiums or breed-based pricing completely out of this pilot.Why: that's an actuarial decision with its own review, and folding it in here answers neither question cleanly.
Don't test only on the easy, well-documented applications because they're faster to run through.Why: a pilot proven only on clean cases hasn't proven anything about the messy real ones it has to survive.
How to answer this, stage by stage
Seven moves. Anchor it to the one segment number that almost got missed, not a general pitch for testing carefully.
1
Scope it to one concrete pilot
Say it like this
"So I'm Baptiste, the PM for PawScan at Fenwood Pet Assurance. It reads the photo submitted with a new dog-insurance application, calls the breed, and flags anything that looks like a pre-existing health issue. I'm not going to talk about AI pilots in general. I'll walk through the actual pilot I'd run: 500 real applications, in shadow mode, checked by segment before anything ships."
Why this works
A scoped example gives the interviewer something to picture and push back on.
2
Say the structure out loud
Say it like this
"Here's how I'll go: why you can't validate this with a demo, the actual pilot design, the number that almost hid the real problem, what I'd change because of it, and what this pilot is not allowed to touch."
Why this works
Two seconds of structure tells the interviewer you have a plan, so they follow instead of guessing where you're going.
3
Reframe what a pilot has to prove
Say it like this
"Most people think a pilot just proves the model basically works. One topline accuracy number, a green light, done. But a single average can hide exactly the failure that matters, if the mistakes aren't spread evenly across every kind of photo. The only way to know is to run it against real applications and look at where the misses actually land, not just how many there are."
Why this works
This is the actual insight being tested. Skip it and you're just saying "run a pilot," which nobody disagrees with and nobody learns from.
4
Give the anchor
Say it like this
"So here's the pilot. Five hundred real incoming applications, over three weeks, in shadow mode: PawScan calls the breed and flags anything that looks like a pre-existing condition, but nothing touches an actual quote until an underwriter reads and confirms it. And the pass bar isn't one number. It's ninety-two percent agreement, checked separately for shelter and mixed-breed dogs against everyone else, because that's exactly the segment where a wrong breed call attaches an exclusion that shouldn't be there."
Why this works
Naming the specific mechanism, out loud, is what separates a real pilot design from a vague "test it thoroughly" gesture.
5
Walk what the pilot actually found
Say it like this
"Here's what the first pass found. Overall, PawScan agreed with the underwriter on 456 of 500, about 91 percent, a hair under our bar. Close enough that another week of tuning looks tempting. But cut by segment, the 118 shelter and mixed-breed applications were only agreeing 74 percent of the time. PawScan called 31 of those dogs a purebred large breed and auto-suggested an exclusion that never should have applied to a mutt."
Why this works
A real number, with the real mechanism attached, does more work than "the model wasn't perfect" ever will.
6
Say what changed because of it
Say it like this
"So I wouldn't ship on the 91 percent. I'd add a second calibration pass for anything flagged shelter or mixed on the intake form, and re-test before the pilot graduates. On the next 150 shelter applications we ran after that fix, only 7 got wrongly tagged, and agreement on that segment went from 74 percent to 95."
Why this works
It shows the pilot actually changing a decision, not just producing a report nobody acts on.
7
Close on what's kept out and the number
Say it like this
"One thing this pilot never touches: whether breed should change the price at all. That's an actuarial call with its own review, not something a photo-flagging pilot decides as a side effect. So: 500 real applications, shadow mode, checked by segment. 91 percent overall almost looked like a pass. The segment that mattered was sitting at 74, and that's the number that would have shipped 31 wrong exclusions if nobody had looked twice."
Why this works
Interviewers remember the last line most, and this one hands them something they can check, not just a mood.
Let's learn
What happens the first time a computer looks at a photo of a scruffy shelter mutt and decides, with total confidence, that it's a purebred?
Fenwood Pet Assurance built PawScan: read the photo submitted with a new dog-insurance application, call the breed, flag anything that looks like a pre-existing health issue, so an underwriter confirms it instead of eyeballing everything cold.
Today, without PawScan
Before PawScan, an underwriter opened each new application, looked at the photo next to the owner's typed breed, and priced the policy by eye, noting anything that looked off. About 6 minutes each. On a normal week, with around 60 new applications, that's fine.
Knowledge spark: what's an exclusion?
A part of the policy that says "we won't pay for this." Insurers often set one exclusion per breed's known risk, like hip trouble in big dogs, decided the day the policy starts.
Baptiste Halsbury, PawScan's PM, didn't push for a rollout plan first. He asked for a pilot: run it against 500 real incoming applications, in shadow mode, so PawScan's calls sit next to what the underwriter decides on their own, but never touch a quote until a person confirms them.
The first read looked close to done. Of 500 applications, PawScan's breed call agreed with the underwriter on 456. That's 91 percent, a hair under the 92 percent bar the team had set. Another week of tuning looked like it would close the gap.
Then someone on the growth team, prepping an unrelated slide about where applications come from, asked for the same agreement number cut by channel: shelter and rescue against breeder-documented purebreds. Of the 118 shelter and mixed-breed applications in the batch, PawScan had called 31 of them a purebred large breed, with a breed-specific exclusion auto-suggested alongside it. Agreement on that one segment: 74 percent.
PawScan's breed-call agreement, three points in the pilot
the topline numberthe segment that carried the riskthe same segment, after the fix
The topline number almost cleared the bar. The segment carrying the risk was seventeen points under it, until a second calibration pass caught up.
The topline number said 91 percent and almost passed. The 31 wrong exclusions were all sitting in one segment nobody had pulled out yet.
At its worst, this doesn't stay small. Fenwood writes about 2,250 shelter and mixed-breed policies a year. At a 26 percent wrong-tag rate, that's roughly 590 policies a year carrying an exclusion that never should have applied, each one a hip claim the company either fights, quietly pays out on a technicality, or wrongly denies to a customer who did nothing wrong. And the actual problem, one breed call trusted a little too far, would still be sitting there unfixed.
The decision that mattered
Check the pilot's pass bar by segment before any go/no-go decision, not only in aggregate. A topline number that clears the bar can still hide a segment failing badly enough to matter.
The choice I would take back. The pilot's original plan was to report one number to the go/no-go review: overall agreement against the 92 percent bar. That's the number leadership asked for, and it's the simplest one to put on a slide. It made sense assuming errors would be spread evenly across every kind of application. It stopped making sense the moment the segment cut showed the errors weren't spread at all. They landed almost entirely on one kind of dog.
What I would leave alone. The applications where the owner's typed breed and PawScan's call already agreed outright, the clear purebreds with documented pedigrees, never needed extra scrutiny. A quick confirm already worked fine for those, in every week of the pilot.
The lesson. A pilot that only reports its average hasn't told you the thing you actually need to know. Grade it by its worst segment, not by the number that fits on a slide.
Now here is the same thing as a story
The short version is above. Read this one when you've got a few minutes, for why it mattered.
The intake desk at Fenwood gets about 60 new dog applications a week, most weeks. An underwriter opens each one, looks at the photo, checks it against the breed the owner typed in, and prices the policy. About 6 minutes a piece, and on a normal week that's a fine way to spend a morning.
Baptiste Halsbury had run product for underwriting tools at Fenwood for two years, long enough to know most applications are unremarkable: a labrador is a labrador, the photo agrees with the form, done. He also knew the ones that weren't so simple. Shelter dogs. Rescues. A dog whose paperwork says "mixed breed, best guess" because nobody who found her knew for certain.
Leadership wanted a rollout plan for PawScan. Baptiste asked for something smaller first: a pilot, run against real applications, checked by real underwriters, before anyone talked about turning it on.
The design was simple on purpose. Five hundred real incoming applications, over three weeks. PawScan would call the breed and flag anything in the photo that looked like a pre-existing condition. But it would run in shadow mode: its calls sat next to what the underwriter decided on their own, and nothing PawScan said would touch an actual quote until a person read it and confirmed it.
The anchor: checked before it ever prices
Three weeks in, the first read looked close to done. PawScan's breed call agreed with the underwriter on 456 of the 500 applications. Ninety-one percent. The bar was 92. Close enough that the plan, quietly, was another week of tuning and then a green light.
Then, on a Wednesday, a data analyst on the growth team messaged Baptiste. She wasn't looking for a problem. She was building a slide about where new applications came from, shelters against breeders against private sellers, and asked if he could send her the same agreement number, cut the same way.
He hadn't cut it that way. He did it that afternoon, mostly out of politeness.
Of the 500 applications, 118 had come in through shelters and rescues, flagged "mixed breed" on the intake form. PawScan had called 31 of those dogs a purebred large breed, high confidence, exclusion auto-suggested right alongside the call. One of them was a lab-shepherd mix named Otis, adopted through a county shelter, that PawScan read as a purebred German Shepherd and flagged for the standard hip-dysplasia exclusion, a condition Otis's actual mixed ancestry made far less certain to matter.
Agreement on that one segment: 74 percent.
The day it's wrong, and the anchor still catches it
The topline number said 91 percent and almost passed. The 31 wrong exclusions were all sitting in one segment nobody had pulled out yet.
Because here's what the 91 percent was hiding. It wasn't 9 percent of mistakes spread evenly across every kind of dog. It was a coin flip on one specific kind of application, the exact kind where getting it wrong meant a real customer's policy carried a real exclusion it never should have. And nobody would have known, because the number the go/no-go review was built to look at said pass.
So here is the decision Baptiste took back.
The original pilot plan reported one number: overall agreement against the bar. That made sense when the team assumed a wrong call was as likely on a golden retriever with a clean pedigree as on a shelter mutt. It stopped making sense the second real data showed the mistakes clustering on exactly the applications where a wrong exclusion does the most damage.
Baptiste didn't scrap PawScan. He added a second calibration pass, specifically for anything the intake form flagged shelter or mixed, and re-ran the segment before letting the pilot graduate. Of the next 150 shelter and mixed-breed applications tested after that fix, only 7 got wrongly tagged a purebred. Agreement on that segment went from 74 percent to 95.
And the thing I'd want to tell myself, back when one clean topline number looked like enough to report to a review committee: an average only tells you the size of a problem. It never tells you whose problem it is.
SPARK, tested against a dog named Otis
This question sounds like it wants a rollout plan, when do we turn this on for everyone, or a metrics answer, how do we know PawScan is working. It's really asking for one concrete pilot design that has to survive contact with real, messy applications, so SPARK fits. A question asking how Baptiste would know PawScan was still working two years into rollout would reach for LEAD instead.
S, situation. Today, without PawScan, an underwriter opens each new application, checks the submitted photo against the owner's typed breed by eye, and prices the policy from that. No tool has checked its own calls against real shelter and mixed-breed dogs yet.
P, payoff. Not "underwrite faster." The habit worth building this pilot: catch a systematic misread, a shelter mix read as a purebred with the wrong exclusion attached, inside a 500-application shadow run, before it's sitting in thousands of live policies waiting for a claim to expose it.
A, anchor. Five hundred real applications, three weeks, shadow mode. PawScan's calls never touch a quote until an underwriter confirms them, and the pass bar is 92 percent agreement checked by segment, shelter and mixed-breed against everyone else, from day one.
R, risk. Too narrow, and the 500 come mostly from clean, well-lit photos with confident breed guesses, so it looks great and then breaks the first time it hits a real shelter phone photo. Too broad, and PawScan gets wired straight into the instant online quote flow, auto-attaching exclusions before the segment number is proven, so a lingering miss ships into thousands of live policies before anyone notices.
K, keep out. Whether breed should change the base premium at all, and by how much, is an actuarial pricing decision. This pilot doesn't touch a single price on a single policy while it's running.
What we left for later, kept visibly separate from day one
Why the anchor survives the risk
Check it against the near miss. Does checking by segment still catch the danger even when the topline number looks fine? Yes, that's the whole point of splitting it out before the go/no-go review, not after. Does it avoid the over-build trap? Yes, because K keeps pricing entirely off this pilot's job, so nobody's tempted to fold two decisions into one green light.
And if you want to be sure it really works, try it somewhere else
A city building department runs on a completely different clock, but the same gap between what a topline number says and what one segment is actually carrying shows up in permit review.
S. Yolande Fentress runs the permit-review pilot for Thornstead's building department. Today, without a tool, an inspector opens each application's submitted photos by eye and decides whether to schedule a site visit or approve the paperwork step from the photos alone. P. The habit worth building: catch a systematic false-flag pattern, the tool reading ordinary aged wiring on pre-1960s homes as an active fire hazard, before it triggers hundreds of unnecessary site visits city-wide, not after contractors start complaining publicly. A. Same shape, a different desk. Yolande runs the flagging tool against 300 real submitted applications over a defined window, an inspector reviewing every flag before any visit gets scheduled, with the pass bar checked separately for homes built before 1960 against newer construction. R. Test it mostly on new-construction photos, clean and code-compliant by default, and it looks perfect, then floods the queue with false call-outs on the city's oldest housing stock the day it goes live. Wire it straight into auto-scheduled site visits before the by-age number is proven, and homeowners in older homes start losing weeks to visits that never needed to happen. K. Whether a flagged home should face a higher permit fee or faster inspection priority is a policy decision for the department, not something a photo-flagging pilot decides on its own.
False-flag rate on pre-1960s homes, week by week during the pilot
before age-aware calibrationcalibration rolling outafter it's tuned to housing age
A citywide average would have called week one "fine." Pre-1960s homes were flagged wrong more than a third of the time until the tool learned to look at a building's age, not just its wiring.
Swap the trigger and it still runs
Speed: even if PawScan scored every application instantly instead of in a couple seconds, that wouldn't fix a shelter mix reading as a purebred. Speed and accuracy are different jobs.
Cost: if running PawScan cost Fenwood nothing at all, that still wouldn't tell you which segment the tool actually misses. You still need real shelter photos, not free compute.
The model gets better: if PawScan's overall breed accuracy climbed to 99 percent, that still wouldn't guarantee the mistakes that remain aren't all clustered on the same segment again.
Where people run it wrong
Testing only on applications that already look easy, clean photos, confident breeder paperwork, and calling a high pass rate proof it's ready for messy real submissions.
Letting the tool auto-attach exclusions the moment it clears the aggregate bar, before an underwriter has read a single flag next to the real photo.
Trying to also settle the pricing question, should mixed breeds cost less to insure, inside the same pilot, so the one real question, is the flag accurate, never gets a clean answer.
How to use it live
If you're asked this cold, pick a real segment your own product already worries about, then ask what a single topline pass rate would hide about that exact segment before you promise a pilot design. That question, asked of yourself out loud, finds the real pass bar faster than trying to write the perfect number up front.
Flashcards (click a card to flip it)
1 · THE SITUATION
What's the situation, before this pilot existed?
Tap to flip
ANSWER
Fenwood's underwriters checked each new dog-insurance photo against the owner's typed breed by hand, about 6 minutes each, with no tool having checked its calls against real shelter and mixed-breed dogs yet.
2 · THE PAYOFF
What's the real habit this pilot is trying to build?
Tap to flip
ANSWER
Catching a systematic misread, a shelter mix tagged as a purebred with the wrong exclusion attached, inside a 500-application shadow run, before it's sitting in thousands of live policies.
3 · THE ANCHOR
What's the one pilot design decision everything else hangs on?
Tap to flip
ANSWER
500 real applications, shadow mode, underwriter confirms every flag before it touches a quote, and the pass bar is checked by segment, shelter/mixed against everyone else, not just in aggregate.
4 · THE RISK
What breaks if the pilot is too narrow, or too broad a commitment?
Tap to flip
ANSWER
Too narrow, tested only on clean easy applications, it passes and then breaks on the first messy shelter photo. Too broad, wired straight into instant pricing before the segment number is proven, and a miss ships into thousands of live policies.
5 · THE PROOF
What did the segment cut find that the topline number never would have?
Tap to flip
ANSWER
Overall agreement was 91%, just under the bar. The 118 shelter/mixed-breed applications were only agreeing 74% of the time, with 31 wrongly tagged a purebred and handed an exclusion that shouldn't apply.
6 · THE NUMBER
___ of ___ shelter applications got the wrong exclusion, and agreement on that segment was ___ percent before the fix, ___ after.
Tap to flip
ANSWER
31 of 118. 74 percent before, 95 percent after.
7 · THE REPLAY
Same pilot, segment-aware calibration added mid-run. What changes?
Tap to flip
ANSWER
Of the next 150 shelter applications tested, only 7 got wrongly tagged instead of the earlier rate, and agreement on that segment rose to 95%, with every remaining miss caught before it shipped.
8 · CROSS-PRODUCT
Section 4 runs SPARK again on a different product. Which one, and what does its anchor test?
Tap to flip
ANSWER
Thornstead's permit-photo code-violation flagging pilot. Its anchor tests whether the flag rate is accurate by housing age, not just overall, before any site visit gets auto-scheduled.
Check yourself Score: 0 / 0
True or false
1. True or false: the 31 wrongly tagged applications were the real cost of running the pilot against real shelter dogs.
True
False
Show hint
Think about what the pilot was actually for, and what would have happened if those 31 had shipped instead of getting caught.
Show answer
False. The real cost would have been letting those 31 wrongful exclusions ship into live policies. Catching them in shadow mode, before a single one touched a real quote, is exactly what the pilot was built to do.
Fill in the blank
2. ___ of ___ shelter and mixed-breed applications got wrongly tagged a purebred, and agreement on that segment went from ___ percent to ___ percent after the fix.
Show hint
This number shows up twice, once in the story, once in the chart.
Show answer
31 of 118; 74 percent to 95 percent. Fewer than a third of shelter applications missed, but the cost of those misses would have landed on real policies, not just a spreadsheet.
Multiple choice
3. Which pilot design matches the anchor this answer argues for?
A. A single 92 percent accuracy bar, checked in aggregate, before graduating the pilot.
B. Shadow mode against 500 real applications, an underwriter confirms every flag, and the pass bar is checked by segment.
C. Auto-attaching exclusions the moment PawScan crosses 90 percent confidence.
D. A spec document with actuarial sign-off, reviewed before any code gets written.
Show hint
The anchor needs both a human check before anything ships, and a pass bar that can't hide behind one average.
Show answer
B. A checks only the aggregate, which is exactly what nearly hid the real problem. C removes the human check that caught the near miss. D is still description, not a real test.
Short answer
4. What old habit does this pilot design take back, and why did it make sense when it was first set up?
Show hint
Think about why reporting one aggregate number sounded fine, back before anyone had cut it by segment.
Show answer
Model answer: The original pilot plan was to report one aggregate agreement number to the go/no-go review, because that's the number leadership asked for and it's the simplest one to put on a slide. That made sense assuming errors would be spread evenly. It stopped making sense once the segment cut showed the errors clustering entirely on shelter and mixed-breed dogs.
Short answer, apply it yourself
5. Pick a product you use yourself, or one your team is building. What's one segment where a topline pass rate could be hiding a much worse number underneath?
Show hint
Look for the group of users or cases that's smallest, or least like the "typical" case a demo would use.
Show answer
Model answer: "Our resume screener says it's 90 percent accurate overall. But almost none of the training data came from career-changers with non-traditional titles. Cutting that 90 percent by 'traditional title' versus 'career-changer' would probably show the real number is a lot worse for exactly the candidates most likely to get screened out unfairly."
Multiple choice
6. Based on this answer's own numbers, if the shelter segment had been agreeing at 90 percent instead of 74, would checking the pass bar by segment before graduating the pilot still have been the right call?
A. Yes, because the point of checking by segment is knowing the real number for the segment that carries the risk, whether it turns out high or low.
B. No, at 90 percent the team should have skipped the segment check and shipped on the aggregate number.
C. No, a higher segment number means the exclusion logic was never a real problem in the first place.
D. Yes, but only because a lower number would have looked worse in front of leadership.
Show hint
Compare what a higher segment number changes about the need to know it, against what it changes about whether a human still needs to confirm each flag.
Show answer
A. A higher segment number doesn't remove the need to know it, or the need for an underwriter to catch what remains. It would still have been the wrong pilot to find that out by looking only at the average.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.