CaseAdvancedShipping & Model Lifecycle / Prototyping with LLMs and rapid POCs / #15
How would you use a prototype to negotiate scope with engineering?
The direct answer
Before agreeing to any scope, run the prototype live on real, messy warranty photos, not the clean demo shot, and let what it gets right and wrong set the order of the negotiation. Prove first that it can name the damage type at all. Then settle the one claim hardest to take back, whether it can be trusted to decide liability on its own, before engineering estimates a single sprint against it.
The ranking, by what breaks first if skipped
Prove the tool can name the damage at all, on real customer photos.Why: dependency. Every other scope claim in the room assumes this part already works.
Pin down which damage calls it actually gets wrong, not just how often.Why: reversibility. The hinge-crack and water-seam calls decide who pays, and that claim is the hardest to undo once a sprint is built around it.
Decide whether the tool auto-decides or waits for a person, before it's a line of code.Why: this is the one scope line engineering will build a whole pipeline around, not a setting you flip later.
Run it live in the room on messy real photos, not a slide of accuracy numbers.Why: cheap to check now, and it settles arguments a slide never will.
Put the nice-to-have polish, an overlay graphic, a summary PDF, at the bottom of the list.Why: none of it changes whether the core scope is even right.
Re-test on new real claims whenever the model underneath changes.Why: a failure shape found this month might not hold after the next update.
How to answer this, stage by stage
Eight moves. The trap in this question is answering it with a list of features for the sprint, when it's really asking which unproven claim gets expensive to reverse the moment engineering attaches an estimate to it.
1
Ground it in one real product
Say it like this
"Let me make this real. Say Barrowmere Appliances is building Dentscope, a tool that reads the photo a customer attaches to a warranty claim and works out what's damaged and whether it's on Barrowmere. Anneke Sorbus is the PM scoping it, Emory Rask leads the engineering team she's negotiating with, and Nolan Trask is the senior claims associate who'd actually use it. I'll answer against that."
Why this works
Grounds an abstract scope question in one real product, so the ranking that follows isn't hypothetical.
2
Name your method before you use it
Say it like this
"I'd use ORDER here. Rank the open questions the prototype has to answer by what's hardest to undo once engineering commits an estimate to it, not just list every feature that could go in the sprint."
Why this works
Signals a method up front, so the answer reads as a plan, not a list made up on the spot.
3
Say what the question is actually testing
Say it like this
"This isn't really asking me what to build first. It's asking whether I know which unproven claim gets expensive to reverse the moment engineering attaches a sprint estimate to it."
Why this works
Separates the real judgment call from a surface reading that treats the question as a build order.
4
Give the ranked answer straight
Say it like this
"Prove the damage call works at all, first, on real photos. Then find the one failure shape that decides who pays. Then run it live in the room. The polish, an overlay graphic, a summary PDF, goes dead last."
Why this works
This is deliverable 0, said out loud, in the order that actually matters.
5
Show what has to be true before anything else
Say it like this
"None of this matters until Dentscope can tell a cracked door from a scorched one on a photo a customer actually took. Every other scope line, the auto-approve limit, the appeal flow, quietly assumes that part already works."
Why this works
Shows the order isn't arbitrary. One thing genuinely has to be true before the next is worth negotiating.
6
Name the claim that's hardest to walk back
Say it like this
"The hardest thing to walk back is whether Dentscope decides a claim on its own or waits for Nolan. Say yes to auto-decide now, and Emory's team builds a whole pipeline around it. Undoing that later means redoing the pipeline, not flipping a setting."
Why this works
Names the one scope claim that turns an estimate into a commitment nobody can casually reverse.
7
Back it with the live evidence test
Say it like this
"Here's what forty real claims showed. On the twenty five obvious ones, a shattered door, a bent panel, Dentscope matched Nolan's call twenty three times. On the fifteen ambiguous ones, a hinge crack, a water stain near a seam, it matched him six times. That gap is what decides whether Dentscope can decide anything on its own near a hinge."
Why this works
A cheap test plus a real number makes the order defensible instead of just tidy.
8
Close on the rule, not the checklist
Say it like this
"So: can it call the damage at all, first, since everything else assumes it can. Then how it's wrong on the cases that decide who pays, since that's the claim hardest to take back. Then how it runs live in the room. The polish goes last, because none of it changes whether the scope is even right."
Why this works
Ends on the literal ranking the question asked for, defended instead of just listed.
Let's learn
Dentscope is a tool that reads the photo a customer attaches to a warranty claim, a cracked dishwasher door, a scorched cooktop, and works out what's damaged and whether it looks like something Barrowmere Appliances has to cover.
Every week, before Dentscope, Barrowmere's claims team pulls about 850 warranty photos and looks at each one by hand. Nolan Trask, who has worked the warranty desk for nine years, spends about 7 minutes on a claim: what broke, does it look like the factory's fault, how bad. That is close to 100 hours a week, spent the same way on the easy claims and the hard ones.
Knowledge spark: what makes a claim "covered"?
A claim is covered when the damage looks like it came from how the appliance was built, a weld that failed, a hinge that was never right. It is not covered when the damage looks like shipping, installation, or normal wear. The photo is often the only evidence anyone has.
Anneke Sorbus, the PM building Dentscope, put together a rough version in a weekend, and tested it on one photo she was proud of: a dishwasher door with a dent the size of a fist, an obvious factory defect. It caught it every time. She could have let engineering estimate the whole build off that one photo.
The extra mistakes were not the problem. Say that plainly, because it is the part most people skip past. What decided the whole shape of the negotiation was not how often Dentscope got a call wrong. It was which calls.
Testing whether Dentscope can name the damage at all comes before testing how it fails. Neither one comes before engineering gets an estimate.
On the 25 obvious claims Anneke eventually tested, a shattered oven door, a bent panel, Dentscope matched Nolan's call 23 times. On the 15 ambiguous ones, the hinge cracks and water seams that decide who pays, it matched him only 6 times.
How often Dentscope matched Nolan's call, obvious damage vs. the ambiguous cases that decide who pays
Obvious: a shattered door, a bent panelAmbiguous: a hinge crack, a water seam
Six of fifteen matched is not a rare miss. It's the shape of the mistake the scope has to plan for, not the one the demo photo ever showed.
We were not negotiating whether Dentscope makes mistakes. We were negotiating whether the scope would tell Nolan to check.
Here is what that costs at its worst. Scope Dentscope to auto-decide every claim off the clean-photo confidence, and a wrongly auto-approved factory defect is a check mailed out with nobody's eyes on the file. A wrongly auto-denied real defect is the kind of complaint that lands in a regulator's inbox, not a support ticket.
The choice I would take back
Early on, the plan was to let Emory's team estimate the whole build off the one dent photo Anneke kept using. I would swap that photo for the messiest real claim in the queue before anyone attaches a number of weeks to anything, because the clean photo only ever proves the easy 92 percent.
What I would leave alone. Whether Dentscope draws a red circle around the damage or just names it in text is not worth a minute of the negotiation yet. Nolan reads either one in about three seconds. Spend the negotiating time on the calls that decide who pays, not on how the answer looks.
The lesson. A model that gets the obvious dents right only proves it can do the easy part. The scope has to be built around the hinge crack, because that is the case where a wrong call costs real money and nobody notices for months.
Now here is the same thing as a story
The short version is above. Keep reading if you want to feel why the one photo Anneke kept using was the wrong one to scope a sprint around.
Anneke Sorbus had scoped four launches at Barrowmere before Dentscope, and what she was good at, from her first year in the job, was turning a vague ask into a spec engineering could build without three rounds of clarifying questions.
Dentscope's first real version could do exactly one trick, and it did it well: read a clean photo of a dented dishwasher door and call it a factory defect, every time. Anneke built it herself over a weekend. She ran it a hundred times against the same photo. It never missed. She carried that photo into every meeting she could find. Leadership liked it. Emory Rask, the engineering lead, watched it call the same door right for the fourth week running and started sketching a sprint on the whiteboard behind her: wired straight into the claims queue, auto-approve, auto-deny, no person needed under a set dollar amount.
For a while, that one photo was the whole conversation. Every scoping meeting opened with it. Nobody asked for a second photo, because the first one kept working.
Then, in the fourth meeting, Emory said something that wasn't really a jab. "We've built the whole plan around one photo you took yourself. What happens on the ones that actually confuse your team?" He wasn't trying to slow her down. He'd just noticed nobody had ever shown him one.
Anneke didn't have an answer. Nobody had run Dentscope against a claim that wasn't hers.
One of these you can move by a week and nothing changes. The other one, once it ships, has already paid or denied someone.
So instead of sending engineering the estimate that week, the way the plan called for, Anneke pulled forty real claims from the queue, already decided by a human, and ran the rough model against every one, the ugly ones included. It took an afternoon, and Nolan's help pulling the files.
Twenty five were obvious, the kind Nolan calls open and shut: a shattered oven door, a panel bent enough to see from across the room. Dentscope matched his call on twenty three of those. Fifteen were the ones Nolan actually slows down for: a hairline crack near a hinge, a water stain creeping along a seam, the kind that could be a factory defect or could be whoever installed it. Dentscope matched his call on six of those fifteen.
It was never really about the ninety two percent it got right. It was about which claims those fifteen ambiguous photos stand in for, and who pays when the model guesses wrong on one.
A month earlier, in the meeting where the team picked which photo to build the demo around, someone had floated pulling a real, messy claim instead. The clean dent photo already worked, and the messy ones felt like something to worry about later, once the model was further along. Nobody wrote down that this one photo would still be the only thing anyone had tested by the time engineering estimated a sprint against it.
I would go back and pick the ugly claim. Not instead of the clean one, first. The dented door proves the model can do the job at all. The hinge crack proves you know how it fails.
With that change, here's the replay. Same forty claims, same afternoon, but now it happens before the scope gets written, not after Emory's team has already built a pipeline around no person needed. Anneke scopes sprint one around the twenty five obvious calls, still glanced at by Nolan for now, and puts the fifteen ambiguous ones into a second, smaller sprint of their own, get the crack-versus-seam call right first, then talk about removing the person. Three weeks later, on a real day of claims, Nolan spends about eleven minutes on the flagged ambiguous photos, instead of seven minutes on every one of that day's roughly 170 claims.
What I'd tell myself, back in that fourth meeting: build the demo around the claim you're least sure about. The scope only has to be honest about that one.
ORDER, before the sprint gets an estimate
BOUND would fit if the question were sizing how many engineers Dentscope needs. This sits earlier than that, which open question the prototype has to answer before a sprint can even be scoped. That's ORDER's job.
O, outcome. Every scope claim in this negotiation protects one thing: that Emory's team estimates against what Dentscope actually does, not against the one photo Anneke is proud of.
R, reversibility. The hardest thing to walk back is whether Dentscope decides a claim on its own or waits for Nolan. Say yes to auto-decide now, and engineering builds a whole pipeline around it. Undo that later and you're not flipping a setting, you're redoing the pipeline.
D, dependency. Nothing else is worth scoping until Dentscope can name the damage at all, on a photo a customer actually took. Every later scope line, the auto-approve limit, the appeal flow, assumes the answer is yes.
E, evidence. Cheap to check first: an afternoon with forty real claims someone already decided, run live in the room instead of a slide of accuracy numbers Emory has to take on faith.
R, rank. Can it call the damage at all, first, since everything else assumes it can. Then the shape of the failure that decides who pays, hinge crack versus factory dent, since that's the claim hardest to reverse. Then how it runs live in front of the team, since a number on a slide and a model choking in the room in front of Emory are two different conversations. The overlay graphic and the claim summary PDF go dead last, because none of it changes whether the first three answers come back right.
The three-week window, and which parts of it were forced versus chosen
Forced by dependencyA genuine choice
Only the middle week was actually up for debate. The order around it was never a negotiation, it was arithmetic.
The check that keeps this ranking honest
Swap the outcome and the order should move. If a wrong hinge-crack call only ever cost a customer a follow-up phone call, the failure test could rank behind the live demo test. It ranks second here because a wrongly auto-denied real defect becomes a regulator complaint, not a phone call.
Same rank, a rental counter instead of a warranty desk
Haldane Tool Rental built Wearcheck: a tool that reads a photo of a returned drill or saw and decides whether the mark on it is normal wear or damage that should cost the renter a repair fee.
O. Every version of Wearcheck's test order protects one thing: that a damage charge reflects what actually happened to the tool, not a guess from a photo that looks like normal wear.
R. A wrongly charged damage fee costs a renter's trust once it's on their card. A wrongly waived one costs Haldane the repair bill. Neither one is a click to undo once the charge has gone through.
D. None of it matters until Wearcheck can tell damage from wear at all, on a tool photographed in a parking lot, not a lab shot.
E. Cheap to check: forty real returns Haldane's counter staff already priced last month, run against the rough model in an afternoon.
R. Same order: prove it can spot damage at all, first. Then learn its failure shape, whether it can tell a normal scuff from a crack that means the motor housing is compromised, the case that decides who pays for a two-hundred-dollar repair. Then test how it runs live at the counter with a renter standing there. The photo gallery for insurance disputes goes last.
Swap the trigger and it still runs
Barrowmere doubles its claim volume during a product recall. The order doesn't move. The failure-shape test matters more, not less, since more claims mean more chances for a hinge crack to slip through unflagged.
Dentscope's underlying model gets noticeably better at reading hairline cracks. Doesn't reorder either. A better model still needs Nolan to confirm the one time it's wrong, or nobody can tell that time from the other twenty-nine.
Barrowmere starts offering Dentscope to a partner retailer's warranty desk. Doesn't reorder. The rule protects the same thing regardless of who's running it.
Where people run it wrong
Treating one clean demo photo as proof the whole model works, when a demo only ever proves the part someone was brave enough to test.
Scoping the live-demo test and the polish features first, because they're easier to schedule than pulling forty real, already-decided claims.
Testing once before the sprint and never re-running the same real claims after the model underneath quietly changes.
How to use it live
Say the outcome out loud before naming a single scope line. "Every test we run before this sprint protects one thing, that engineering estimates against what the model actually does." Then rank from there. Naming the outcome first turns a scope argument into something you can defend line by line.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits ranking what a prototype has to prove before scope talk starts, and why not BOUND?
Tap to flip
ANSWER
ORDER, for ranking which open question is hardest to undo if you get the order wrong. BOUND is for sizing something once you already know what to build, not for ranking which unproven claim to test first.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Anneke Sorbus, the product manager building Dentscope at Barrowmere Appliances, negotiating scope with engineering lead Emory Rask.
3 · THE HABIT
What did Anneke stop doing once the first clean demo kept working?
Tap to flip
ANSWER
She stopped pulling new photos to test with. Every scoping meeting ran off the same one dented dishwasher door.
4 · THE DEPENDENCY
What has to be tested before any scope line means anything?
Tap to flip
ANSWER
Whether Dentscope can name the damage at all, on a real customer photo, not a studio shot. Every later scope claim assumes the answer is yes.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Building the whole scope estimate around one clean demo photo instead of a messy real claim. It made sense because the photo proved the model could do the job at all, and nobody had planned to test anything else before engineering estimated the sprint.
6 · THE NUMBER
Dentscope matched Nolan's call on ___ of 25 obvious claims, and ___ of 15 ambiguous ones.
Tap to flip
ANSWER
23 of 25, and 6 of 15. That gap is what decides whether Dentscope can be trusted to decide anything on its own near a hinge.
7 · THE REPLAY
Same forty claims, tested before the scope gets written instead of after. What changes?
Tap to flip
ANSWER
Sprint one covers only the twenty-five obvious calls, still glanced at by Nolan. The fifteen ambiguous ones get their own smaller sprint. Three weeks later, Nolan spends about eleven minutes on flagged photos instead of seven minutes on every claim.
8 · THE TRANSFER
Section 4 runs ORDER again on a different product. Which one, and what plays the role of the hinge-crack gap there?
Tap to flip
ANSWER
Haldane Tool Rental's Wearcheck. The equivalent gap is telling a normal scuff from a crack that means the motor housing is compromised, not just spotting damage at all.
Check yourself Score: 0 / 0
Fill in the blank
1. Dentscope matched Nolan's call on ______ of 25 obvious claims, and ______ of 15 ambiguous ones.
Show hint
It's the number the whole ranking argument turns on.
Show answer
23 of 25, and 6 of 15. That gap is why the scope couldn't treat every claim the same way.
Multiple choice
2. Which question does this answer say has to be tested first, before anything else about Dentscope?
A. Whether the claim summary PDF looks clean
B. Whether the model can name the damage at all, on a real customer photo
C. How long a big batch of claims takes to process
D. Whether Nolan likes the interface
Show hint
Ask what every other scope claim quietly assumes is already true.
Show answer
B. Every later scope line assumes the core task already works. Test that first, or the rest is a plan for a feature that might not exist.
True or false
3. True or false: since Dentscope rarely misses an obvious claim, the gap between obvious and ambiguous accuracy didn't need its own line in the scope negotiation. Say why.
True
False
Show hint
Think about what happens when something uncommon also can't be undone. Now check how uncommon 6 of 15 actually is.
Show answer
False. Six of fifteen matched is not rare at all, and a wrongly auto-decided ambiguous claim, once it's paid or denied, isn't something you can quietly take back. Common and unrecoverable together is exactly what needs its own written rule.
Multiple choice
4. What does the reversibility step argue in this answer's ORDER?
A. Engineering should decide the interaction pattern, since they'll build it anyway
B. The live-demo test should happen before the accuracy test
C. Whether Dentscope decides a claim on its own is the scope line hardest to undo once a pipeline is built around it
D. The scope should be written before any prototype exists
Show hint
Ask which decision, once built, stops being a setting and becomes a rebuild.
Show answer
C. Auto-decide versus flag-for-a-person is the one claim that turns an estimate into a commitment nobody can casually reverse.
Short answer, apply it yourself
5. Pick an AI feature you use or are building. What's one thing you'd test on the ugliest real input you have, before letting anyone estimate a sprint around it?
Show hint
Look for the input closest to your actual mess, not your best demo.
Show answer
Model answer: "A receipt-scanning expense app, tested against a crumpled, faded gas station receipt, not the crisp restaurant receipt used in every product demo."
Short answer, the number question
6. If Dentscope had matched 12 of 15 ambiguous claims instead of 6, would the scope still need to flag every ambiguous claim for Nolan? Say what changes and what doesn't.
Show hint
Reversibility is about whether a wrong call can be undone, not about how rare it is.
Show answer
Model answer: "What changes: how often Nolan has to stop and confirm one, so the flagged pile gets smaller. What doesn't change: one missed hinge crack is still a real, unrecoverable payout or denial once it ships, so a confirmation step still belongs in the scope, just a lighter one."
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.