CaseAdvancedAI Opportunity & Model Strategy / Roadmapping under model uncertainty / #5
Explain how you would sequence features so that a model improvement unlocks rather than obsoletes work.
ORDERthe bulk pass that rewrote three years of a photographer's own corrections in one afternoon
The feature that looks most impressive in a demo is often the one you should build last. Ledgeport is a personal budgeting app that auto-categorizes transactions for freelancers and small business owners. Otto Reindl is Head of Product, and the question in front of him is how to order four candidate features so a better model adds to what's already built instead of quietly wrecking it.
The direct answer
Build the parts that never need to be rebuilt first: a place to persist every correction a user makes, then a confidence-tiered queue that uses it, then a control that lets big imports run in small slices. Only after those three hold do you build the flashy one-shot feature, applying an improved model to years of old history at once, because that's the one piece of work a bad rollout can't take back.
Build in this order
Persist every correction a user makes.Why: nothing else can improve safely without a record of what a person already fixed by hand.
Ship a confidence-tiered review queue on top of it.Why: this is the piece that gets better for free every time the model does, no rebuild required.
Add a slice-size control for any import.Why: fixes the actual unit-of-work mistake, and makes every later bulk operation reversible in small pieces instead of one giant one.
Only then, build bulk historical re-categorization, gated by evidence.Why: it's the most tempting demo and the hardest one to undo if the model isn't ready, so it goes last, not first.
Pilot the bulk pass on a small slice before running it on anyone's full history.Why: cheap evidence beats an expensive apology after three years of tags get quietly overwritten.
How to answer this, stage by stage
Nobody is scoring whether you can name four features. They're scoring whether you can say which one is safe to leave for later, and why the order isn't just a guess.
Stage 1
Ground it in one real product and one real feature list
Say it like this
"Let's use Ledgeport, a budgeting app that auto-tags transactions. Four candidates: correction history, a confidence queue, slice-size controls, and bulk historical re-categorization."
Why this works
Keeps the sequencing question from turning into an abstract debate about roadmaps in general.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as ORDER. Outcome, what all four are actually competing to move. Reversibility, which one is hardest to undo. Dependency, what unblocks what. Evidence, what I could learn cheaply first. Rank, the actual order."
Why this works
Signals a repeatable way to sequence any feature list, not just this one.
Stage 3
Name the outcome before ranking anything
Say it like this
"All four features are competing to move one thing: how much a freelancer trusts the app enough to stop double-checking it by hand. Whichever order gets there without ever having to un-teach that trust wins."
Why this works
Without a stated outcome, ranking is just opinion, this is what makes the rest of the answer checkable.
Stage 4
Find the one that's hardest to undo
Say it like this
"Correction history, the confidence queue, and slice controls are all cheap to adjust later. Bulk re-categorization of three years of history in one pass is not, because a bad rollout burns trust that took years to build, in one afternoon."
Why this works
This is the actual judgment call ORDER is testing, stated in plain, checkable terms.
Stage 5
Say what unblocks what
Say it like this
"You can't build a trustworthy confidence queue without correction history to calibrate it. You can't run a safe bulk pass without a slice-size control that lets you stop halfway through if something looks wrong. The order isn't a preference, it's forced by what each piece actually needs."
Why this works
Shows the sequence is dictated by real dependencies, not just a hunch about priority.
Stage 6
Give the rank, and defend the top pick
Say it like this
"Correction history first, confidence queue second, slice controls third, bulk re-apply last, gated by a small pilot. Correction history goes first because every other feature on this list silently depends on it existing."
Why this works
This is the direct answer, restated as a concrete, defensible order.
Stage 7
Prove it with the compressed failure
Say it like this
"Ledgeport actually built bulk re-categorization first, because it demoed the best. It ran an improved model over three years of a freelancer's history in one pass, with no way to slice it and no memory of her own corrections. It overwrote work she'd spent hours getting right."
Why this works
Compresses the whole failure into the exact scenario ORDER's ranking was built to prevent.
Stage 8
Close on the rule that survives it
Say it like this
"The rule I'd take from this: never let the most demoable feature also be the hardest one to undo. If it's both, it goes last, and it gets a pilot before it ever touches someone's full history."
Why this works
Ends on a portable principle, not just a fact about one company's mistake.
Let's learn
Here is what happens when the feature that demos best gets built first, and the ones that would have protected it get built last, if at all.
Before Ledgeport, a freelancer categorizing business expenses for taxes did it by hand in a spreadsheet, roughly two hours a month for a typical year of transactions. With Ledgeport, a model auto-tags each transaction as it comes in, and most users check maybe one in ten by hand. Ledgeport's roadmap had four candidate features competing for the next quarter: correction history, a confidence-tiered review queue, a slice-size control for imports, and a splashy bulk re-categorization tool that could apply a newly improved model to years of old transactions in one pass.
Each box quietly needs the one before it. The last box is the one everyone wanted to build first.
Here's the turn: the team built bulk re-categorization first, since it was the easiest to demo and the most exciting line on the roadmap. It shipped without correction history and without a slice control, because neither of those existed yet either. The first time it ran against a real user's three years of transactions, it didn't fail loudly. It just quietly replaced her own corrections with the new model's fresh guesses, transaction by transaction, with no record that anything had changed.
Manual overrides Colette re-applied, before and after the bulk pass
One afternoon's bulk pass generated more rework than two years of ordinary use ever had.
Minutes Colette spent reviewing categorized transactions per week, before the pass through six weeks after
Six weeks later, Colette still spent three times as long reviewing her books as she had before the bulk pass. Trust doesn't reset to zero cost once it's broken.
At its worst, the feature meant to showcase how much smarter the model had gotten instead erased the one thing that had been quietly making the product trustworthy all along: a person's own corrections.
The choice I would take back
Ledgeport shipped bulk re-categorization as the only way to apply a model improvement to old data, all-or-nothing, no slice, no check against prior corrections. That made sense when the team was small and the feature was the fastest way to show the model's progress to leadership. It stopped making sense the moment the feature touched a real user's years of careful, hand-made fixes.
What I would leave alone: the live, transaction-by-transaction categorization for new purchases. It was never the problem, and slowing it down to add extra safeguards would cost more in daily friction than it would prevent.
The lesson: the feature that best proves a model got smarter is often the one most capable of proving it in a way nobody can undo.
Now here is the same thing as a story
The short version above is what you'd say scoping this live in an interview. Read this one for how a routine audit turned one freelancer's ruined afternoon into a sequencing rule for the whole company.
Colette Fairweather has shot weddings and small-business headshots for five years, and she treats her books the way she treats her camera settings: checked, adjusted, trusted only once it's proven itself. She'd used Ledgeport for two years, mostly hands-off, correcting maybe two or three miscategorized transactions a month out of a few hundred.
For those two years, Ledgeport worked the way she expected: a transaction came in, got tagged, and if she fixed it, the fix stuck. She trusted it enough to stop keeping her old spreadsheet backup at all.
Same roadmap quarter, two very different kinds of feature hiding inside it.
Then Ledgeport announced an improved categorization model and, to show it off, ran the new model over every user's full transaction history in a single overnight batch, replacing old tags with new ones wherever the model's confidence had changed. Colette's three years of tax-ready records, including every correction she'd made by hand, got quietly recomputed from scratch, all at once, with nothing preserved.
Knowledge spark: why does "the model got better" still break old, correct work?
A newer model doesn't know which of the old tags were the model's own guess and which ones a person had fixed by hand. Without a record telling them apart, applying the new model to everything treats a careful person's correction exactly the same as a guess that was never checked.
Colette didn't notice for three weeks, until she pulled her books for a client invoice and found categories that made no sense: a camera-gear purchase tagged as a meal, a mileage reimbursement tagged as software. Nobody at Ledgeport had decided to overwrite her corrections specifically. The bulk pass applied to everyone's history the same way, and hers happened to have more corrections built up than most, so it broke louder.
The same freelancer, the same app, and a trust that now has to be rebuilt in much smaller pieces.
The bulk pass didn't cost Colette three weeks of noticing. It cost her the two years of not having to think about her books at all, the exact thing Ledgeport was supposed to give her.
After that, Colette stopped trusting any bulk operation in the app. She started re-uploading disputed months in batches of about twenty transactions she could fully review by hand before approving, giving up most of the time savings she'd built up trust for in the first place.
The bulk pass sits exactly where a risky feature shouldn't: high impact and almost impossible to walk back.
It took a routine quarterly audit, someone on Otto's team pulling ten random users' correction logs, to notice Colette's override count had jumped from a handful a month to over two hundred in a single month, far outside anything the team expected from ordinary use.
Rerun in this order, the bulk pass never runs until the pieces that make it safe already exist.
Rerun the same model improvement with this order already built: bulk re-categorization checks correction history first, skips any transaction a person already fixed, and runs in slices small enough to pause and review. Colette's corrections stay hers. The model improvement adds new coverage to transactions nobody had touched yet, instead of quietly undoing the ones she'd already handled.
What I'd tell myself, watching that audit uncover Colette's ruined afternoon: building the flashiest feature first was never wrong because it was ambitious. It was wrong because it was also the one feature a mistake couldn't be quietly walked back from.
ORDER, the rank that survives being pushed on
O
Outcome. What all four candidates compete to move.
Whether a freelancer trusts the app enough to stop double-checking it, without that trust ever needing to be rebuilt from scratch.
Without naming this first, ranking four features is just a guess dressed up as a plan.
R
Reversibility. Which decision is hardest to undo.
Correction history, the confidence queue, and slice controls are all cheap to adjust later. A one-shot bulk pass over years of history is not.
Flips don't flip back, and neither does a user's trust after a bad bulk rewrite.
D
Dependency. What unblocks what.
The confidence queue needs correction history to calibrate against. A safe bulk pass needs slice controls to be interruptible. The order is forced by these, not chosen freely.
This is the hardest step, and the one that turns a preference into a real sequence.
E
Evidence. What you could learn cheaply first.
Pilot the bulk re-apply on a small slice of low-stakes transaction types before ever running it on a full history.
Cheap evidence here would have caught Colette's problem before it reached her books at all.
R
Rank. The actual order, defended.
Correction history, confidence queue, slice controls, then bulk re-apply last, gated by a pilot.
Correction history goes first because every other item on the list silently depends on it.
The recap, one line per letter: outcome is trust that never has to be rebuilt, reversibility is spotting that a bulk pass is the one item a mistake can't undo, dependency is the confidence queue and safe bulk pass both needing correction history and slice controls first, evidence is a small pilot before any full-history run, and rank is the concrete order: history, queue, slices, then bulk re-apply last.
And if you want to be sure it really works, try it somewhere elseSame five letters, a court-translation service instead of a budgeting app. Different flip family entirely, the same tempting one-shot feature.
Brellin Language Services runs a live transcription and translation tool for court interpreters, with a roadmap candidate to auto-replace flagged uncertain phrases across old session transcripts once a new language model ships. Mapped onto ORDER: outcome is interpreters trusting live flagged terms enough to glance rather than re-listen to a whole recording. Reversibility flags the same danger as Ledgeport's bulk pass: auto-replacing terms across historical transcripts, including ones an interpreter already corrected on the record, is far harder to undo than a live confidence flag ever is. Dependency is a correction log for past interpreter fixes needing to exist before any historical reprocessing is safe. Evidence is trialing historical reprocessing on transcripts with no prior human corrections first. Rank puts live confidence-flagging first, historical reprocessing last, gated by evidence, same shape as Ledgeport's order. The flip here is input, not scope: Ingrid Kastrup, a court translator, started speaking more slowly and simply into her headset mic once she noticed the model mishearing certain accented terms, shaping her own natural speech around the tool's weak spot instead of translating the way she actually would in the room.
The same four parts keep a translation tool's improvements from quietly overwriting a court record.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "build what never needs rebuilding first, and put the one-shot feature last, gated by a pilot," and stop.
Cost: no time to build a full slice-size control this quarter. Say so honestly, and cap the maximum bulk-operation size hard, at code level, until a real control exists.
The model gets better, for real: if the improved model genuinely clears its bar on a small pilot slice, that's the evidence step working exactly as designed, and the honest move is expanding the bulk pass gradually, not all at once out of excitement.
Where people run it wrong.
They rank features by how good they'll look in a demo, not by how hard a mistake would be to undo.
They build the flashy one-shot feature first because leadership wants to see model progress, and treat the safety infrastructure as a "later" problem.
They discover the damage from a support ticket or an angry review instead of a standing audit that would have caught it in week one.
How to use it live. The moment someone asks how to sequence features around a model that's still improving, ask yourself which candidate is hardest to walk back if you're wrong. That one goes last, no matter how good it looks on a slide.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Scope flip: after the bulk pass overwrote her corrections, Colette stopped trusting any large operation and started feeding the app disputed months in batches of twenty she could fully review by hand.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Colette Fairweather, a five-year freelance photographer who treats her books with the same care she gives her camera settings.
3 · THE HABIT
What did Colette stop doing during her first two years on Ledgeport?
Tap to flip
ANSWER
She stopped keeping her old spreadsheet backup at all, trusting that a correction she made once would simply stick.
4 · THE SEQUENCE, IN THIS STORY
What's the correct build order this answer argues for?
Tap to flip
ANSWER
Correction history first, confidence-tiered queue second, slice-size control third, bulk historical re-categorization last, gated by a small pilot.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Shipping bulk re-categorization as the only way to apply a model improvement to old data, all-or-nothing, a call that made sense when the team wanted the fastest way to show progress to leadership.
6 · THE NUMBER
Fill in the blank: in the single month after the bulk pass, Colette had to redo ___ manual overrides.
Tap to flip
ANSWER
210, more than five times the roughly 40 total overrides she'd made in the two years before it.
7 · THE REPLAY
Same model improvement, same rollout, but the four features built in the recommended order. What changes?
Tap to flip
ANSWER
The bulk pass checks correction history first and skips anything Colette already fixed, running in small slices instead of all at once. Her prior corrections stay hers, and the model improvement only adds coverage to transactions nobody had touched.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Brellin Language Services' court transcription tool. The flip is input: Ingrid Kastrup started speaking more slowly and simply into her headset to avoid the model mishearing accented terms.
Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: this answer's recommended build order is correction history, confidence queue, slice controls, then ___ last.
Show hint
Look at the priority list and the rank step.
Show answer
Bulk historical re-categorization. Gated by a small pilot before it ever touches a full history.
Multiple choice
2. Why does this answer rank bulk re-categorization last instead of first?
A. Because it's the most expensive feature to build technically.
B. Because users never asked for it.
C. Because it's the hardest of the four to undo if the model isn't ready.
D. Because it requires a bigger engineering team than the others.
Show hint
Look at the "reversibility" step and the quadrant diagram.
Show answer
C. A bad bulk pass overwrites years of careful corrections in one afternoon, with no way to walk it back.
True or false
3. True or false: this answer argues Ledgeport should never build bulk historical re-categorization at all.
True
False
Show hint
Look at the rank step and "what I would leave alone."
Show answer
False. It should still get built, just last, gated by a pilot, and checked against correction history so it never overwrites a user's own fixes.
Short answer, where it wouldn't matter
4. Name a part of Ledgeport where this sequencing concern genuinely doesn't apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Live categorization of new incoming transactions. It was never the problem, and adding extra sequencing safeguards there would only add daily friction for no real benefit.
Short answer, apply it yourself
5. Think of a product update you've seen that "improved" something automatically and broke work you'd already done by hand. What would have made that update safer to sequence?
Show hint
Think of a photo app's auto-organize feature, a spreadsheet's auto-formatting, or a note-taking app's re-tagging feature.
Show answer
Model answer: A photo app's "smart album" reorganization that ignored manually renamed albums, which would have been safer if it checked for and skipped anything a person had already touched.
Short answer, work the number
6. If Ledgeport has 50,000 users with an average of 3 years of history, and even 2 percent see a Colette-sized override spike from an unsequenced bulk pass, roughly how many users is that?
Show hint
2 percent of 50,000.
Show answer
Model answer: About 1,000 users, each facing hours of unexpected rework, entirely avoidable with the right build order.
Before you close the answer
Why this works
Tests whether you'll rank features by how good they'd look in a demo, or by which one a mistake in it can't be walked back from once the model improves.
Follow-up traps
"Doesn't this order just delay the most valuable feature?" Response: it delays the riskiest one, not the most valuable one, and delivers most of the trust benefit earlier through the confidence queue, which improves for free as the model does.
"What if leadership demands the flashy feature for a launch deadline?" Response: ship it gated by a small pilot and a hard cap on scope, so the deadline gets met without betting someone's full history on an unproven pass.
If pressed
The version that shipped required any bulk re-apply to skip transactions with a correction timestamp newer than the model's training cutoff, so a user's own fix could never be silently overwritten by a model that hadn't even seen it yet.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.