ConceptIntermediateAI Opportunity & Model Strategy / Data strategy as product strategy / #2
What is a data flywheel and what are its preconditions?
ORDERthe six months a coffee mistag stopped getting fixed at Corvid Bank
Corvid Bank's budgeting app is a single swipe on a phone, checked three or four times a day. Tobias Renke leads the categorization feature there, the one that sorts every purchase into "Groceries" or "Dining" or "Transport" automatically. He built the correction flow that was supposed to make the whole thing get smarter over time.
The direct answer
A data flywheel is a loop where using the product makes data, that data makes the product better, and a better product pulls in more use, so each turn is easier than the last instead of resetting to zero. It only spins if three things are true from day one: the feedback gets captured the moment it happens, giving that feedback stays cheap enough to keep it representative, and something downstream actually consumes it to make the product better. Skip any one of those and you have transactions, not a flywheel.
The preconditions, ranked
Capture the feedback signal at the moment it happens, built in from day one.Why: this is the one precondition you can't retrofit. A correction nobody captured is gone forever, unlike a UI friction problem you can fix anytime.
Keep giving feedback cheap enough that it stays representative of every case, not just the ones people care most about.Why: costly feedback gets rationed toward the easy or high-stakes cases, and the training signal quietly stops matching reality.
Have enough usage volume for the feedback to mean something, not just be anecdotal.Why: a flywheel needs enough spin per turn to actually move the wheel forward.
Build a real pipeline that consumes the feedback and retrains the product.Why: feedback that piles up unused isn't a loop. It's a warehouse.
Watch the feedback for bias by segment, not just by raw volume.Why: the correction count can look perfectly healthy while the mix of what's being corrected quietly skews.
Decide, before scaling the loop, what "good enough" looks like.Why: without a target, a spinning wheel and a working one look identical from a distance.
How to answer this, stage by stage
Nobody is scoring whether you can define "flywheel." They're scoring whether you can name the exact thing that breaks it, out loud, before it happens.
Stage 1
Define it in one breath, with a real loop
Say it like this
"A data flywheel is when using the product makes data, the data makes the product better, and the better product pulls in more use. Corvid Bank's spend categorizer is a clean example: people correct a mistag, that correction retrains the model, and the model gets better at everyone's transactions, not just theirs."
Why this works
Gets the definition out fast, grounded in one concrete loop, instead of a paragraph of abstraction.
Stage 2
Say your structure out loud
Say it like this
"I'll run the preconditions as ORDER. Outcome, what we're actually optimizing for. Reversibility, which precondition can't be fixed later. Dependency, what has to exist before what. Evidence, what's cheap to check first. Rank, the actual order that matters."
Why this works
Signals you have a repeatable way to answer "what are the preconditions," not just a list you're making up live.
Stage 3
Reframe: "preconditions" means "what can't be retrofitted"
Say it like this
"The interesting question isn't which preconditions exist. It's which one you can't go back and add later if you missed it. That's the one that actually deserves the word precondition."
Why this works
This is where the answer stops being a checklist and starts being a real argument about sequencing.
Stage 4
Give the direct answer
Say it like this
"The wheel spins only if feedback is captured at the moment it happens, stays cheap enough to be representative, and gets consumed downstream. Miss any one, and you don't have a flywheel."
Why this works
This is the sentence an interviewer should be able to write down and check every follow-up question against.
Stage 5
Prove it with the compressed failure
Say it like this
"At Corvid Bank, a redesign turned the one-tap correction into a three-tap detail modal. Correction on transactions under fifteen dollars fell from 38 percent to 6 percent in six months, while corrections on purchases over a hundred dollars barely moved. The model kept learning from big purchases and quietly stopped learning about coffee and subscriptions."
Why this works
Compresses the whole failure into one number split two ways, which is the number the argument would fall apart without.
Stage 6
Say what you'd measure going forward
Say it like this
"I'd track correction rate broken out by transaction size, not just overall accuracy, since overall accuracy stayed flat the entire time this was quietly going wrong."
Why this works
Shows you'd catch this kind of drift before six months pass, not after.
Stage 7
Close on one line
Say it like this
"A flywheel doesn't fail because nobody's using the product. It fails quietly, one small correction at a time, when giving feedback stops being cheap for the cases that matter most."
Why this works
Restates the direct answer with the story's own number doing the closing work.
Let's learn
Here is what a data flywheel actually needs, and what a redesign can quietly take away from it.
Corvid Bank's budgeting app auto-sorts every purchase into a category, and if it gets one wrong, a user can fix it with a single tap. That correction was supposed to retrain the categorizer, so the whole system got smarter with every mistake anyone caught. Before a redesign changed the correction flow, 38 percent of miscategorized transactions got corrected within a week, regardless of whether the purchase was five dollars or five hundred.
The five letters, held up as one page. Reversibility is the step that separates a real precondition from a nice-to-have.
Then a redesign added a detail modal to the correction flow, asking for merchant type and whether the purchase was recurring, to feed a future "smart budgets" feature. What used to be one tap became three, and one of those taps meant typing.
Here's the turn: the extra taps themselves weren't the real problem. The real problem is what people do when giving feedback gets three times as expensive. They don't stop giving feedback altogether. They start rationing it toward the corrections that feel worth the extra effort, and quietly stop bothering with the ones that don't.
Correction rate for small transactions, by month
Overall correction volume looked steady the entire time, since big-purchase corrections held their ground. This is the number that was actually collapsing underneath it.
At its worst, the categorizer keeps getting sharper about rare, big-ticket purchases while it quietly forgets how to handle the recurring five-dollar coffees and eleven-dollar subscriptions that make up most of what people actually buy, which is exactly the part of the product a budgeting app is supposed to be good at.
The flywheel doesn't need everyone to keep giving feedback. It needs every kind of case to keep giving feedback. The moment one kind gets expensive to correct, the loop starts learning about a world that isn't the whole picture anymore.
The choice I would take back
Merging the quick one-tap fix with a mandatory detail modal, to capture richer labels for a future feature. That made sense when the team wanted better merchant-type data for "smart budgets." It stopped making sense once it meant every small, everyday correction had to pay the same three-tap tax as a rare, large purchase.
What I would leave alone: I wouldn't add friction anywhere near the correction for genuinely one-off, high-value purchases, since people already correct those reliably and a detail prompt there barely costs them anything they weren't already willing to spend.
The lesson: a data flywheel isn't powered by how much people use the product. It's powered by how cheap it stays to tell the product it's wrong, for every kind of case, not just the ones people happen to care about.
Now here is the same thing as a story
The short version above is what you'd say defending the correction flow's design in a roadmap review. Read this one for how the skew built up over six months with nobody deciding it on purpose.
Every night before bed, half of Corvid Bank's users open the app for thirty seconds, scroll the day's spending, and tap to fix whatever got mistagged. It's a small habit, barely worth noticing, and it was the entire reason the categorizer kept improving.
One extra step doesn't sound like much, until it's the step standing between a cheap fix and no fix at all.
Tobias shipped the detail modal in March, to power a "smart budgets by merchant" feature planned for later that year. The team tested it, watched correction volume, and it looked fine. People were still tapping to fix things. Nobody split the number by transaction size, so nobody saw what was actually happening underneath the average.
Nobody decided, on any single day, to stop teaching the model about small purchases. The habit just quietly thinned out, one skipped modal at a time.
By month four, a five-dollar coffee mistagged as "Entertainment" instead of "Dining" sat uncorrected for weeks, because who's going to spend fifteen seconds filling out a merchant-type form to fix five dollars. A three-hundred-dollar mistagged furniture purchase still got fixed the same day, because at that size, the extra taps felt worth it.
Knowledge spark: why does cheap feedback matter more than a lot of feedback?
A model trained mostly on corrections from expensive, rare purchases learns a world where most spending looks like that. But most real spending is small and recurring. If the cheap, everyday corrections dry up first, the model gets more confident about the cases it sees least often, and quietly worse at the ones that make up most of a real budget.
Two clusters, not one blended average. Small, recurring spending fell into the corner nobody was watching.
By month six, Tobias pulled the categorizer's accuracy by segment for an unrelated review and noticed something odd: accuracy on recurring subscriptions had actually dropped four points over the same period the model's overall accuracy metric looked completely flat.
Corvid Bank's loop wasn't a graveyard and it wasn't starved for volume. It landed in the third branch: technically alive, quietly biased.
When the modal was proposed, someone in the room said, "let's just capture merchant type while we have the user's attention," and it sounded completely reasonable, since the future feature genuinely needed that data and asking felt like the obvious moment to do it.
Correction rate by transaction size, before and after the modal
Large purchases barely noticed the extra friction. Small purchases lost nearly all of theirs, which is exactly the segment the flywheel needed most.
Rerun the same six months with the quick one-tap fix kept intact, and the detail modal offered only as an optional follow-up for people who want to add it: small-transaction corrections stay near 35 percent instead of collapsing to 6 percent, subscription-category accuracy holds steady instead of slipping four points, and the "smart budgets" feature gets its merchant-type data from the smaller group of users willing to fill in the extra detail, without taxing everyone else's easy fixes to get it.
What I'd tell myself, watching six months of aggregate accuracy hide a real, growing hole: the modal was never wrong to want. It was wrong to make everyone pay for it, including the people fixing a five-dollar coffee who never needed to answer a question about merchant type at all.
ORDER, the ranking that would have caught this in month twoNot a definition exercise. ORDER is what tells you which precondition is the one you can't add back later.
O
Outcome. What all candidates compete to move.
A categorizer that keeps getting better at every kind of spending, not just the rare, expensive kind.
Without a shared outcome, ranking preconditions is just a list of nice-sounding words.
R
Reversibility. Which precondition can't be retrofitted.
Cheap, low-friction feedback capture has to exist from day one. A missed correction from six months ago cannot be recovered. A UI change can be shipped anytime.
This is the hardest step, and the one the "let's just add one more question" instinct always runs straight past.
D
Dependency. What unblocks what.
You need cheap capture before volume matters, and enough volume before a retraining pipeline is worth building at all.
Some order is forced by what the loop actually requires, not by what's easiest to build first.
E
Evidence. What's cheap to learn first.
Splitting correction rate by transaction size for one week would have shown the small-purchase collapse in month two, not month six.
Cheap evidence, checked early, is what would have caught this before four extra months of drift.
R
Rank. State the order, defend the top.
Cheap, universal feedback capture first. Volume and a real retraining pipeline second. Bias monitoring by segment third, always running.
The ranking follows from what can't be undone, not from which precondition sounds most impressive in a slide deck.
The recap, one line per letter: outcome is a categorizer that keeps improving on every kind of spending, reversibility is that cheap universal feedback capture can't be added back after the fact, dependency is that capture has to exist before volume matters and volume before a pipeline is worth building, evidence is that a one-week segment split would have shown the collapse four months early, and rank puts cheap capture first, a real pipeline second, and segment monitoring third.
And if you want to be sure it really works, try it somewhere elseSame five letters, a precision-farming app instead of a banking app. Different flip family entirely, the same missing precondition.
Anselm Rurik runs product at Hallowfield Agritech, where farmers photograph a crop leaf and the app diagnoses disease, with a correction option if the app got it wrong. Mapped onto ORDER: outcome is a diagnosis model that keeps improving across every regional pest strain, not just the common ones photographed most. Reversibility says regional photo representation, captured from the start, can't be retrofitted; a "why this diagnosis" explainer screen can be added anytime. Dependency says you need enough regional volume before a local specialization feature is worth building. Evidence favors checking, cheaply, whether a given region's pest strains are even represented in the current training set before promising local accuracy gains.
The flip here is verification, not substitution. For one full growing season, the app's diagnoses were accurate enough that farmers stopped spot-checking them against their own field experience. Then a new pest strain arrived in one region, and the app confidently misdiagnosed it as a known, treatable disease for several farmers in a row, since that region's strain had never been in the training data at all. Once word spread between neighboring farms, farmers started manually re-verifying every single diagnosis against a physical field guide before treating anything, which undid the entire point of a fast, automated first read.
The third step is where Hallowfield's loop actually broke: corrections on the new strain were captured, but nothing fed them back before the damage spread between farms.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "a flywheel needs cheap capture, enough volume, and a pipeline that actually uses the feedback, in that order," and stop.
Cost: no budget to build a full bias-monitoring dashboard before the next feature ships. Say so honestly, and manually split correction rate by segment once a month until there's budget for more.
The model got better, for real: Corvid Bank's overall accuracy metric genuinely held steady the whole six months. That's exactly why nobody caught the skew. A healthy average is not proof every segment underneath it is healthy too.
Where people run it wrong.
They treat "we added a feedback button" as proof the flywheel exists, without checking whether giving that feedback stayed cheap for every kind of case.
They watch an aggregate accuracy number and assume it represents every segment equally.
They add a data-collection step to serve a future feature, without asking what it costs the everyday corrections happening right now.
How to use it live. The moment an interviewer asks about preconditions, ask yourself: which one of these, if missing today, could I never get back later? Answer that first, rank the rest under it, and the response builds itself from there.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Substitution flip: once correcting got costlier, users kept fixing big, rare purchases and quietly stopped fixing cheap, everyday ones, rationing their effort toward the cases that felt worth it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Tobias Renke, who leads the spend-categorization feature at Corvid Bank and built the correction flow meant to power its data flywheel.
3 · THE HABIT
What did users stop doing because it no longer felt worth the effort?
Tap to flip
ANSWER
Correcting small, everyday mistagged purchases, like coffee or subscriptions, once the fix required a three-tap detail modal instead of one tap.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Correcting a mistake versus letting it ride. No middle setting once the extra taps made small corrections feel not worth doing at all.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Merging the one-tap correction with a mandatory detail modal to capture merchant-type data for a future feature, taxing every correction to serve one use case.
6 · THE NUMBER
Fill in the blank: correction rate for transactions under $15 fell from 38% to ___ over six months.
Tap to flip
ANSWER
6 percent, while correction on purchases over $100 barely moved, from 41% to 39%.
7 · THE REPLAY
Same six months, one-tap fix kept intact and the detail modal made optional. What changes?
Tap to flip
ANSWER
Small-transaction corrections stay near 35% instead of collapsing to 6%. Subscription-category accuracy holds steady instead of slipping four points.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Hallowfield Agritech's crop-disease diagnosis app. The flip is verification: farmers went from trusting every diagnosis to re-checking every single one against a field guide, after a new pest strain got misdiagnosed for several farmers in a row.
Check yourself Score: 0 / 0
Multiple choice
1. According to this answer, which precondition of a data flywheel can't be fixed after the fact?
A. A polished UI for the correction screen.
B. Capturing cheap, representative feedback from the start.
C. A marketing plan to get more users.
D. A dashboard showing overall accuracy.
Show hint
Look at the "reversibility" step of ORDER.
Show answer
B. A missed correction from months ago can't be recovered later, unlike a UI change, which can ship anytime.
True or false
2. True or false: Corvid Bank's overall accuracy metric dropped noticeably during the six months this problem was building.
True
False
Show hint
Look at "the model got better, for real" in the swap-the-trigger section.
Show answer
False. Overall accuracy stayed flat the whole time, which is exactly why nobody noticed the skew building underneath it.
Fill in the blank
3. Fill in the blank: after the redesign, correction on purchases over $100 held at ___ %, almost unchanged from before.
Show hint
Look at the grouped bar chart, "Correction rate by transaction size, before and after the modal."
Show answer
39%. Barely different from the 41% baseline, since large purchases were already worth the extra effort to correct.
Short answer, where it wouldn't matter
4. Name a place in this same product where adding friction to feedback genuinely would NOT hurt the flywheel, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Corrections on rare, high-value purchases. People already correct those reliably, so a small amount of added friction there costs almost nothing they weren't already willing to spend.
Short answer, apply it yourself
5. Think of a product you use that asks for your feedback or corrections. What would make giving that feedback so costly that you'd quietly stop bothering for the small stuff?
Show hint
Think about a feedback flow that used to be one click and later grew extra steps.
Show answer
Model answer: A music app that used to let you thumbs-down a bad recommendation in one tap, then added a "tell us why" survey first. Most people would keep flagging a truly awful suggestion but stop bothering for a merely mediocre one.
Short answer, work the number
6. If small-transaction correction had held at 35% instead of falling to 6%, roughly how many more corrections would you expect out of 10,000 small mistagged transactions in a month?
Show hint
Compare 35% of 10,000 against 6% of 10,000.
Show answer
Model answer: About 3,500 corrections at 35%, versus about 600 at 6%. That's roughly 2,900 more corrections a month, which is 2,900 more lessons the model never got to learn.
Before you close the answer
Why this works
Tests whether you'll define a flywheel as a slogan or as a set of specific, checkable preconditions, and whether you know which one of those preconditions can't be added back once it's missing.
Follow-up traps
"Couldn't Corvid Bank just remove the modal once they noticed the problem?" Response: yes, and they did, but the six months of missing small-transaction corrections are gone. The fix stops new damage; it doesn't recover the old signal.
"Isn't asking for merchant-type data a reasonable thing to want?" Response: yes, which is why the fix isn't deleting the modal, it's making it optional so it doesn't tax the corrections that were already working.
If pressed
The fix Corvid Bank shipped split the flow: the one-tap fix stayed the default action, and the detail modal became a separate, optional "add more detail" button shown only after the quick fix already saved.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.