ConceptAdvancedDesigning for Uncertainty & Trust / Feedback loops and data flywheels / #13

Describe the flywheel for a product where more usage genuinely improves quality.

SPARK the scenario: Cobblestead Property Group's Fixwell, a maintenance triage tool for forty rental buildings

Interviewer's question: "Describe the flywheel for a product where more usage genuinely improves quality." Cobblestead Property Group runs Fixwell to triage tenant maintenance requests. Tomasz Okonkwo has dispatched repairs there for five years.

The direct answer
The flywheel runs on corrections, not clicks. Every time a dispatcher overrides Fixwell's vendor guess, or a vendor's invoice later confirms or contradicts it, that becomes one labeled example fed into the next weekly retrain, so routing gets sharper exactly where it was weakest. The loop only stays real as long as those corrections keep happening on cases that are genuinely wrong, not skipped because everyone's started trusting the button.
Do this, in order
  1. Feed the model real corrections, not raw volume.Why: more clicks with no true accept-or-override signal behind them is just more data, not a flywheel.
  2. Retrain on a fixed weekly cadence tied to the correction queue.Why: a correction sitting unused for two months isn't part of a loop, it's a filed complaint.
  3. Watch the override rate itself as the health signal of the flywheel.Why: a falling override rate is either real improvement or dispatchers going quiet, and you can't tell which without checking.
  4. Keep a person confirming the category on day one, no matter how good routing gets.Why: full automation would remove the very correction signal the flywheel needs to keep working.
  5. Check the vendor's actual finished repair against the dispatcher's call, not just the dispatcher's confidence.Why: dispatchers can be confidently wrong together, and only the finished job tells you which category was right.
  6. Leave the refund and invoice math exactly as it works today.Why: this design is about who confirms the category, not how the bill gets calculated.

How to answer this, stage by stage

Nobody is grading whether you can define "flywheel." They're grading whether you can name the exact signal that keeps it spinning, and the exact way it stalls.

Stage 1
Scope it to one concrete product
Say it like this
"I'll answer this for Fixwell, Cobblestead Property Group's maintenance triage tool that reads a tenant's text and photo and suggests which vendor to call."
Why this works
Turns an abstract flywheel question into one loop the interviewer can actually inspect.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, how dispatch works without it. Payoff, the habit I want. Anchor, the actual loop. Risk, what breaks it. Keep out, what's not built yet."
Why this works
Signals a method before a single claim about the loop gets made.
Stage 3
Ground it in today, without the feature
Say it like this
"Right now Tomasz reads every tenant text and photo by hand and guesses the vendor category himself, about four minutes a ticket, worse on a Monday after a weekend of messages."
Why this works
Shows the loop is replacing a real, strained process, not a hypothetical one.
Stage 4
Give the anchor, the loop itself
Say it like this
"Every override Tomasz makes, and every vendor invoice that confirms or contradicts the guess, becomes one labeled example. Those feed a weekly retrain, so the model gets sharper exactly where dispatchers keep correcting it."
Why this works
This is the direct answer: the loop is corrections in, sharper routing out.
Stage 5
Name the failure mode that breaks it
Say it like this
"If dispatchers start rubber-stamping every suggestion instead of really checking, the corrections stop being true corrections, and the model quietly learns from confirmed-wrong labels instead."
Why this works
Answers the real follow-up: what makes a flywheel fake instead of real.
Stage 6
Close on the one line
Say it like this
"A real flywheel runs on corrections, not clicks. The moment people stop actually checking, the loop is still turning, it's just turning the quality down instead of up."
Why this works
Restates the direct answer in one breath, ready for whatever gets pushed on next.

Let's learn

The tool is a shared tablet bolted to the dispatch desk at Cobblestead Property Group, the same one three different dispatchers pass around across a week.

Cobblestead manages forty rental buildings. Fixwell is the feature behind that tablet: a tenant texts what's wrong and a photo, and Fixwell reads it and suggests which kind of vendor to call, plumber, electrician, appliance repair, or pest control.

Before Fixwell, Tomasz read every message by hand, checked the photo, and picked the vendor himself, about four minutes a ticket. On a typical Monday, thirty eight messages had piled up over the weekend, and he wasn't done until well past lunch.

Now Fixwell suggests a vendor instantly, with a note on how sure it is, and Tomasz accepts it or overrides it.

Dispatcher override rate on Fixwell's suggestions, by week
70% 35% 0% Week 1 Week 4 Week 8 Week 12 62% 19%
A falling override rate looks like success. It only means the loop is healthy if the corrections that remain are still honest ones.

Here's the turn: the falling override number is not, by itself, the win. What matters is whether the corrections still happening are real ones, made by someone who actually looked, not someone who's learned to tap "accept" out of habit.

Hand sketched metaphor scene titled Today, without Fixwell. Left, a person icon labeled By hand, caption reads photo, guesses vendor, 4 minutes. Right, a gauge icon labeled Fixwell, caption suggests vendor, under 1 minute.
Same ticket, same dispatcher. What changed is how many of them ever need real judgment at all.
The decision that mattered Every accept or override, plus the vendor's finished invoice, becomes one labeled example. Those examples feed a fixed weekly retrain, so routing keeps improving exactly where dispatchers keep correcting it.
Hand sketched labeled parts diagram titled The anchor, close up. Center gauge icon labeled Fixwell Suggestion, with four callouts: predicted vendor, confidence note, override button, feeds retrain.
Four small parts. The fourth one is the only reason this is a flywheel and not just a shortcut.

At its worst: the override rate keeps falling, everyone reads it as success, and nobody notices that dispatchers have quietly stopped actually checking the photo before tapping accept. The model keeps "improving" on corrections that were never true corrections, and quality drifts the wrong way in silence.

What I would leave alone: the refund and invoice math, the actual dollar amount owed for a repair, doesn't need touching. This design is only about who confirms the category, not how the bill gets calculated.

The lesson: a flywheel isn't proven by usage going up. It's proven by the correction signal staying honest while usage goes up, and those are two different things that happen to move together for a while.

Now here is the same thing as a story

The short version above is what you'd say defending this design to Cobblestead's operations lead. Read this one for how the quiet failure almost happened.

Tomasz Okonkwo has dispatched repairs at Cobblestead for five years. Ask him which building has a temperamental boiler and he'll tell you before you finish the question.

For its first two months, Fixwell mostly agreed with him, and when it didn't, Tomasz caught it fast: a glance at the photo, a tap to override, done in ten seconds.

Then someone on the product team noticed dispatchers were moving faster than expected and added a quiet convenience: if nobody tapped anything within ten seconds, Fixwell's suggestion auto-accepted, so the queue kept moving during a rush.

Nothing about that felt wrong at the time. Tomasz was already fast. The timeout just caught the cases where he'd glance and agree anyway.

The near miss came on a Thursday. A tenant's message said the kitchen smelled like gas, with a blurry photo of the stove. Fixwell read "stove" and suggested appliance repair. Tomasz was mid-call with another tenant when the ten-second timeout hit, and the suggestion auto-accepted before he'd even read it.

Hand sketched comparison diagram titled The day it's wrong. Left panel, a question mark box icon labeled Rubber stamp, caption wrong vendor sent, no correction logged. Right panel, a gauge icon labeled Override caught, caption dispatcher corrects, model learns.
Same ticket, two different desks. Only one of them ever produces a real correction.

The tenant called back forty seconds later, scared, and Tomasz caught it on the second ring, before the appliance repair vendor was even dispatched. He rerouted it to the gas utility's emergency line himself.

We didn't just risk a wrong vendor that day. We risked teaching the model that a gas smell and a broken stove looked the same, because the ten-second timeout had already logged it as a confirmed correct guess.

Here's the decision I'd take back: the ten-second auto-accept default. It made sense when it shipped, dispatchers really were fast, and the queue really did move better. Nobody meant to hide anything. But it meant the loop was quietly counting silence as agreement, and silence is not a correction.

Replayed with the default removed: the same blurry gas-smell photo comes in during the same busy Thursday. Fixwell still suggests appliance repair. But nothing auto-accepts, so it sits flagged until Tomasz explicitly taps something, and when he does look, thirty seconds later, he overrides it himself, correctly, and that override, a real one, becomes the labeled example that teaches the model to weigh "smell" words above "stove" next time.

The old design let the clock make the call. The new one makes sure a person's actual attention is the only thing that ever counts as a correction.

I approved that auto-accept timeout because it felt like a small kindness to a busy dispatcher, ten seconds nobody would miss. It took a scared phone call and a fast second guess to see that those ten seconds were exactly where the flywheel's real fuel was quietly leaking out.

SPARK, in one screenNot a UI exercise. SPARK is what tells you why a flywheel needs a genuine correction, not just a fast one.

S
Situation. How this happens today, without the feature.
Tomasz reads every tenant message and photo by hand and picks the vendor himself, about four minutes a ticket, worse on a Monday.
Grounds the flywheel in a real, strained process before naming the loop.
P
Payoff. The habit this should build.
Trusting Fixwell on the routine, low-stakes cases while genuinely pausing on the ones flagged low-confidence or high-stakes, instead of skimming every ticket the same shallow way.
Names the real behavior change the loop depends on.
A
Anchor. The loop everything hangs on.
Every override, and every vendor invoice that confirms or contradicts a guess, becomes one labeled example, folded into a fixed weekly retrain.
This is the hardest step and the direct answer: the loop only works if the correction is real.
R
Risk. What breaks the first time it's wrong.
A convenience default, like auto-accepting after a timeout, can quietly turn silence into a fake correction, teaching the model something false with real confidence.
Names the exact silent-degradation failure mode a flywheel can hide inside itself.
K
Keep out. What we won't build, day one.
No auto-dispatch without a person confirming the category, no tenant sentiment scoring, no sharing the model across other property companies yet.
Shows judgment about scope, not a wish list of everything Fixwell could eventually do.
Hand sketched icon list titled What we left for later. Three items: a question mark box icon labeled no auto dispatch day one, a gauge icon labeled no tenant sentiment scoring yet, a document icon labeled no sharing model cross property.
Each of these is a real feature Cobblestead could build eventually. None of them is this design's job today.
Hand sketched quadrant titled Sorting maintenance requests. Axes Fixwell confidence and stakes if wrong. Leaky faucet with photo and clogged drain sit high confidence low stakes, bottom right. Gas smell and water near outlet sit low confidence high stakes, top left.
Only the bottom-right corner is where trusting the suggestion is actually safe. The top-left is exactly where a person has to look.
Hand sketched timeline titled How the anchor came together. Four milestones: idea proposed after a dispatch backlog review, pilot with 3 dispatchers, loop wired in with corrections feeding retrain highlighted, full rollout across all 40 properties.
The loop wasn't the whole idea from day one. It's the piece that turned a suggestion tool into an actual flywheel.

The recap, one line per letter: situation is a dispatcher reading every ticket by hand, payoff is trusting the routine case while pausing on a real flag, anchor is the correction-to-retrain loop, risk is a convenience default quietly faking a correction, and keep out is holding back on full automation until the loop is proven solid.

Fixwell's training set: seed data versus dispatcher corrections
1000 500 0 Launch 500 seed Week 12 950 total
Nearly half the training set by week twelve is dispatcher corrections. That's the flywheel showing up as an actual number, not just a metaphor.

And if you want to be sure it really works, try it somewhere elseSame five letters, a smallholder farming app instead of a maintenance desk. A completely different domain, the same correction-fuels-the-loop shape.

SproutLens is a photo-based crop disease app used by smallholder farmers. Priya Adjei is an agricultural extension worker who visits farms and checks the app's diagnoses against what she actually finds on the leaf.

Mapped onto SPARK: situation is a farmer today, waiting days for an extension worker's visit to find out if a spotted leaf means blight or just sun damage; payoff is farmers trusting SproutLens's common, clear-photo diagnoses while still flagging the blurry or unusual ones for Priya's next visit. The anchor is structurally the same loop, aimed at a different correction: every time Priya's on-farm diagnosis confirms or overturns the app's guess, that becomes one labeled example, folded into the next regional retrain, so the model sharpens fastest on exactly the diseases it currently misreads. Risk is the same shape too: if farmers start trusting every diagnosis without sending Priya the unusual cases, the flywheel starves, not from a bad default this time, but from nobody bothering to flag anything anymore. Keep out is no automatic pesticide-purchase recommendations yet, no farm-to-farm disease alerts without a human check, no fully unattended diagnosis in the field.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "the loop runs on corrections, not clicks, protect the correction and the flywheel takes care of itself," and stop.
Cost: if a weekly retrain is too expensive to run at first, batch it monthly, and say so honestly, a real loop that turns slowly still beats a fast loop that isn't real.
The model gets better, for real: if Fixwell's routing gets so accurate that overrides become rare, that's not a reason to stop watching, it's exactly when a fake "no override needed" default becomes the easiest mistake to make.

Where people run it wrong.
They count total usage as proof the flywheel is working, without checking whether the corrections behind that usage are still honest ones.
They add a convenience default, an auto-accept, a skip button, without asking what it quietly teaches the model to treat as agreement.
They let retraining slip to "whenever there's time," which turns a loop into a one-time training run with extra steps.

How to use it live. When someone asks you to describe a flywheel, ask yourself one question first: what specific human action produces the label, and what happens to the loop the day people stop doing that action carefully. Build the anchor around protecting that one action.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "describe the flywheel for a product where usage improves quality"?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. The anchor here is the correction-to-retrain loop itself.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Tomasz Okonkwo, a five-year maintenance dispatcher at Cobblestead Property Group who can name a building's problem boiler on sight.
3 · THE SITUATION
How does dispatch happen today, without Fixwell?
Tap to flip
ANSWER
Tomasz reads every tenant message and photo by hand and picks the vendor himself, about four minutes a ticket, worse on a Monday.
4 · THE ANCHOR
What exactly turns into a labeled training example?
Tap to flip
ANSWER
Every dispatcher override, plus every vendor invoice that confirms or contradicts the guess. Both feed a fixed weekly retrain.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
The ten-second auto-accept timeout, added to keep the queue moving, which quietly let silence count as a correction.
6 · THE NUMBER
Fill in the blank: the override rate fell from 62 percent in week 1 to ___ percent by week 12.
Tap to flip
ANSWER
19 percent. By then, dispatcher corrections made up nearly half the entire training set.
7 · THE REPLAY
Same busy Thursday, redesigned default. What changes?
Tap to flip
ANSWER
The gas-smell ticket doesn't auto-accept. It sits flagged until Tomasz taps something, and his real override becomes a true labeled example instead of a fake one.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the anchor there?
Tap to flip
ANSWER
SproutLens, a crop disease app. Same loop: an extension worker's on-farm confirmation or correction feeds the next regional retrain.

Check yourself Score: 0 / 0

True or false
1. True or false: a falling override rate always means Fixwell's routing is genuinely getting better.
  • True
  • False
Show hint
Look at the line chart's caption.
Show answer
False. A falling override rate can also mean dispatchers have quietly stopped checking. Only a real correction signal tells the two apart.
Multiple choice
2. What made the ten-second auto-accept timeout dangerous to the flywheel specifically?
  • A. It made Fixwell's suggestions slower to appear.
  • B. It annoyed tenants who wanted a faster response.
  • C. It let silence count as a confirmed-correct label, teaching the model something false with real confidence.
  • D. It required Tomasz to log in more often.
Show hint
Look at the highlight line about the gas-smell ticket.
Show answer
C. The flywheel only works if a correction reflects a real check. A timeout that fakes agreement poisons the training signal, not the response time.
Fill in the blank
3. Fill in the blank: by week 12, Fixwell's training set held 500 seed examples plus about ___ dispatcher corrections.
Show hint
Look at the stacked bar chart.
Show answer
450. Nearly half the entire training set, by week twelve, was dispatcher corrections rather than the original seed data.
Short answer, where it wouldn't matter
4. Name a maintenance request in Fixwell where this flywheel risk wouldn't really apply.
Show hint
Look at the quadrant diagram's bottom-right corner.
Show answer
Model answer: A routine leaky faucet with a clear photo, high confidence and low stakes, where an occasional missed override barely costs anything.
Short answer, apply it yourself
5. Pick a product you use yourself. What's one habit it built in you that you'd stop doing if it got a little worse?
Show hint
Think of a tool where you stopped double-checking its suggestion after it was right often enough in a row.
Show answer
Model answer: Many people stop double-checking a map app's suggested route, or a spell-checker's fix, once it's been right so consistently that checking starts to feel pointless.
Short answer, name the reversal
6. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "here's the decision I'd take back."
Show answer
Model answer: The ten-second auto-accept default. It made sense because dispatchers really were fast and the queue really did move better, and it stopped making sense once silence started counting as a real correction.
Before you close the answer
Why this works
Tests whether you can name the exact human action that fuels a flywheel, and the exact convenience feature that quietly breaks it. Most candidates describe the loop but never say what kills it.
Follow-up traps
"Isn't removing the auto-accept default just going to slow the dispatchers back down?" Response: only on the cases that genuinely need a look, the routine, high-confidence cases were never the ones the timeout was protecting anyway.

"How do you know the vendor invoice is a more honest signal than the dispatcher's own tap?" Response: it isn't automatically, which is why both get logged, and a mismatch between the two is itself a useful flag worth reviewing.
If pressed
Fixwell's actual retrain job only includes a correction if the vendor invoice within seven days matches or contradicts the dispatcher's call, a lone override with no invoice yet gets held for the next cycle instead of used half-confirmed.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more