ConceptAdvancedAI Opportunity & Model Strategy / Data strategy as product strategy / #19
How do you measure whether your data flywheel is actually turning?
LEADthe number that stopped moving three releases before anyone checked it
Here is what happens when a loop that is supposed to feed itself quietly stops feeding itself, and every dashboard keeps saying it's fine. Cascade Mobile is a regional wireless carrier. ReplyLine is its tool that drafts replies to customer support chats, and is supposed to get better each release as support agents' edits feed back into training. Selin Marek is the senior AI PM who has to prove, every quarter, that this loop is actually turning and not just spinning.
The direct answer
Don't measure the flywheel by whether agents accept drafts more often. Measure how much they still have to rewrite before sending, tracked release over release, and pair it with an audit of what fraction of logged edits actually made it into the last training run. If the rewrite number stops falling while acceptance keeps climbing, the loop has stalled and something upstream is broken, no matter how good the headline metric looks.
Do this, in order
Track edit-depth, not accept rate, as the flywheel's real gauge.Why: accept rate rises when agents get tired of checking just as easily as it rises when the model gets better, and you can't tell the two apart from that number alone.
Audit what fraction of logged edits actually reach the training set each release.Why: a flywheel can look busy at the logging end and be empty at the retraining end, and nothing on a usage dashboard shows you that gap.
Set a real threshold: no drop in edit-depth across two releases means stop and audit before shipping another one.Why: a metric nobody acts on is decoration, not a gauge.
Watch the lagging outcome too, but don't wait for it to move before acting.Why: cost per resolved contact can sit still for months while the real damage is already locked in upstream.
Say plainly where accept rate is still a fine number to watch.Why: on a healthy loop, accept rate and edit-depth move together, and it would be wrong to throw the number out everywhere just because it can be gamed.
How to answer this, stage by stage
Nobody is scoring whether you can name a metric. They're scoring whether you can catch the one that's lying to the room.
Stage 1
Scope it to one loop
Say it like this
"I'll ground this in ReplyLine at Cascade Mobile, a support-reply drafting tool that's supposed to improve as agents edit its drafts and those edits get folded back into training."
Why this works
Stops the answer from turning into a lecture about flywheels in general.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as LEAD. Link, the business outcome underneath it. Early signal, the number that moves first. Abuse, how that number gets gamed. Decision, what I'd actually do at each reading."
Why this works
Shows a repeatable way to pick a metric instead of naming the first one that comes to mind.
Stage 3
Reframe: "is it improving" isn't one question, it's two
Say it like this
"'Is the flywheel turning' is really two separate questions people mash together: is the model getting better, and are agents behaving as if it is. Those can move in opposite directions, and most teams only ever check the second one."
Why this works
This is the line that separates a candidate who's thought about gamed metrics from one reciting "north star metric."
Stage 4
Give the one decision
Say it like this
"I'd watch edit-depth ratio, how much of each draft's characters an agent actually rewrites before sending, release over release. If it isn't falling, I don't care what accept rate says, I go audit whether logged edits are actually reaching the training set."
Why this works
This is the direct answer stated as something you would actually go check on a Tuesday.
Stage 5
Prove it with the compressed failure
Say it like this
"At Cascade Mobile, edit-depth sat at 42, 41, 40, 39 percent across four releases, basically flat, while accept rate climbed from 55 to 74 percent over the same stretch. Turned out a pipeline bug meant only about one in ten logged edits from the last three releases ever reached the model. Agents weren't rewriting less because drafts got better. They were rewriting less because they'd stopped expecting to need to."
Why this works
Compresses the whole failure into the one gap, between a gamed number rising and a real one standing still.
Stage 6
Say what you'd measure past launch
Say it like this
"Past the launch metrics, I'd keep a standing ingestion audit: what percent of qualifying edits from the last release actually appear in this release's training data. That number alone would have caught the bug in week one instead of month nine."
Why this works
Shows you think about the pipeline itself as a thing that can silently fail, not just the model's output.
Stage 7
Name the AI-specific reasoning, then close
Say it like this
"The reason this isn't just generic analytics advice is that the model's own output shapes what data comes back to it. A metric that only counts what agents did, not what the model actually learned from, will always be gameable, because agent behavior and model behavior can drift apart without either dashboard noticing on its own."
Why this works
Closes on the judgment call, not a summary, and restates the direct answer in one breath.
Let's learn
Here is what happens when the metric everyone watches can go up for a reason that has nothing to do with the thing it's supposed to measure.
Before ReplyLine, a Cascade Mobile support agent read each incoming chat and typed a reply from scratch, about ninety seconds a ticket, several hundred tickets a shift split across a team. With ReplyLine, a draft reply appears the moment a ticket opens, and the agent edits it or sends it as is. The whole pitch of the tool was that every edit an agent makes gets logged and folded into the next training run, so the drafts should need less and less rewriting over time. That loop, working, is the flywheel.
The third step is where Cascade Mobile's loop actually broke, quietly, for three releases running.
Here's the turn: the metric Cascade Mobile actually watched, the share of drafts agents sent without opening the editor at all, kept climbing. Release over release it looked like exactly the story everyone wanted: the model learning, agents trusting it more. Nobody was watching the number underneath it.
Edit-depth ratio across five releases
The early signal had already told the whole story by release two. Nobody was reading it.
At its worst, a flywheel that has quietly stopped turning still ships releases, still shows a rising accept rate, and still gets treated as a success story right up until a lagging number finally moves and forces the question nobody asked earlier.
The choice I would take back
Cascade Mobile never logged, release over release, what fraction of agent edits actually made it into that release's training data. That made sense when the pipeline was new and simple enough to trust by eye. It stopped making sense once the pipeline grew several steps and nobody was checking the middle of it.
What I would leave alone: I wouldn't throw out accept rate everywhere. On a genuinely healthy loop it moves together with edit-depth, and watching it costs nothing extra once edit-depth is already being tracked properly.
The lesson: a flywheel metric has to measure the wheel, not the person standing next to it. The moment your number can rise because someone got tired instead of because something got better, you're not measuring the loop anymore.
Now here is the same thing as a story
The short version above is what you'd say defending a metric choice in a planning review. Read this one for how a whole floor's trust quietly outran the thing it was trusting.
Every shift on Cascade Mobile's support floor starts the same way: headsets on, queue open, the first ticket usually a billing question that answers itself in the first line.
Bryndis Halvorsen has worked that floor for three years. When ReplyLine launched, it was, for months, the best thing that had happened to her shift. A draft appeared, she'd read it, fix a word or two, hit send. Ninety-second tickets became twenty-second tickets. She still read every draft closely those first weeks, the way you'd watch a new hire before trusting them with the register.
Then she started reading a little less closely. Then barely at all. By the fourth release, she was hitting send on most drafts the instant they appeared, the way you stop double-checking a coworker who's never once let you down.
Nobody decided, on any one day, to stop checking whether the number underneath was actually moving.
Nobody at Cascade Mobile noticed a single Tuesday where anything changed. There wasn't one. The floor's trust in ReplyLine had been quietly climbing for a year, and the loop it was supposedly built on top of had been quietly flat for almost as long.
Knowledge spark: what is a data flywheel, really?
A setup where the model's own output produces the data that trains its next version. Agents editing drafts should make future drafts need less editing. It's a loop, not a one-time launch, and like any loop, it can spin without actually gripping anything, the same way a wheel can turn while a belt has already slipped off it.
What finally surfaced it wasn't a complaint. It was Marcus Teal, a data analyst three weeks into the job, asking a question in a release review that should have been trivial: "has edit-depth actually gone down since launch?" Nobody in the room had the number ready. When they pulled it, it hadn't moved.
Accept rate and edit-depth both move early. Only one of them is telling you the truth.
The audit that followed found the real break: a retraining pipeline bug had been silently deduplicating agent edits by an old ticket ID scheme, discarding roughly nine out of every ten logged edits from the last three releases before they ever reached the training set. The model had barely trained on real corrections in nine months. Agents had simply gotten used to a tool that had quietly stopped improving.
Accept rate wasn't measuring whether the model got better. It was measuring how long it takes a floor to stop checking something that used to earn it.
All three of these are true at once, and only the middle one is the actual root cause.
Two years earlier, when the pipeline was first built, someone had said, "let's keep the ingestion logic simple, we can always add monitoring once we see how it behaves in practice," and it sounded reasonable, since the whole system was new and nobody had edits to lose yet.
One clock had already rung. Nobody was listening to it, because the other one hadn't rung yet.
Cost per resolved contact, before and after the audit
The lagging cost didn't move until release 5, three releases after edit-depth had already flatlined and could have warned the team.
Rerun the same nine months with an ingestion audit in place from day one: the dropped-edit bug gets caught at release two, when edit-depth first fails to fall as expected. The retrain pipeline gets fixed within a sprint. Edit-depth resumes its slide, cost per contact never spikes, and Bryndis keeps a tool that's actually still earning the trust she's giving it.
What I'd tell myself, watching a new hire ask the one question the whole room should have been asking for a year: the rule was never wrong to trust agents' own behavior as a signal. It was wrong to trust it alone, with nothing underneath checking whether the thing they were trusting was still true.
LEAD, the signal that rang firstNot a dashboard of everything you could track. LEAD is what tells you which single number to actually watch.
L
Link. The business outcome that actually matters.
Cost per resolved contact staying low as ReplyLine genuinely gets better at drafting, not just at being trusted.
Without naming the real outcome, "is it improving" has no anchor at all.
E
Early signal. The thing that moves weeks before the outcome does.
Edit-depth ratio had already flatlined by release two, three releases before cost per contact finally moved.
This is the hardest step, and the one that separates a real flywheel metric from a vanity one.
A
Abuse. How this metric gets gamed.
Draft accept rate climbed for a year without the model actually improving, because agents were just growing more willing to skip checking.
Every metric has a way to be hit without the underlying thing actually being true.
D
Decision. What you'd do differently at each threshold.
No drop in edit-depth across two releases triggers an ingestion audit before the next model ships, full stop.
A metric with no attached action is a dashboard tile, not a gauge.
The recap, one line per letter: link is cost per contact staying low as the model genuinely improves, early signal is edit-depth ratio catching a stall three releases before the cost number did, abuse is accept rate quietly rising while nothing underneath actually got better, and decision is a two-release flat reading triggering a mandatory pipeline audit.
And if you want to be sure it really works, try it somewhere elseSame four letters, a fishing fleet instead of a support floor. Different flip family entirely, the same lying accept rate.
Saltmark Fisheries runs an onboard catch-classification tool that sorts a haul by species and estimated weight from a conveyor-belt camera, so deckhands can log the catch faster. Mapped onto LEAD: link is accurate catch logs without slowing the sort. Early signal is the correction rate, how often a deckhand overrides the model's species call, tracked per trip, not the raw percent of hauls confirmed without any override. Abuse is that confirmed-without-override climbed for months as hauls got bigger, not because the model got better at identifying species, but because deckhands started trusting a busy conveyor belt they no longer had time to watch closely. Decision is auditing the correction-logging step itself the moment confirmed-without-override rises faster than haul volume does. The flip underneath it is a scope one, not an over-trust one: as catch volume grew past what one deckhand could review haul by haul, Callum Bardsey stopped checking each fish and started spot-checking one crate in five, and once he did, the model's actual error rate on rare species had no way to reach him at all.
The same five-step loop, wearing a conveyor belt instead of a headset. This time the break is at step two, not step three.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "don't trust the number that measures behavior, trust the one that measures the thing behavior is supposedly responding to," and stop.
Cost: no budget for a full ingestion-audit pipeline this quarter. Say so honestly, and at minimum log a monthly manual sample of ten edits checked against the live training set.
The model actually got better, for real: if edit-depth genuinely falls alongside rising accept rate, that's the flywheel working as designed, and the honest move is to say so plainly rather than assume every good number is hiding a bug.
Where people run it wrong.
They watch acceptance or usage as if it were quality, instead of watching what the model would need to be doing to earn that acceptance.
They treat a flat lagging metric as proof nothing is wrong, instead of checking whether the early signal underneath it has already broken.
They build the logging step and assume it's also the training step, without ever auditing that edits actually cross that gap.
How to use it live. The moment someone asks how you'd know a flywheel is turning, ask yourself: what number would rise even if the model never actually improved, and what number could only rise if it genuinely did? Name the second one, out loud, and the rest of the answer follows.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust flip: Bryndis stopped checking drafts almost entirely as her trust climbed, right as the model underneath had quietly stopped improving.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Bryndis Halvorsen, a three-year support agent at Cascade Mobile who used ReplyLine's drafts on every ticket.
3 · THE HABIT
What did Bryndis stop doing because the tool worked, early on?
Tap to flip
ANSWER
She stopped reading every draft closely, then stopped checking almost at all by the fourth release, the same way you stop double-checking a coworker who's never let you down.
4 · THE SIGNAL, IN THIS STORY
What's the two-number gap this whole answer turns on?
Tap to flip
ANSWER
Accept rate climbing from 55 to 74 percent while edit-depth stayed flat at 39 to 42 percent the whole time. One measures behavior, the other measures whether the model actually got better.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Never logging, release over release, what fraction of agent edits actually reached the training set, leaving no history to catch the pipeline silently dropping nine out of ten of them.
6 · THE NUMBER
Fill in the blank: edit-depth ratio held at about ___ percent across four straight releases while accept rate climbed nearly 20 points.
Tap to flip
ANSWER
39 to 42 percent, essentially flat.
7 · THE REPLAY
Same nine months, an ingestion audit running from day one. What changes?
Tap to flip
ANSWER
The dropped-edit bug gets caught at release two instead of release five. The pipeline gets fixed within a sprint. Cost per contact never spikes to $5.30, and Bryndis's trust in the tool stays earned instead of assumed.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Saltmark Fisheries' catch-classification tool. The flip is scope: Callum Bardsey went from checking every fish to spot-checking one crate in five once haul volume outgrew what he could review by hand.
Check yourself Score: 0 / 0
Multiple choice
1. According to this answer, what's the strongest sign that ReplyLine's data flywheel had actually stalled?
A. Draft accept rate stopped climbing.
B. CSAT scores dropped sharply in a single week.
C. Edit-depth ratio stayed flat across four releases while accept rate kept rising.
D. Agents started filing complaints about draft quality.
Show hint
Look at the line chart and the gap it shows against accept rate.
Show answer
C. A flat edit-depth ratio against a climbing accept rate is exactly the gap between behavior changing and the model actually improving.
True or false
2. True or false: this answer argues you should stop tracking draft accept rate entirely, since it can be gamed.
True
False
Show hint
Look at "what I would leave alone."
Show answer
False. On a genuinely healthy loop, accept rate and edit-depth move together, and it's still a fine number to watch alongside edit-depth, just never alone.
Fill in the blank
3. Fill in the blank: the retrain pipeline bug meant only about ___ out of every ten logged edits from the last three releases ever reached the training set.
Show hint
Look at the paragraph right after the icon list about "three ways the flywheel number got faked."
Show answer
One. Roughly nine out of ten logged edits were silently discarded before they ever reached the model.
Short answer, where it wouldn't matter
4. Name a place in ReplyLine where a rising accept rate really would mean the model got better, not just that trust outran it.
Show hint
Think about what accept rate looks like on a loop where the ingestion pipeline is actually working.
Show answer
Model answer: On a loop with a working ingestion audit, a rising accept rate that moves together with a falling edit-depth ratio is a genuinely healthy signal, since both numbers are then telling the same true story.
Short answer, apply it yourself
5. Think of a product you use that's supposed to learn from your behavior. What's one number that could look healthy even if it had completely stopped learning?
Show hint
Look for a usage number that could rise for reasons that have nothing to do with quality improving.
Show answer
Model answer: A music app's "skip rate" falling could mean better recommendations, or it could just mean you've given up skipping and are half-listening in the background.
Short answer, work the number
6. If edit-depth had actually been falling in step with the rising accept rate, roughly what value would you expect by release 5, starting from 42 percent at release 1?
Show hint
Accept rate rose about 19 points over 4 releases. A genuinely improving model would show a comparable fall in edit-depth, not a 3-point one.
Show answer
Model answer: Somewhere well under 30 percent, not the 39 percent actually recorded. A near-flat number over the same stretch accept rate moved 19 points is the whole tell.
Before you close the answer
Why this works
Tests whether you'll trust a rising usage number at face value, or go looking for the number underneath it that the usage number is supposed to be a proxy for.
Follow-up traps
"Isn't tracking edit-depth just extra overhead for a number accept rate already implies?" Response: no, because the whole failure here is that they moved in opposite directions for nine months. They imply each other only when the loop is actually healthy.
"What if edit-depth falls but agents are just rubber-stamping worse drafts?" Response: that's exactly what the abuse step is for, pair edit-depth with a periodic quality sample so a falling number can't hide agents disengaging instead of the model improving.
If pressed
The ingestion audit that followed set a specific bar: any release where fewer than 80 percent of qualifying edits from the prior cycle appear in the new training set blocks that release from shipping until someone signs off on the gap.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.