ConceptAdvancedDesigning for Uncertainty & Trust / Feedback loops and data flywheels / #15

Explain why many claimed data flywheels do not actually exist.

AUDIT the scenario: Meridian Leasing Analytics' SwiftScreen, a tenant-screening tool sold on the claim that it "gets smarter with every application"

Interviewer's question: "Explain why many claimed data flywheels do not actually exist." Meridian Leasing Analytics sells SwiftScreen, a tenant-screening tool, to property companies. Naledi Vance is evaluating it for Halden Ridge Properties, a mid-size portfolio manager.

The direct answer
Most claimed flywheels aren't a lie, they're a claim about volume dressed up as a claim about quality. Before you believe one, check three things: a named eval set behind the number, a version and date pin so it can be checked, and proof the correction loop actually closes on a real cadence. Missing any one of those, what you have is more raw data, not a flywheel.
Do this, in order
  1. Ask who ran the claimed number, and whether they had a reason to inflate it.Why: a vendor grading its own homework in its own sales deck is not independent evidence.
  2. Demand the eval set's name, not just the headline score.Why: a score with no stated eval set behind it can't be checked, repeated, or trusted by anyone outside the company that made it.
  3. Demand a version and date pin on any before-and-after number.Why: models get updated quietly, and a number with no pin can't be tied to anything real.
  4. Check whether the correction signal is a real label, not just more raw usage.Why: volume without a ground-truth label doesn't compound, it just repeats whatever bias was already there.
  5. Check whether the retraining loop closes on any real cadence at all.Why: a claim can be true in principle and still false in practice if nobody ever actually presses retrain.
  6. Test the claim yourself, on a slice of your own real data, before it reaches a signature.Why: a number that only exists inside the vendor's own environment is a sales pitch, not evidence.

How to answer this, stage by stage

Nobody is grading whether you sound suspicious. They're grading whether you know exactly which three things to check before you believe a claim like this.

Stage 1
Scope it to one concrete claim
Say it like this
"I'll answer this for SwiftScreen, a tenant-screening tool Meridian Leasing Analytics sells on the claim that it 'gets smarter with every application.'"
Why this works
Turns an abstract skepticism question into one specific claim the interviewer can watch you take apart.
Stage 2
Say your structure out loud
Say it like this
"I'll use AUDIT. Ask who paid for the claim, uncover the eval set, demand the version pin, isolate what's missing, and test it myself."
Why this works
Signals a method before you've made a single accusation about the vendor.
Stage 3
Ask who's actually behind the number
Say it like this
"This claim comes straight from Meridian's own sales deck. Nobody outside the company ran this number."
Why this works
The simplest, most overlooked question in the whole method.
Stage 4
Uncover what's actually being measured
Say it like this
"There's no named eval set anywhere, just a slide claiming accuracy improves over time, with no benchmark I could go check myself."
Why this works
Separates a real, checkable score from a headline number with nothing underneath it.
Stage 5
Demand the version pin
Say it like this
"There's no date and no model version attached to any of their numbers, so there's no way to know which SwiftScreen actually produced which score."
Why this works
Shows you know models change quietly, and an unpinned number can't survive that.
Stage 6
Name what's actually missing
Say it like this
"What's really missing is proof the loop closes: no failure cases shown, no evidence flagged wrong screenings ever get folded back into a retrain."
Why this works
This is the direct answer: the gap isn't dishonesty, it's an unclosed loop.
Stage 7
Say how you'd test it yourself
Say it like this
"Before signing, I'd ask Meridian to run their model against a sample of our own past applicants, and compare it to what we already know happened with those cases."
Why this works
Moves from criticism to a concrete, fair test either side could actually run.
Stage 8
Close on the line that matters
Say it like this
"A believable flywheel claim needs a named eval set, a version pin, and proof the loop closes. Missing any one of those, it's just a bigger pile of data, not a flywheel."
Why this works
Restates the direct answer in one breath, ready for whatever gets pushed on next.

Let's learn

Here's what a lot of "it gets smarter over time" claims actually mean, once you look closely: more data went in, and nobody checked whether any of it was the right kind.

SwiftScreen is a tool property companies pay to screen rental applicants. It reads credit history, past evictions, and income documents, and flags each applicant as low, medium, or high risk for a leasing agent to review.

Meridian's sales deck says SwiftScreen "gets smarter with every application processed," a claim that sounds exactly like a flywheel and gets repeated in nearly every sales call.

Claimed accuracy versus a replicated test on Halden Ridge's own applicants
100% 50% 0% 92% Vendor's claimed number 74% Naledi's own replicated test
Eighteen points is not a rounding error. It's the gap between an eval set nobody can see and a test run on real applicants.

Here's the turn: the claim probably isn't false in the sense of being made up. SwiftScreen genuinely does process more applications every month than it did last year. But more applications processed is not the same thing as more correct labels learned from, and Meridian's claim quietly treats the two as one thing.

Hand sketched labeled parts diagram titled The vendor's claim, taken apart. Center document icon labeled SwiftScreen Claim, with five callouts: who paid, what eval set, which version, what's missing, test it yourself.
Five questions. The vendor's sales deck answers none of them.
The decision that mattered Treat "processes more applications" and "learns from more correct labels" as two separate claims, and ask for evidence of the second one specifically, not just a bigger version of the first.

At its worst: a property company signs a multi-year contract on the strength of "it gets smarter," the false-decline rate for legitimate applicants never actually improves, and eighteen months later nobody can even say what changed, because there was never a version pin to compare against in the first place.

Hand sketched icon list titled What a believable claim needs. Four items: a document icon labeled a named eval set, a gauge icon labeled a version and date pin, a question mark box icon labeled failure cases shown, a scale icon labeled independent replication.
SwiftScreen's sales deck has none of these four. Most vendor claims that fold under a second look are missing at least two.

What I would leave alone: the idea that usage data can genuinely improve a model isn't wrong, it's just incomplete. A real flywheel is entirely possible here. The problem isn't the concept, it's this specific unverified claim about it.

The lesson: "gets smarter with usage" and "gets smarter with usage that includes a real correction signal, retrained on a real schedule, and checked against a real eval set" are two different sentences. Most sales decks only ever say the first one.

Now here is the same thing as a story

The short version above is what you'd say in the vendor selection meeting. Read this one for how the eighteen-point gap actually got found.

Naledi Vance can smell a padded sales deck before the second slide. Six years of picking vendors for Halden Ridge Properties will do that.

For most of a quarter, SwiftScreen's pitch looked solid. The demo was clean, the sample screenings matched what her own team would have flagged, and the "gets smarter with every application" line sat right there on slide four, exactly the kind of thing that makes a busy operations director stop asking questions.

She was close to signing. Then a peer at another mid-size property company mentioned, almost as an aside at an industry lunch, that they'd been using SwiftScreen for over a year, and their false-decline rate on legitimate applicants hadn't moved at all in that time, despite the same "gets smarter" line in their renewal pitch.

That one comment is what made Naledi go back and actually ask the three questions she'd skipped the first time.

Hand sketched flow diagram titled Where the loop actually breaks. Four steps: usage happens, correction question mark circled in red, retrain scheduled question mark, version updated question mark.
Meridian could answer the first box. Nobody at the company could answer the next three with anything more than "probably."

She asked for the eval set behind the "gets smarter" claim. There wasn't a named one, just an internal dashboard nobody outside Meridian had ever seen. She asked which model version the 92 percent figure came from, and on what date. Nobody could say. She asked whether flagged wrong screenings actually got used to retrain the model, and on what schedule. The honest answer, once she pushed past the sales rep to an actual engineer, was "it happens sometimes, when someone has time."

The claim wasn't a lie. It was a sentence about volume, wearing the confidence of a sentence about quality.

So she ran her own test. Halden Ridge had eighteen months of past applicant outcomes on file, cases where they already knew who'd actually paid rent reliably and who hadn't. She fed 200 of those cases through SwiftScreen's current version and compared its calls against what had actually happened.

Hand sketched timeline titled SwiftScreen's actual update history. Four milestones: launch version 1 ships, 16 months pass no public changelog, v2 ships quietly circled in red no version pin no notice, claim persists same slide unchanged.
Sixteen months between two things that actually looked different, and no way for a customer to have noticed either one.

74 percent. Not a disaster, but eighteen points below the number on slide four, and nowhere near the kind of gap a genuinely improving model should have produced over sixteen months of real usage.

Here's the replay that matters: with the eval set, version pin, and retrain-cadence questions asked upfront, Naledi never gets to the point of nearly signing on a claim she couldn't check. She runs the 200-case test in the evaluation phase instead of after a near miss at an industry lunch, and either SwiftScreen's real number holds up, or Halden Ridge walks away from the contract eighteen months and one renewal cycle earlier.

The old habit, in her own team, was treating a clean demo and a confident slide as evidence. It took a stranger's offhand comment about a stalled false-decline rate to see that a flywheel claim needs the same scrutiny as any other unverified number, no matter how good the pitch sounds.

AUDIT, one letter at a timeNot a fraud investigation. AUDIT is what tells you a flywheel claim and a volume claim are not the same sentence.

A
Ask who paid for it.
Meridian's own sales deck, with no independent party involved in producing the "gets smarter" claim.
The simplest question, and the one most buyers skip when the demo looks good.
U
Uncover the eval set.
No named benchmark anywhere, just an internal dashboard nobody outside the company has ever seen.
A score with no stated eval set behind it can't be checked by anyone but the person who made the claim.
D
Demand the version pin.
No date, no model version attached to the 92 percent figure, so it can't be tied to anything reproducible.
Models update quietly, and an unpinned number is unfalsifiable by design, not by accident.
I
Isolate what's missing.
No failure cases shown, and no evidence that flagged wrong screenings actually get folded back into a real retrain on any schedule.
This is the hardest step and the direct answer: the gap is a loop that never closes, not a lie that was told.
Hand sketched comparison diagram titled Real flywheel versus claimed flywheel. Left panel, a document icon labeled Real, caption labeled corrections versioned replicable. Right panel, a question mark box icon labeled Claimed, caption more raw usage no version pin.
Both panels describe a model touched by lots of usage. Only one of them describes a model that actually learned something from it.
Months since SwiftScreen's last verifiable update, by quarter asked
15mo 7mo 0 Q1 Q2 Q3 Q4 14 months
A real retrain cadence would reset this line every few months. Instead it just climbs, quarter after quarter, with nothing ever resetting it.
T
Test it yourself.
200 of Halden Ridge's own past applicants, run through SwiftScreen's current version, checked against what actually happened.
Replication on your own data is the only test a vendor can't quietly control.
Hand sketched quadrant titled Which claims are worth believing. Axes eval set named and version pinned. SwiftScreen and a typical vendor claim sit low on both, bottom left. A believable claim sits high on both, top right.
Most vendor claims live in the same crowded corner. Believing one means asking it to move.

The recap, one line per letter: ask is that this claim comes only from Meridian's own sales deck, uncover is that there's no named eval set behind it, demand is that there's no version pin on the 92 percent figure, isolate is that no failure cases or retrain cadence are ever shown, and test is Naledi's own 200-case replication that found the real number.

And if you want to be sure it really works, try it somewhere elseSame five letters, a résumé-screening vendor instead of a tenant-screening one. A completely different hiring context, the same unclosed loop underneath.

BrightArc Talent sells a résumé-screening tool that recruiting teams use to rank applicants. Conrad Isby, an HR operations lead, is evaluating it for his company's hiring pipeline, and BrightArc's pitch deck says the same kind of thing SwiftScreen's does: it "improves with every résumé it reviews."

Mapped onto AUDIT: ask is that the claim comes entirely from BrightArc's own case studies, with no independent hiring outcome data behind it. Uncover is that there's no named benchmark, just a chart of BrightArc's own internal accuracy metric with no definition of what counts as a correct call. Demand is that no version or date accompanies any of the numbers in the deck, so a claim from two years ago and one from last month are indistinguishable. Isolate is that BrightArc shows no examples of a résumé it initially mis-ranked and then corrected, and no description of how often, if ever, hiring-manager overrides get folded back into training. Test is Conrad running the tool against 150 of his own company's past hires and rejections, where the actual outcome, whether that person succeeded in the role, is already known.

Swap the trigger and it still runs.
Speed: an interviewer caps you at thirty seconds. Say "check the eval set, the version pin, and whether the loop actually closes, most claimed flywheels fail on at least one," and stop.
Cost: if there's no time to run a full replication test before a purchase decision, at minimum demand the version pin and eval set in writing, since that alone turns an unfalsifiable claim into a checkable one.
The model gets better, for real: if a vendor's model genuinely does improve, that's still not a reason to skip the check, a real improvement should hold up to a version pin and an outside replication test without flinching.

Where people run it wrong.
They accept a clean product demo as proof of a claim about long-term learning, when a demo only proves the model works on the cases the vendor chose to show.
They confuse "processes more data" with "learns from more correct labels," treating the first as if it implies the second automatically.
They never ask for a version pin, so an old number and a current number end up sitting in the same sentence with no way to tell them apart.

How to use it live. When someone hands you a flywheel claim, ask yourself one question first: could I go find the eval set, the version, and the failure cases behind this number right now. If the honest answer is no, treat the claim as unverified, not as false, and say exactly what would turn it into a real one.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "why do many claimed data flywheels not actually exist"?
Tap to flip
ANSWER
AUDIT: ask who paid, uncover the eval set, demand the version pin, isolate what's missing, test it yourself. Built for judging a claim or artifact.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Naledi Vance, an operations director at Halden Ridge Properties, evaluating Meridian Leasing Analytics' SwiftScreen tool.
3 · THE CLAIM
What exact claim is being tested in this answer?
Tap to flip
ANSWER
That SwiftScreen "gets smarter with every application processed," a sentence about volume dressed up as a sentence about quality.
4 · WHAT'S MISSING
What did Naledi find missing when she actually asked?
Tap to flip
ANSWER
No named eval set, no version or date pin on the 92 percent figure, and no proof the correction loop closes on any real schedule.
5 · THE OLD HABIT
What habit almost let Naledi sign the contract?
Tap to flip
ANSWER
Treating a clean product demo and a confident sales slide as if they were evidence, instead of asking for a checkable eval set and version pin.
6 · THE NUMBER
Fill in the blank: Meridian claimed 92 percent accuracy, but Naledi's own replicated test on 200 past applicants found ___ percent.
Tap to flip
ANSWER
74 percent. Eighteen points below the vendor's number, and nowhere near what sixteen months of real improvement should have produced.
7 · THE REPLAY
Same evaluation, questions asked upfront. What changes?
Tap to flip
ANSWER
Naledi runs the 200-case replication test during evaluation, not after a near miss at an industry lunch, and Halden Ridge decides a full renewal cycle earlier.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's missing there?
Tap to flip
ANSWER
BrightArc Talent's résumé-screening tool. Same gaps: no named benchmark, no version pin, and no shown examples of a correction actually happening.

Check yourself Score: 0 / 0

Short answer, name the reversal
1. What habit does this answer say Naledi's team should give up, and why did that habit make sense before?
Show hint
Look at "the old habit" near the end of the story.
Show answer
Model answer: Treating a clean demo and a confident slide as evidence. It made sense because the demo genuinely matched what her own team would have flagged, which felt like proof until she checked further.
Multiple choice
2. What's the core difference between a real flywheel and SwiftScreen's claimed one?
  • A. A real flywheel uses a more expensive model.
  • B. A real flywheel is built from labeled corrections on a real cadence, not just more raw usage volume.
  • C. A real flywheel never makes mistakes.
  • D. A real flywheel doesn't need a human reviewer at all.
Show hint
Look at the "real versus claimed" comparison diagram.
Show answer
B. Volume alone doesn't compound. A labeled correction, checked and retrained on a schedule, is what actually makes a flywheel real.
True or false
3. True or false: Meridian's "gets smarter" claim was most likely an outright fabrication.
  • True
  • False
Show hint
Look at the highlight line about volume versus quality.
Show answer
False. The claim wasn't a lie, it was a sentence about volume wearing the confidence of a sentence about quality, which is a subtler and more common problem.
Fill in the blank
4. Fill in the blank: SwiftScreen went about ___ months between its version 1 launch and a quiet version 2 update, with no changelog either time.
Show hint
Look at the update-history timeline diagram.
Show answer
16 months. Long enough that a real improvement should have shown up clearly, and short enough that nobody would have blamed the vendor for not yet retraining.
Short answer, apply it yourself
5. Pick a product or vendor claim you've encountered yourself. What's one number in it you'd want a version pin and an eval set for before believing it?
Show hint
Think of a "up to X percent faster" or "improves over time" claim you've seen in a pitch or an ad.
Show answer
Model answer: Many people can point to an "up to 40 percent more accurate" claim in an ad with no stated comparison point, no date, and no way to check it against their own use case.
Short answer, where it wouldn't matter
6. Name a situation where a vendor's "gets smarter over time" claim would actually be easy to verify.
Show hint
Think about what would need to be true for the claim to already be checkable.
Show answer
Model answer: If the vendor published dated, versioned scores on a named public benchmark every quarter, a buyer could track the real trend themselves without needing to run a replication test at all.
Before you close the answer
Why this works
Tests whether you can separate a genuinely improving system from a system that just has more usage sitting in a database. Most candidates either accept the claim or reject it outright, instead of naming exactly what would prove it.
Follow-up traps
"Isn't demanding a version pin just going to slow down every vendor negotiation?" Response: it costs one email, and a vendor unwilling to provide it has already answered the real question.

"What if the vendor's internal eval set really is good, just not public?" Response: then a replication test on your own data should confirm it easily, which is exactly the check that costs the vendor nothing if the claim is true.
If pressed
Naledi's 200-case replication used a stratified sample, not a random one, weighted toward the applicant types Halden Ridge actually sees most, since a random sample from a different market mix would have understated the real gap.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more