Explain why many claimed data flywheels do not actually exist.
Interviewer's question: "Explain why many claimed data flywheels do not actually exist." Meridian Leasing Analytics sells SwiftScreen, a tenant-screening tool, to property companies. Naledi Vance is evaluating it for Halden Ridge Properties, a mid-size portfolio manager.
- Ask who ran the claimed number, and whether they had a reason to inflate it.Why: a vendor grading its own homework in its own sales deck is not independent evidence.
- Demand the eval set's name, not just the headline score.Why: a score with no stated eval set behind it can't be checked, repeated, or trusted by anyone outside the company that made it.
- Demand a version and date pin on any before-and-after number.Why: models get updated quietly, and a number with no pin can't be tied to anything real.
- Check whether the correction signal is a real label, not just more raw usage.Why: volume without a ground-truth label doesn't compound, it just repeats whatever bias was already there.
- Check whether the retraining loop closes on any real cadence at all.Why: a claim can be true in principle and still false in practice if nobody ever actually presses retrain.
- Test the claim yourself, on a slice of your own real data, before it reaches a signature.Why: a number that only exists inside the vendor's own environment is a sales pitch, not evidence.
How to answer this, stage by stage
Nobody is grading whether you sound suspicious. They're grading whether you know exactly which three things to check before you believe a claim like this.
Let's learn
Here's what a lot of "it gets smarter over time" claims actually mean, once you look closely: more data went in, and nobody checked whether any of it was the right kind.
SwiftScreen is a tool property companies pay to screen rental applicants. It reads credit history, past evictions, and income documents, and flags each applicant as low, medium, or high risk for a leasing agent to review.
Meridian's sales deck says SwiftScreen "gets smarter with every application processed," a claim that sounds exactly like a flywheel and gets repeated in nearly every sales call.
Here's the turn: the claim probably isn't false in the sense of being made up. SwiftScreen genuinely does process more applications every month than it did last year. But more applications processed is not the same thing as more correct labels learned from, and Meridian's claim quietly treats the two as one thing.
At its worst: a property company signs a multi-year contract on the strength of "it gets smarter," the false-decline rate for legitimate applicants never actually improves, and eighteen months later nobody can even say what changed, because there was never a version pin to compare against in the first place.
What I would leave alone: the idea that usage data can genuinely improve a model isn't wrong, it's just incomplete. A real flywheel is entirely possible here. The problem isn't the concept, it's this specific unverified claim about it.
The lesson: "gets smarter with usage" and "gets smarter with usage that includes a real correction signal, retrained on a real schedule, and checked against a real eval set" are two different sentences. Most sales decks only ever say the first one.
Now here is the same thing as a story
The short version above is what you'd say in the vendor selection meeting. Read this one for how the eighteen-point gap actually got found.
Naledi Vance can smell a padded sales deck before the second slide. Six years of picking vendors for Halden Ridge Properties will do that.
For most of a quarter, SwiftScreen's pitch looked solid. The demo was clean, the sample screenings matched what her own team would have flagged, and the "gets smarter with every application" line sat right there on slide four, exactly the kind of thing that makes a busy operations director stop asking questions.
She was close to signing. Then a peer at another mid-size property company mentioned, almost as an aside at an industry lunch, that they'd been using SwiftScreen for over a year, and their false-decline rate on legitimate applicants hadn't moved at all in that time, despite the same "gets smarter" line in their renewal pitch.
That one comment is what made Naledi go back and actually ask the three questions she'd skipped the first time.
She asked for the eval set behind the "gets smarter" claim. There wasn't a named one, just an internal dashboard nobody outside Meridian had ever seen. She asked which model version the 92 percent figure came from, and on what date. Nobody could say. She asked whether flagged wrong screenings actually got used to retrain the model, and on what schedule. The honest answer, once she pushed past the sales rep to an actual engineer, was "it happens sometimes, when someone has time."
So she ran her own test. Halden Ridge had eighteen months of past applicant outcomes on file, cases where they already knew who'd actually paid rent reliably and who hadn't. She fed 200 of those cases through SwiftScreen's current version and compared its calls against what had actually happened.
74 percent. Not a disaster, but eighteen points below the number on slide four, and nowhere near the kind of gap a genuinely improving model should have produced over sixteen months of real usage.
Here's the replay that matters: with the eval set, version pin, and retrain-cadence questions asked upfront, Naledi never gets to the point of nearly signing on a claim she couldn't check. She runs the 200-case test in the evaluation phase instead of after a near miss at an industry lunch, and either SwiftScreen's real number holds up, or Halden Ridge walks away from the contract eighteen months and one renewal cycle earlier.
The old habit, in her own team, was treating a clean demo and a confident slide as evidence. It took a stranger's offhand comment about a stalled false-decline rate to see that a flywheel claim needs the same scrutiny as any other unverified number, no matter how good the pitch sounds.
AUDIT, one letter at a timeNot a fraud investigation. AUDIT is what tells you a flywheel claim and a volume claim are not the same sentence.
The recap, one line per letter: ask is that this claim comes only from Meridian's own sales deck, uncover is that there's no named eval set behind it, demand is that there's no version pin on the 92 percent figure, isolate is that no failure cases or retrain cadence are ever shown, and test is Naledi's own 200-case replication that found the real number.
And if you want to be sure it really works, try it somewhere elseSame five letters, a résumé-screening vendor instead of a tenant-screening one. A completely different hiring context, the same unclosed loop underneath.
BrightArc Talent sells a résumé-screening tool that recruiting teams use to rank applicants. Conrad Isby, an HR operations lead, is evaluating it for his company's hiring pipeline, and BrightArc's pitch deck says the same kind of thing SwiftScreen's does: it "improves with every résumé it reviews."
Mapped onto AUDIT: ask is that the claim comes entirely from BrightArc's own case studies, with no independent hiring outcome data behind it. Uncover is that there's no named benchmark, just a chart of BrightArc's own internal accuracy metric with no definition of what counts as a correct call. Demand is that no version or date accompanies any of the numbers in the deck, so a claim from two years ago and one from last month are indistinguishable. Isolate is that BrightArc shows no examples of a résumé it initially mis-ranked and then corrected, and no description of how often, if ever, hiring-manager overrides get folded back into training. Test is Conrad running the tool against 150 of his own company's past hires and rejections, where the actual outcome, whether that person succeeded in the role, is already known.
Swap the trigger and it still runs.
Speed: an interviewer caps you at thirty seconds. Say "check the eval set, the version pin, and whether the loop actually closes, most claimed flywheels fail on at least one," and stop.
Cost: if there's no time to run a full replication test before a purchase decision, at minimum demand the version pin and eval set in writing, since that alone turns an unfalsifiable claim into a checkable one.
The model gets better, for real: if a vendor's model genuinely does improve, that's still not a reason to skip the check, a real improvement should hold up to a version pin and an outside replication test without flinching.
Where people run it wrong.
They accept a clean product demo as proof of a claim about long-term learning, when a demo only proves the model works on the cases the vendor chose to show.
They confuse "processes more data" with "learns from more correct labels," treating the first as if it implies the second automatically.
They never ask for a version pin, so an old number and a current number end up sitting in the same sentence with no way to tell them apart.
How to use it live. When someone hands you a flywheel claim, ask yourself one question first: could I go find the eval set, the version, and the failure cases behind this number right now. If the honest answer is no, treat the claim as unverified, not as false, and say exactly what would turn it into a real one.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the vendor's internal eval set really is good, just not public?" Response: then a replication test on your own data should confirm it easily, which is exactly the check that costs the vendor nothing if the claim is true.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Feedback loops and data flywheels
- #1 Design the feedback mechanism for an AI feature where users rarely click thumbs down.
- #2 Explain the difference between explicit and implicit feedback signals.
- #3 What implicit signals tell you an output was bad?
- #4 How do you avoid a feedback loop that only captures complaints?
- #5 Describe how you would turn user edits into a quality signal.
- #6 What is the latency between collecting feedback and improving the product, and how do you shorten it?