ConceptIntermediateAI Opportunity & Model Strategy / Data strategy as product strategy / #21
What data would you stop collecting, and why?
ORDERthe ten-second sample that stopped earning its keep two years ago
Two years ago, keeping a phone's diagnostics running cost Hallowick Wireless about $40,000 a month. Hallowick Wireless is a regional carrier. SignalScope is its tool that predicts which devices are heading toward failure so a technician can reach out before the phone actually dies. Deniz Okafor is the AI PM who has to decide, every quarter, which of the dozens of signals it collects still earn their keep.
The direct answer
Stop collecting the ten-second signal-strength samples and the dense cell-tower handoff logs. Both were set to maximum frequency when the model was new and needed everything it could get. Two years later, model accuracy has plateaued while the fields have tripled the storage bill and are the leading cause of battery-drain complaints. Keep the rare crash-dump snapshots and the fault-linked thermal readings; those are still improving predictions and can't be reconstructed once a failure has already happened.
Do this, in order
For each collected field, check whether it has moved model accuracy in the last two quarters.Why: without this, "should we keep collecting X" is a guess dressed up as a policy.
Ask which fields are cheap to resume if you're wrong, and which aren't.Why: a rare crash dump can't be recreated after the fact. A frequent signal sample can always be turned back up.
Check what else depends on the field before cutting it.Why: some data feeds a labeling or eval step downstream, and cutting it blind breaks something you didn't mean to touch.
Run a cheap holdout test before cutting anything company-wide.Why: turning a field off for a sample of devices for two weeks tells you the real cost of stopping, before you commit to it everywhere.
Rank the cuts by storage and battery cost saved per point of accuracy risked.Why: the goal isn't collecting less for its own sake, it's collecting less of what was never paying for itself.
How to answer this, stage by stage
Nobody is scoring whether you can name one field to cut. They're scoring whether you can tell the difference between data that's still working and data that's just still there.
Stage 1
Scope it to one fleet, one decision
Say it like this
"I'll ground this in SignalScope at Hallowick Wireless, and the actual quarter where the diagnostic storage bill tripled while accuracy barely moved."
Why this works
Stops the answer from becoming a generic essay about data minimalism.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as ORDER. Outcome, what all the data is competing to serve. Reversibility, what's hardest to get back once it's gone. Dependency, what breaks downstream if I cut blind. Evidence, what a cheap holdout would tell me. Rank, the actual list."
Why this works
Signals a repeatable method instead of a gut call about which fields "feel unnecessary."
Stage 3
Reframe: not "delete data," but "which field lost its reason to exist"
Say it like this
"This isn't about being frugal with data for its own sake. Every field earned its place once, when the model needed it. The question is which ones stopped earning that place and nobody ever went back to check."
Why this works
This is where a strong answer separates from "cut whatever seems unused."
Stage 4
Give the ranked list
Say it like this
"Cut ten-second signal-strength sampling back to sixty seconds. Thin the cell-tower handoff logs, since that signal already did its job in year one. Keep crash-dump snapshots, rare and irreplaceable. Keep fault-linked thermal reads, still moving accuracy every quarter."
Why this works
This is the direct answer, applied to four real candidates instead of stated as an abstract rule.
Stage 5
Prove it with the audit
Say it like this
"Failure-prediction accuracy went from 81 to 89 percent in year one, when the dense signal sampling was new. In year two, doubling down on the same signal only moved it from 89.0 to 89.1, while the monthly storage bill went from 40,000 dollars to 126,000. That's the whole case in two numbers."
Why this works
Compresses the entire argument into the one comparison a written cut-off rule would have caught a year earlier.
Stage 6
Name what's hard to undo
Say it like this
"Crash-dump snapshots only exist because a device happened to fail while diagnostics were on. If we stop collecting those and later need them for a new fault pattern, that data is just gone, there's no going back and recording last month's crash."
Why this works
Shows the reasoning isn't "less data always," it's reversibility-aware.
Stage 7
Say what you'd never cut
Say it like this
"I wouldn't touch fault-linked thermal readings even under budget pressure. That's the one field still moving prediction accuracy every single quarter."
Why this works
Shows judgment instead of a blanket instinct to cut everything under a cost review.
Stage 8
Name the AI-specific reasoning, then close
Say it like this
"The honest reason this isn't just a storage-cost exercise is that a field's value to a model changes as the model matures. Data that was essential at launch can plateau, and data nobody thought mattered can quietly become the thing catching a new failure pattern. I'd cut the signal-strength sampling back, and I'd keep the rare crash dumps no matter what the storage bill says."
Why this works
Closes with the real judgment call and restates the direct answer in one breath.
Let's learn
Here is what happens when a setting that made sense on day one never gets revisited as the thing it was feeding grows up.
Before SignalScope, a customer only found out their phone was failing when it actually stopped working, and a technician diagnosed it from scratch in the store, about twenty minutes per visit with no advance warning at all. With SignalScope, the model watches a running stream of diagnostics and flags a device days before it's likely to fail, so a proactive repair or replacement offer can go out first.
Four kinds of data, collected the same way today as they were the day the model launched.
Here's the turn: the extra data isn't the problem by itself. The problem is that nobody ever asked, field by field, whether each one was still buying anything. Two of these four stopped moving the model a year ago and nobody noticed, because the storage bill climbing quietly doesn't set off any alarm the way a crashing model would.
Failure-prediction accuracy versus monthly diagnostic storage cost
Accuracy did almost all its climbing in year one. The storage bill kept climbing anyway.
At its worst, a diagnostic pipeline can quietly become a tax: real money spent, real batteries drained, on data that stopped teaching the model anything new a year before anyone checked.
The choice I would take back
When SignalScope launched, the team set signal-strength sampling to every ten seconds and tower handoffs to full precision, because the model was new and needed every signal it could get. That made sense in month one. It stopped making sense once accuracy plateaued and nobody ever revisited the default.
What I would leave alone: I wouldn't touch crash-dump snapshots or fault-linked thermal readings, even though they're a much smaller share of total data volume. Both are still moving accuracy, and one of them can't be recreated after the fact.
The lesson: a data field's value isn't fixed at the moment you start collecting it. It changes as the model matures, and if nobody ever checks back, you end up paying full price for data quietly worth less than it used to be.
Now here is the same thing as a story
The short version above is what you'd say defending a data cut in a cost review. Read this one for how a battery complaint turned into a hidden bias nobody was looking for.
The network operations floor at Hallowick Wireless never really empties out, but it thins out after midnight, down to two people and a wall of dashboards nobody's watching too closely.
Foday Kanneh has worked field diagnostics for five years. He's the one dispatched when a premium account calls in about a phone that's suddenly always hot and always at eight percent battery by lunchtime. For the first year of SignalScope, he'd check the phone, confirm the diagnostic pinging was part of the problem, and file a ticket asking the platform team to dial it back for that device. Slow, but it worked, and the platform team usually got to it within a week.
The second step is why you can't cut a field blind. Some raw signal is quietly load-bearing for how faults get labeled in the first place.
As Hallowick's premium tier grew, so did the ticket queue. By month sixteen, Foday was seeing three or four battery complaints a week, and the platform team's backlog for "dial back this device's pinging" stretched past a month. So he stopped filing tickets. He wrote his own small script, run from his own laptop, that flipped a hidden diagnostic-frequency flag directly on any premium device he visited, no ticket, no review, no record anywhere except his own local log.
Knowledge spark: why does a workaround like this quietly bias a model?
Once premium devices are collecting less frequent data than everyone else, the training set stops representing them the way it used to. The model doesn't know that happened. It just sees a fleet where premium devices look, on paper, a little different than they used to, for no reason the labels can explain.
Nobody at Hallowick decided to let this happen. It built up quietly, over eight months, one battery complaint at a time, until roughly five thousand devices, nearly all of them premium accounts, were collecting diagnostics at a fraction of the standard rate, invisibly, with no flag anywhere saying so.
The two fields in the bottom-left corner are the ones actually worth cutting back. The other two are the ones worth keeping no matter the bill.
The gap surfaced when a network engineer, reviewing the quarterly infrastructure spend, asked a question that should have been routine: why had diagnostic storage tripled with no matching jump in prediction accuracy anywhere on the roadmap. Pulling the numbers open, the team found both problems at once, a plateaued signal nobody had thinned back, and a private workaround nobody had reviewed.
The storage bill wasn't the real cost. The real cost was a slice of the fleet quietly going dark to the model, one battery complaint at a time, with nobody choosing for it to happen.
Nobody decided on any single day to keep collecting at full frequency forever. The setting from launch just never got a second look.
What's eating the diagnostic storage bill
Signal-strength sampling alone is more than half the bill, and the field accuracy stopped needing a year ago.
Rerun the same two years with a standing per-field review in place: signal-strength sampling gets thinned to sixty seconds the quarter accuracy plateaus, saving most of that $71,000 a month. Foday's battery tickets get a real fast-track instead of a month-long queue, so the workaround never gets built, and no slice of the fleet quietly disappears from the training data. Crash dumps and thermal reads keep collecting at full rate, since both are still worth every dollar they cost.
What I'd tell myself, reading Foday's private script for the first time: the platform team's backlog was never a small thing to shrug off. It was the reason a careful, well-meaning technician built exactly the kind of shadow system a shared review process exists to prevent.
ORDER, the rank that separates cheap data from earned dataNot a rule to hoard everything or delete indiscriminately. ORDER is what tells you exactly which field to touch first.
O
Outcome. What all the candidate fields compete to serve.
Catching device failure early, without draining batteries or paying for signal the model no longer needs.
Without a shared outcome, deciding what to cut is just a guess about what feels excessive.
R
Reversibility. What's hardest to get back.
A signal-strength sample can always be turned back up next quarter. A crash dump from a device that already failed cannot be recreated.
This is the hardest step, and the one a blanket "cut anything unused" rule always skips.
D
Dependency. What unblocks what.
Raw signal feeds how faults get labeled, which feeds the eval set the model is graded against. Cut the wrong field and the eval set breaks quietly too.
Some fields aren't just inputs, they're the reason a downstream step exists at all.
E
Evidence. What's cheap to learn first.
A two-week holdout, sampling ten seconds down to sixty on a slice of devices, would have shown the real cost of thinning that field before committing fleet-wide.
Cheap evidence gathered first is what makes cutting a field a decision instead of a hunch.
R
Rank. State the order, defend the top.
Thin signal-strength sampling first, it's the biggest cost and the flattest lift. Thin tower handoffs second. Leave crash dumps and thermal reads alone.
The ranking follows directly from reversibility and current lift, not from which field looks biggest on a storage report.
The recap, one line per letter: outcome is catching failure early without wasting battery or storage, reversibility is a signal sample you can always turn back up against a crash dump you can never recreate, dependency is raw signal quietly feeding the eval set downstream, evidence is a two-week holdout proving the real cost before a fleet-wide cut, and rank is thinning the flattest, cheapest-to-reverse fields first while leaving the still-climbing ones alone.
And if you want to be sure it really works, try it somewhere elseSame five letters, a farm's soil-sensor network instead of a phone fleet. Different flip family entirely, the same over-collected default.
Deep Furrow Farms runs an irrigation-recommendation tool that reads soil-moisture sensors across several hundred acres and tells a grower when and where to water. Mapped onto ORDER: outcome is using water efficiently without under-watering a field that actually needs it. Reversibility says a sensor reading every fifteen minutes can always be turned back up if a field looks risky, while a full season of a newly planted crop's moisture curve cannot be replayed once it's gone. Dependency says the recommendation model needs at least one full season of a field's data before it trusts its own advice there. Evidence favors a one-field pilot before changing sensor frequency everywhere. But the flip here is a substitution one, not a workaround: once Deep Furrow's data-processing fees rose with sensor count, Agnes Thorne started paying for dense readings only on her highest-value fields and let her lowest-margin fields run on sparse, once-a-day readings, exactly the fields where a slow leak in an irrigation line was most likely to go unnoticed.
The same four branches sort a phone fleet's diagnostics and a farm's soil sensors exactly the same way.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "cut the field that's stopped moving accuracy and is cheap to bring back, keep the one that's rare and can't be recreated, that's the whole rule," and stop.
Cost: no time to build a full per-field review this quarter. Say so honestly, and start with the single biggest line item on the storage bill.
The field turns out to still matter, for real: if a holdout test shows accuracy actually drops when a field is thinned, that's a legitimate reason to keep collecting it, not a failure of the review.
Where people run it wrong.
They treat "we've always collected this" as a reason to keep collecting it, instead of checking whether it's still buying anything.
They cut a field based on volume or cost alone, without checking whether it's cheap or impossible to get back later.
They let a well-meaning workaround around a slow process quietly bias who's actually represented in the data.
How to use it live. The moment an interviewer asks what you'd stop collecting, ask yourself: has this actually moved anything lately, and could I get it back if I turned out to be wrong? Answer those two, and the ranking writes itself.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Workaround flip: Foday built his own undocumented script to quietly throttle diagnostics on premium devices, rather than wait out a month-long ticket queue.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Foday Kanneh, a five-year field technician at Hallowick Wireless who handled battery-drain complaints on premium accounts.
3 · THE HABIT
What did Foday stop doing once the ticket queue grew too long?
Tap to flip
ANSWER
He stopped filing tickets asking the platform team to dial back a device's diagnostic frequency, and started flipping the setting himself with a private script instead.
4 · THE SWITCH, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Waiting for the official ticket process versus quietly running a private, unreviewed toggle. Once the queue outgrew what he could wait on, there was no in-between.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Setting signal-strength sampling to maximum frequency at launch and never scheduling a check-in on whether that default still made sense once the model matured.
6 · THE NUMBER
Fill in the blank: prediction accuracy moved from 81 to 89 percent in year one, then only to ___ percent in year two, despite storage costs tripling.
Tap to flip
ANSWER
89.1 percent, essentially flat.
7 · THE REPLAY
Same two years, a standing per-field review in place from the start. What changes?
Tap to flip
ANSWER
Signal-strength sampling gets thinned the quarter accuracy plateaus, saving most of the $71,000 monthly cost. Battery tickets get fast-tracked, so Foday's private script never gets built, and no slice of the fleet quietly vanishes from the training data.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Deep Furrow Farms' irrigation-recommendation tool. The flip is substitution: Agnes Thorne rationed dense sensor readings toward her highest-value fields once fees rose, leaving her riskiest low-margin fields the least monitored.
Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: monthly diagnostic storage cost rose from $40,000 at launch to about $___ two years later.
Show hint
Look at the line chart comparing accuracy and storage cost.
Show answer
$126,000. Roughly tripling, while accuracy moved less than half a point over the same year.
Multiple choice
2. According to this answer, what actually determines whether a data field should be cut?
A. How much it costs to store, regardless of anything else.
B. Whether it's still moving model accuracy, and how hard it would be to get back if you're wrong.
C. Whether the field has been collected for more than a year.
D. Whether any single engineer requests it be removed.
Show hint
Look at the outcome and reversibility steps.
Show answer
B. Cost alone doesn't decide it. Plateaued lift plus easy reversibility is what actually makes a field safe to cut.
True or false
3. True or false: this answer recommends cutting crash-dump snapshots, since they're a small share of total data volume.
True
False
Show hint
Look at "what I would leave alone" and the quadrant diagram.
Show answer
False. Crash dumps are rare, still lifting accuracy, and impossible to reconstruct once a device has already failed, so they're kept regardless of their small size.
Short answer, apply it yourself
4. Think of a report or dashboard your team still updates out of habit. What would it take to check whether it's still earning its place?
Show hint
Think about the last time anyone actually acted on that report's numbers.
Show answer
Model answer: Ask when a decision last actually changed because of that report. If nobody can name one in the last two quarters, it's a candidate to cut or automate down to a glance.
Short answer, work the number
5. If thinning signal-strength sampling saves most of its $71,000 monthly cost, roughly how much would Hallowick save in a year?
Show hint
Multiply a monthly savings estimate by twelve.
Show answer
Model answer: Somewhere near $700,000 to $800,000 a year, assuming most but not all of the $71,000 monthly cost goes away once sampling is thinned rather than eliminated.
Short answer, where it wouldn't matter
6. Name a case where you'd keep collecting a field at full frequency even under real budget pressure, and say why.
Show hint
Look at "what I would leave alone" and the fault-linked thermal readings.
Show answer
Model answer: Fault-linked thermal readings, since they're still moving prediction accuracy every quarter. Cutting a field that's still earning its keep just to save money is exactly the mistake this answer is arguing against.
Before you close the answer
Why this works
Tests whether you can look past total storage cost to which specific fields are still doing real work, instead of treating "cut data" or "keep everything" as a blanket rule.
Follow-up traps
"What if you're wrong and accuracy quietly drops after you thin the signal?" Response: that's what the two-week holdout is for, catching the real cost before committing fleet-wide, and the setting can always be turned back up since it's cheap to reverse.
"Isn't cutting data just an excuse to save money at the model's expense?" Response: no, because the two fields kept in this answer, crash dumps and thermal reads, are the ones still moving accuracy. The cuts target exactly the fields that stopped paying for themselves.
If pressed
The fix that followed the audit didn't just thin the signal field, it also fast-tracked battery-complaint tickets to a 48-hour turnaround, which is what actually made Foday's private script unnecessary going forward.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.