InterviewIntermediateAI Opportunity & Model Strategy / Opportunity identification for AI / #10

Which is the better first AI feature: one that saves users time or one that unlocks a new capability? Defend it.

ORDER · why Karlin Fessler shipped Farlight's plainest idea before its boldest one

Cadeby Health builds Farlight, a wristband and coaching app that reads a runner's heart rate, sleep, and pace and turns it into plain-language advice. With the core dashboards live, the team had two features ready to be Farlight's first real AI feature: a weekly recap that read three screens back as one short paragraph, and a bolder flag that told a runner whether their training put them at real risk of injury. Leadership wanted the injury flag for a launch event. Product manager Karlin Fessler had to pick which one actually deserved to go first, and defend it to a room that wanted the exciting answer.

The direct answer
Ship the time-saving feature first: something like a plain-language recap of a week's workout data, not a feature that predicts something new like injury risk. It has a real number to prove itself against, and if it disappoints, you can fix or quietly drop it for almost nothing. A new capability that flops is loud and hard to walk back, and earning the right to build it well needs exactly the trust and evaluation habits the time-saving feature builds first.
Do this, in order
  1. Ship the time-saving feature first, not the one that promises something new.Why: it has a real, already-measured baseline and a far smaller cost if it disappoints.
  2. Check that a real, already-measurable baseline actually exists for the time-saving candidate before committing to it.Why: "saves time" is only a real claim if there's a number to save time against.
  3. Hold any new-capability feature back until it has a genuinely validated need and enough real confirmed cases to trust the model's judgment.Why: a capability the model can't yet earn calls into question everything you build after it.
  4. Weigh reversibility honestly: a pitched new capability that flops is a public, hard-to-unsay miss, while an underwhelming time-saver can be quietly reworked.Why: this asymmetry, not raw impact, is what should decide the order.
  5. Run any new-capability feature silently against real outcomes before it ever tells a user anything.Why: this is how you find out whether the model actually works on the case that matters, before a person trusts it with something serious.
  6. Once the time-saving feature has built real evaluation habits, revisit the new-capability feature with actual evidence behind it.Why: the point was never to kill it, just to stop it from going first.

How to answer this, stage by stage

Nobody is grading whether Karlin can make the bold feature sound thrilling for five minutes. They're grading whether the pick she names would survive an actual runner trusting it with her own body.

1
Scope it to one product, one team, one fork in the road
Say it like this
"Let's make this concrete. I'm product lead for Farlight at Cadeby Health, we turn wearable data into coaching. We had two features ready to be our first real AI feature, and I had to pick which one goes first."
Why this works
Naming the real product and the real fork stops the answer from staying an abstract debate about time versus capability.
2
Name your method out loud before you use it
Say it like this
"I'd run this as ORDER. Outcome, what a first AI feature actually has to prove. Reversibility, which pick is hardest to walk back if I'm wrong. Dependency, what has to be true before either one is even ready. Evidence, what's cheap to check first. Rank, the actual pick, defended."
Why this works
Two seconds of structure signals a method, not a preference, before the pressure of the room sets in.
3
Reframe what the question is really testing
Say it like this
"This isn't really asking me to pick the more exciting feature. It's asking whether I'll protect the team's first real shot at trust, or spend it on something we can't back up yet."
Why this works
Separates a real answer from a pitch for whichever feature sounds better in a demo.
4
Commit to the pick, before any evidence
Say it like this
"Here's the pick. We ship the time-saving feature first, the weekly recap. Not because it's safer for its own sake. Because it's the one we can actually prove, and proving something is the whole point of going first."
Why this works
This is the direct answer, said plainly, before a single number arrives to back it up.
5
Back it with the one number that actually decides it
Say it like this
"The recap has a real baseline. People were spending about twenty-one minutes every Sunday piecing three screens together by hand. The injury flag has forty-six confirmed cases to learn from, company-wide, almost all sudden ones. That's not a dataset, that's a handful of stories."
Why this works
A real number beats a feeling about which feature is more impressive, every time it's said out loud.
6
Prove it with the failure, compressed
Say it like this
"I'll tell you what happens if you skip this. We shipped the injury flag first, for a launch event. It told a runner she was fine for six Sundays straight. On the seventh, she couldn't finish a long run, and an X-ray found a stress fracture the model had never once seen confirmed in its own training data."
Why this works
Shows the real cost of skipping the Dependency check in four sentences, not a slide deck.
7
Name the alternative you're turning down
Say it like this
"We talked about shipping the injury flag anyway, with a big disclaimer on it. I turned that down. In testing, people stopped reading the disclaimer by the second week. The color was the whole message, caveat or not."
Why this works
Shows a real judgment call was made, not just the safe path that happened to look obvious in hindsight.
8
Close on what ships next, and when
Say it like this
"Ship the recap now. Hold the injury flag in shadow mode until it's been checked against real, confirmed cases across every injury type we claim to catch, not just the obvious ones. That's the whole answer."
Why this works
Ends on something concrete the interviewer can hold the candidate to later, not just a confident closing line.

Let's learn

Farlight is a wristband and app from Cadeby Health. It watches a runner's heart rate, sleep, and pace, and turns it into coaching.

Before any AI feature shipped, a Farlight wearer had three separate screens open every Sunday night: a heart-rate zone chart, a sleep-score trend, and a table of pace splits. Working out what the week actually did to their body took about 21 minutes of comparing all three by hand.

Hand sketched flow diagram titled A Farlight wearer's Sunday night, before any AI. Four boxes connected by arrows, left to right, the fourth box highlighted: Open heart-rate zones. Open sleep score. Open pace splits. Piece it together by hand, 21 minutes.
Three screens, one tired brain, every single Sunday.

Two features were ready to fix that, and only one of them could go first. One read those same three screens back as a plain paragraph in about 3 minutes, what the week did, in words a runner would actually use. The other was bolder: a flag that told a runner whether their training load put them at real risk of injury, something no dashboard had ever tried to say out loud.

What does a "confirmed case" mean here? A real injury that a doctor or physical therapist actually verified and logged, not just a runner reporting soreness in an app. Cadeby's whole multi-year user base had only 46 of these on file, and most were sudden ones: a fall, a twisted ankle. Almost none looked like a slow injury building up over weeks.
Hand sketched decision tree titled Farlight's first real AI feature, root box reads Which one ships first. Two branches: reads a known task back, has a real baseline leads to Weekly recap. Claims a new, unproven judgment call leads to Injury-risk flag.
Two candidates, one small team, one launch slot.

Here is the turn. Cadeby's leadership wanted the bold one for a launch event, and it shipped first. For six Sundays it told a runner named Yasmeen Trenning, in a calm green line, that she wasn't at risk. On the seventh, she couldn't finish a long run. An X-ray four days later found a stress fracture that had been building for weeks, exactly the kind of injury the model had almost never seen confirmed.

We did not tell Yasmeen she was fine. We told her she was fine, in writing, for six Sundays running, on a case the model had never once seen confirmed.

At its worst, a confident wrong answer on something this serious doesn't just cost one runner's trust. It costs the whole feature's credibility the moment her story reaches anyone training the same way she was.

The choice I would take back Cadeby greenlit the injury flag as Farlight's first AI feature under pressure for a launch moment, before the team had anywhere near enough confirmed cases to trust its judgment. I would ship the plain recap first, same team, same model skills, a real number to prove it against, and time to build the confirmed-case data the injury flag actually needed before anyone's body was on the line.

What I would leave alone: the sleep-score number itself doesn't need any of this. It's plain arithmetic on sensor data, not a model's judgment call, and it's been right every time anyone has checked it against a lab test.

The lesson: a first AI feature's job is not to win a launch. It's to prove, cheaply, that a team can tell when their model is right, before they ever hand it something a person's body depends on.

Sunday night: reviewing a week of training, before and after the recap
21 10 0 21 min Before, three screens 3 min After, one recap
Manual reviewAI recap
The recap's whole claim rests on this one number, and the number was already sitting there, measured, before a single line of the model got written.

Now here is the same thing as a story

What you'd actually say sits above. Read this one for the eight Sundays nobody at Cadeby was watching what a calm green line was quietly costing Yasmeen.

Yasmeen Trenning can tell you, before she's laced her shoes, exactly how her left shin is going to feel three miles into a long run. Sixteen years of marathons will do that.

She joined Farlight's beta cohort that spring, three hundred and forty runners training with a regional running club, all wearing the band months before the injury flag went live for anyone else. For the first two weeks after it shipped, the green line at the top of her app felt like a second coach. She still checked her own splits table too, out of habit, just to see if the flag agreed with what she already knew.

It always did. By week three she'd stopped opening the splits table at all. Why would she. The line said fine, and the line had been right every week so far.

Hand sketched two panel comparison titled How Yasmeen's Sunday nights changed in six weeks. Left panel, a person icon labeled Week 2, caption: checks all three screens herself, calls the coach if unsure. Right panel, a person icon labeled Week 7, caption: glances at the green line, laces up, doesn't dig any deeper.
Nobody told her to stop checking. A line that's never once been wrong does that on its own.

Nothing dramatic happened in week five or six. That part matters. No bad morning, no sharp pain she ignored. Her training load had been climbing for a month, the ordinary way marathon training climbs in its final stretch, and the model had nothing in its own history that looked like this kind of slow build. So it stayed green, the whole time, because it genuinely didn't know any better.

In week seven, on a sixteen-mile long run, a dull ache in her left shin turned sharp around mile eleven, and she walked the last three. Four days later, an X-ray found a tibial stress fracture that had likely been forming for weeks.

Yasmeen wasn't careless. She did exactly what any sensible person does with a tool that's never once been wrong: she trusted it a little more each week, and stopped double-checking something it had already checked for her. That's not a flaw in her. It's what a calm, confident green line trains anyone to do.

The message reached Karlin on a Wednesday, forwarded from the run club's coach, not angry, just confused. Why didn't it catch this. Karlin didn't have a clean answer that day. What she had, once she pulled the model's own confirmed-case list, was forty-six examples, mostly falls and sprains, and not one that looked like a slow overuse injury building for a month. The model hadn't missed Yasmeen's case. It had never been shown a case shaped like it.

Two years earlier, on a different feature, Karlin had made the opposite mistake. She'd held a genuinely useful sleep-coaching tool back for four extra months waiting for a level of certainty the feature never actually needed, and a competitor shipped something similar first. She wasn't going to overcorrect into that habit either.

We didn't ship something slower than the launch deck wanted. We shipped something that had to earn the right to speak before it was allowed to.

So this time, she made a different call. The recap went out first, quietly, no launch event behind it, just a plain paragraph reading three screens back to a runner in three minutes instead of twenty-one. The injury flag went into shadow mode: it kept running, kept guessing, but it stopped telling anyone anything, while the team built a real, varied set of confirmed cases behind it, month by month, case type by case type.

Here's what I'd tell myself, standing in the room where the launch date first got picked: a disclaimer doesn't change what a scared runner does with a calm green line. Only real evidence does.

ORDER, the five checks that kept Farlight's boldest idea from shipping first

This isn't a straight tradeoff between two sides of one coin. Farlight had two different features competing for one small team's next quarter, and the job was ranking them, not picking a lane. That's what ORDER is built for, not PICK.

Hand sketched labeled parts diagram titled Five checks before a feature gets to go first. A center document icon labeled First AI Feature, with five callouts around it: Outcome, what it protects. Reversibility, hardest to unsay. Dependency, is it even ready. Evidence, cheap to check first. Rank, the pick, defended.
Five checks a feature clears before it earns the right to go first.
OOutcome. What a first AI feature's real job actually is.
Not the biggest headline, and not a launch event's applause. A first AI feature's real job is to build genuine internal confidence, real evaluation habits, and a runner's trust, fast, more than to maximize any one feature's own raw impact. Everything else Karlin checked exists to serve that one thing.
Name the outcome before ranking anything. Skip this and "which feature is more exciting" quietly becomes the real ranking rule.
RReversibility. Which pick is hardest to walk back if it's wrong.
A time-saving feature that underperforms is easy to quietly rework or retire, almost nobody outside the team ever hears about it. A new-capability feature that flops is a far louder, harder signal to walk back, because it was pitched as something genuinely new, not just a small efficiency tweak. Once Farlight told Yasmeen she was fine in writing, that couldn't be unsaid.
This is why order matters, not preference. One miss is a quiet fix. The other is a story that travels.
Hand sketched two panel comparison titled Which one can Karlin still take back? Left panel, a scale icon labeled Recap underperforms, caption: quietly rework it, almost nobody outside the team ever knows. Right panel, a box icon labeled Injury flag flops, caption: already told a runner she was fine, hard to unsay after an X-ray.
Reversibility isn't a reason to avoid the bold idea. It's a reason to check the evidence before shipping it.
DDependency. What has to already be true for each one to succeed.
A time-saving feature needs an existing, well-understood workflow with a clear, measurable "before" state to prove itself against, which the recap had: twenty-one minutes, three screens, every Sunday. A new-capability feature needs a genuinely validated need, harder to prove exists before shipping, since nobody had ever tried to answer "am I at risk" out loud before. Forty-six confirmed cases, almost none shaped like Yasmeen's, was nowhere near enough to trust that judgment.
This is the step a normal feature list doesn't have. A new report or a new screen doesn't need a labeled outcome dataset behind it before anyone can build it. A risk judgment does.
Hand sketched flow diagram titled What has to happen before either feature gets ranked. Five boxes connected by arrows, left to right, the second box highlighted: Idea lands. Real baseline or real cases. Evidence checked. Given a rank. Built.
Every box after the second one only means something once the second one is actually true.
EEvidence. What's cheap to check before committing.
Whether a real, already-measurable baseline time exists for the time-saving candidate, versus how much unproven belief the new-capability feature rests on. The twenty-one-minute number cost nothing to confirm, a few timed sessions with real users. The injury flag's readiness cost nothing to check either, and what it turned up was forty-six thin, lopsided examples, no real way yet to know if the model's judgment could be trusted on the case type that actually mattered.
Cheap to check, and it settled the whole question, not a guess about which feature "felt" more ready.
Sunday-night app opens, ten weeks, Farlight's beta cohort
100% 50% 0 Word spreads through the cohort, week 9 Wk 1 Wk 5 Wk 10
Weekly Sunday-night opensWord reaches the cohort
The rate held near 93% for seven weeks straight. It never recovered gently, it fell off once trust actually broke.
RRank. The actual pick, defended.
Time-saving first. It has a clear existing baseline to measure against, a lower risk of a highly visible miss, and it builds the exact evaluation muscle and trust a new-capability feature would need to succeed anyway. The injury flag doesn't get killed, it gets held in shadow mode, checked silently against real confirmed outcomes, until it clears a real bar across every injury type it claims to catch, not just the easy majority case.
If this rank would be identical no matter what the Outcome in step O was, it was picked by instinct, not judgment. Swap the outcome to "win the loudest launch demo" and the rank flips completely, which is exactly why naming Outcome first matters.

One alternative is worth naming and rejecting directly: shipping the injury flag anyway, wrapped in a bold on-screen disclaimer warning that it was an early beta and might miss real risks. It lost, because in the two-week pilot test Cadeby actually ran, every runner who read the disclaimer on day one had stopped reading it by day three. A calm green line is the whole message a person absorbs, caveat text or not. The AI-specific failure worth naming is a training-data gap treated as an accuracy problem: the model wasn't slightly wrong about injury risk, it had simply never been shown a case shaped like a slow, weeks-long overuse injury, so it stayed calm on exactly the case type it was least equipped to judge. The guardrail is the shadow-mode requirement: any new-capability feature runs silently against real outcomes, logged and scored, before it ever tells a user anything, and it only speaks once it clears a real bar across the case types it claims to cover, not just the common one. And the trade-off is accepted on purpose: a slower path to the flashy "unlocks something new" launch moment, in exchange for a real trust foundation and a far lower chance of a highly public, hard-to-walk-back miss.

And if you want to be sure it really works, try it somewhere else

Same five letters, a fishing cooperative instead of a wearable company, and the honest answer doesn't change.

Tidecross Cooperative runs Haulnote, a tool that reads a trawler's onboard sensors and radio call-ins and helps a deckhand log the day's catch. Product manager Gwilym Bramridge faced the same fork: ship a feature that turns a full day of scrawled catch logs into one clean summary for the co-op office, or ship a feature that flags when a boat is at real risk of going over its species quota before the trip ends, something no tool at Tidecross had ever tried to predict.

Hand sketched quadrant chart titled Sorting Tidecross's two feature options. X axis, how ready is the evidence, from barely any to well tested. Y axis, hard to walk back if wrong, from easy to adjust to already announced. Announce quota-risk at the annual meeting sits in the upper left, barely any evidence and hard to walk back. Quota-risk flag, quiet beta sits mid-left. Auto-summarize catch logs sits lower right, well tested and easy to adjust. Keep the paper log as-is sits furthest right and low.
The option in the top-left corner is the one worth saying no to, not the one worth rushing.

Same steps, mapped onto Tidecross. Outcome: protect whether a boat actually trusts what Haulnote tells it, not whether the co-op looks impressive at the annual meeting. Reversibility: announcing a quota-risk feature to the whole fleet at that meeting is far harder to walk back than a quiet internal pilot; the catch-log summary can be quietly reworked if deckhands don't like it. Dependency: the quota-risk flag needs real logged cases of boats that actually went over quota, and Tidecross had only 9 confirmed overage cases on file, across three species that behave nothing alike; the summary needs a well-understood existing task, and every deckhand already fills out a paper log by hand, about 25 minutes a trip. Evidence: Gwilym could check the paper-log time cheaply, right away, by timing three trips; the quota-risk flag had no such cheap check, only those 9 thin cases. Rank: the same call. Ship the catch-log summary first, hold the quota-risk flag until real overage cases build up, and notice that the summary feature itself sharpens the co-op's own record-keeping, which is exactly what the quota-risk flag will need to learn from later.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to Rank, name the pick and why, everything else is support.
Cost: there's no budget to build a real confirmed-case set before the meeting. Say so plainly, and use whatever's cheap and real, informal outcome notes the coaches already keep, rather than assuming the data is probably fine.
The model got better, for real: a newer base model turns out to need far less labeled data to hit a reliable bar. Say that too, plainly, and move the new-capability feature up in the queue. The method never says never build it, it says decide from real evidence either way.

Where people run it wrong.
They ship the bold capability because leadership wants a launch story, and skip the evidence check entirely.
They assume a bigger model fixes a data problem, when it's a capability the model has simply never been shown.
They never revisit the shelved feature once trust is built, so it just quietly dies instead of shipping with real evidence behind it.

How to use it live. Before picking either feature, ask out loud: "do we actually have enough real, confirmed cases to trust this model's judgment, or are we hoping it figures it out?" If the honest answer is "we don't know," that's the whole Dependency check, and it's reason enough to let the humbler feature go first.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits deciding which AI feature should ship first?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. Built for ranking real choices by what's hardest to undo, and for checking whether a feature is even ready to compete at all.
2 · THE CAST
Who holds each role in this story, and where do they work?
Tap to flip
ANSWER
Karlin Fessler is the product lead for Farlight at Cadeby Health. Yasmeen Trenning is a masters runner in Farlight's 340-person beta cohort whose stress fracture the injury flag never caught.
3 · THE OUTCOME
What is a first AI feature's real job supposed to protect?
Tap to flip
ANSWER
Building genuine internal confidence, real evaluation habits, and a user's trust, fast, more than maximizing any single feature's raw impact.
4 · REVERSIBILITY
Which kind of feature miss is hardest to walk back once it ships?
Tap to flip
ANSWER
A new-capability feature that flops, since it was pitched as something genuinely new. A time-saving feature that underperforms can be quietly reworked or retired instead.
5 · THE OLD DECISION
What decision would Karlin take back?
Tap to flip
ANSWER
Greenlighting the injury flag as Farlight's first AI feature under pressure for a launch event, before the team had enough confirmed cases to trust its judgment on the case type that actually mattered.
6 · THE NUMBER
Fill in the blank: the confirmed-injury dataset had only ___ cases, while manual weekly review took about ___ minutes before the recap shipped.
Tap to flip
ANSWER
46 confirmed cases; 21 minutes, cut to about 3 minutes once the recap shipped.
7 · THE RANK
State the actual pick, in one line.
Tap to flip
ANSWER
Ship the time-saving recap first, with a real baseline behind it. Hold the injury flag in shadow mode until it clears a real bar on real confirmed cases across every injury type it claims to catch.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs ORDER again on a different product. Which one, and who runs it?
Tap to flip
ANSWER
Haulnote, Tidecross Cooperative's catch-log tool for fishing trawlers. Product manager Gwilym Bramridge runs the same method against a predictive quota-overage flag.

Check yourself Score: 0 / 0

Fill in the blank
1. Yasmeen's stress fracture was confirmed by X-ray ___ days after she couldn't finish her long run. The injury flag had told her she was fine for ___ straight Sundays before that.
Show hint
Check stage 6 of the walkthrough and the story section.
Show answer
4 days; 6 Sundays. The model had never once seen a confirmed case shaped like a slow, weeks-long overuse injury.
Multiple choice
2. Why did Farlight's injury flag never turn amber for Yasmeen, even as a stress fracture was building?
  • A. The wearable's sensor malfunctioned that week.
  • B. It was trained on only 46 confirmed cases, almost all sudden ones, so it had never seen a slow overuse case like hers.
  • C. Yasmeen had turned the flag off by accident.
  • D. Cadeby had disabled the flag for beta cohort users.
Show hint
Check the Dependency letter and the spark box in Let's learn.
Show answer
B. The gap came from a real, checkable data shortage, not a bug, a mistake, or a disabled setting.
True or false
3. True or false: Karlin's plan is to cancel the injury-risk flag for good.
  • True
  • False
Show hint
Check the Rank letter and the priority list's sixth bullet.
Show answer
False. It goes into shadow mode, checked silently against real outcomes, and ships once it clears a real bar across the case types it claims to catch.
Short answer, name the rejected option
4. What alternative did Karlin consider instead of simply holding the injury flag back, and why did it lose?
Show hint
Check stage 7 of the walkthrough and the closing paragraph of the framework recap.
Show answer
Model answer: Shipping the injury flag anyway with a bold on-screen disclaimer. It lost because in a two-week pilot, every runner who read the disclaimer on day one had stopped reading it by day three. The green line was the whole message, caveat or not.
Short answer, apply it yourself
5. Think of an AI feature you've used yourself. Was its first real capability something that saved you time on a task you already understood, or something that claimed to do a new kind of judgment call? What would you have wanted to see checked first?
Show hint
Check the Dependency and Evidence letters, what has to be true and what's cheap to check.
Show answer
Model answer: A feature that claimed a new judgment call (like flagging a risk or a fraud) is worth asking: how many real, confirmed examples did the team actually have before it started telling anyone anything, and did those examples cover the case that would matter most, not just the common one.
Short answer, work the number
6. If Cadeby had 1,200 confirmed injury cases instead of 46, spread evenly across injury types including slow overuse cases, would the same rank still hold? Why or why not?
Show hint
Check the Dependency and Rank letters together.
Show answer
Model answer: not necessarily in the same order, but the process still holds. With a real, varied golden set, the injury flag might clear the Dependency bar and move up. But the Rank step's real point, that nothing ships past shadow mode without a genuine evidence check, would still apply exactly the same way.
Before you close the answer
Why this works
Tests whether a candidate will resist shipping the flashiest AI feature first just because it demos well, and instead protect the harder, less visible thing an AI team actually needs early: real evaluation habits and earned trust.
Follow-up traps
"Isn't the time-saving feature just the safe, boring choice?" Response: it's the provable choice, which is different. Going first means earning the right to be trusted, and only a feature with a real baseline can actually prove itself either way.

"What if a competitor ships the bold capability first and wins the headline?" Response: a headline built on a model that hasn't earned its judgment yet is a liability with a delay on it, not a real lead. The trade-off is accepted on purpose: slower to the exciting launch, in exchange for a foundation that doesn't fail in public.
If pressed
Shadow mode isn't a flat waiting period. It's a rolling check against a golden set stratified by case type, sudden injuries, overuse injuries, and joint-specific injuries scored separately, so the flag can go live on the case types it's actually proven on before it's proven on all of them, instead of waiting for one single bar to clear across everything at once.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more