CaseAdvancedAI Opportunity & Model Strategy / Opportunity identification for AI / #14

Your company has no AI features. Propose a sequence of three, and explain the order.

ORDER · the sequence Idelle Kohlmann held back at Thrumwell Health

Thrumwell Health runs virtual urgent care and primary care, and its intake team leans on Sillmark, a set of AI tools built into the app that reads what a patient types or says the moment they reach out. With zero AI features live, three were ready to compete for the first launch slot: a plain summarizer for staff, a severity flag for patients, and a full pre-visit recommendation that would need nobody's judgment but its own. Product lead Idelle Kohlmann had a planned order. Leadership wanted the boldest one moved up, for a launch event two months out, and an eager engineering pod nearly gave it to them.

The direct answer
Ship three features in this order: a low-stakes internal tool that summarizes intake transcripts for staff, then a patient-facing severity flag built on the confirmed data the first feature quietly creates, then a full pre-visit triage recommendation last. The first two have real baselines and a small cost if they disappoint. The last one needs real confirmed outcomes across every major symptom category before it earns the right to speak to a patient alone.
Do this, in order
  1. Ship the summarizer, then the severity flag, then the full recommendation, in that order, never the reverse.Why: this is the actual answer, not a wish list of nice features.
  2. Check that a real, measured baseline exists for the summarizer before committing to it.Why: "saves staff time" is only a real claim if there's a number to save time against.
  3. Treat the severity flag as the feature that turns the summarizer's byproduct into real labeled data.Why: coordinators correcting drafts is free, honest data the third feature will actually need.
  4. Hold the full recommendation back until it clears a real, confirmed-outcome bar across every major symptom category.Why: a bar measured only in total volume hides the exact gap that nearly hurt a patient.
  5. Weigh reversibility honestly: a public miss from the full recommendation is far harder to walk back than a quiet internal tool underwhelming.Why: this asymmetry, not raw ambition, is what should decide the order.
  6. Once the first two features have built real evaluation habits, revisit the full recommendation with actual evidence behind it.Why: the point was never to kill it, just to stop it from going first.

How to answer this, stage by stage

Nobody is grading whether Idelle can name three cool features in five minutes. They're grading whether the order she gives them would survive a patient trusting the boldest one with something serious.

1
Scope it to one team, one launch fight
Say it like this
"Let's make this concrete. I'm the product lead for Sillmark at Thrumwell Health, we turn patient intake calls and chats into something a clinician can act on fast. We had no AI features live, three candidates ready, and leadership wanted the boldest one moved to the front."
Why this works
Naming the real product and the real fight stops the answer from staying an abstract debate about caution versus ambition.
2
Name your method out loud before you use it
Say it like this
"I'd run this as ORDER. Outcome, what the first three features together need to prove. Reversibility, which pick is hardest to walk back if it's wrong. Dependency, what has to be true before the next one can even work. Evidence, what's cheap to check before committing. Rank, the actual sequence, defended."
Why this works
Two seconds of structure signals a method, not a preference, before the pressure of the room sets in.
3
Reframe what the question is really testing
Say it like this
"This isn't really asking me to name three cool features. It's asking whether I'll protect the team's first real shot at trust, or spend it proving something we can't back up yet."
Why this works
Separates a real answer from a feature wish list dressed up as a roadmap.
4
Give the sequence, before any evidence
Say it like this
"Here's the order. First, an internal tool that summarizes the intake transcript for staff. Second, a patient-facing flag for how urgent something sounds. Third, a full pre-visit recommendation, on its own, no person checking it first. That last one only goes third."
Why this works
This is the direct answer, said plainly, before a single number arrives to back it up.
5
Back it with the one chain that actually decides it
Say it like this
"Feature one costs us almost nothing to get wrong, and it happens to build the exact labeled data feature three would need. Feature three, if we skipped to it, had about 70 confirmed cases behind it at the time, and not one shaped like a real chest pain call."
Why this works
A real number and a real dependency beat a feeling about which feature is more impressive, every time it's said out loud.
6
Prove it with the failure, compressed
Say it like this
"I'll tell you what happens if you skip the order. An early build of the full recommendation told a patient with chest pain spreading down his arm that it was routine, next slot in five days. His wife didn't buy it. She called for care instead, and he needed a stent that night."
Why this works
Shows the real cost of skipping the Dependency check in four sentences, not a slide deck.
7
Close on what ships next, and the bar the last one has to clear
Say it like this
"Ship the summarizer now. Move the severity flag up next. Hold the full recommendation until it's cleared a real bar, confirmed outcomes across every major symptom category, not just the common ones. That's the whole answer."
Why this works
Ends on something concrete the interviewer can hold the candidate to later, not just a confident closing line.

Let's learn

Sillmark is the set of AI tools built into Thrumwell Health's telehealth app. It reads what a patient types or says the moment they reach out, before a clinician ever sees the case.

Before any of it existed, an intake coordinator took every call or chat live, then spent about eight minutes typing up a clean note for the clinician: what hurts, since when, what's already been tried. Thrumwell runs about 410 of these a day.

Hand sketched flow diagram titled Before Sillmark, every intake case. Four boxes connected by arrows, left to right, the third box highlighted: Call or chat comes in. Coordinator reads it live. Writes the summary by hand. Clinician reads it cold.
Eight minutes, every case, all day, before a clinician ever sees it.

Three features were ready to compete for the first launch slot. One would read the transcript back as a clean note for staff in under a minute. One would flag how urgent a case sounded, so a coordinator could triage the queue faster. One would skip people entirely: read the intake, decide what the patient should do next, and tell them, on its own.

Hand sketched labeled parts diagram titled Three features, one launch slot. A center document icon labeled Sillmark's first AI feature, with three callouts around it: Summarize the transcript. Flag how severe it is. Recommend the whole visit.
Three candidates, one small team, one launch slot.
What does a "confirmed case" mean here? A case where a real clinician later checked what the model said against what actually happened, and logged whether it was right. Not a patient just closing the app. At the time of the near miss, Sillmark had 70 of these on file, and almost all were common complaints: colds, rashes, a sprained ankle. None looked like a real chest pain call.

Here is the turn. Leadership wanted the full recommendation for a launch event, and an early, unreviewed build of it went into a quiet pilot ahead of schedule. For six weeks it worked fine on the cases it saw. Then a patient named Marrek Streit typed that he had crushing pressure in his chest, spreading down his left arm, that had started twenty minutes before. The model answered, calm as anything: "Routine. Next available telehealth slot in 5 days."

The model was not a little wrong about Marrek's chest pain. It had never once been shown a confirmed case shaped like it.

At its worst, a confident wrong answer on something this serious doesn't just cost one patient's trust. It costs the whole feature's credibility the moment his story reaches anyone who hears how close it came.

The choice I would take back Thrumwell let an engineering pod pilot the full recommendation ahead of the planned order, because it was the feature leadership was most excited to show off. I would hold it back until the first two features had built real, confirmed outcomes across every symptom category the model would need to judge, not just the easy majority.

What I would leave alone: the appointment-scheduling logic underneath all three features doesn't need any of this. It's plain arithmetic against a calendar, not a model's judgment call, and it's been right every time anyone has checked it.

The lesson: a first AI feature's job is not to win a launch. It's to prove, cheaply, that a team can tell when their model is right, before they ever hand it something a patient's life depends on.

Writing an intake summary, before and after the summarizer
Before, by hand 8 min After, summarizer under 2 min
Manual write-upAI-drafted summary
This is the number the summarizer's whole claim rests on, and it was already sitting there, timed, before a line of the model got written.

Now here is the same thing as a story

What you'd actually say sits above. Read this one for the six quiet weeks nobody at Thrumwell was watching what a fluent, confident sentence was risking.

Every Tuesday, before the coordinators' shift started, Idelle Kohlmann walked the intake floor with a spare headset and listened to three or four calls live. Six years in healthcare product work had taught her that a real call tells you more than a dashboard ever will.

The summarizer shipped in March. For most of the spring it was the easy win people pointed to in the all-hands. Coordinators went from typing a note for eight minutes to reading a draft and fixing a line or two, two minutes, sometimes three on a bad connection. Clinicians liked it. Idelle let herself enjoy that for a few weeks, which she almost never did.

Then leadership started asking, gently at first, when the real feature was coming. The one where Sillmark would just tell a patient what to do, nobody in the loop, the thing that would make the investor deck sing. Idelle's plan had it third, behind a severity flag that still needed real confirmed outcomes behind it. An engineering pod, eager and not unreasonable, built a rough version anyway, on their own time at first, to see if it was even possible.

It worked, on the cases it saw. A cold. A rash. A sprained ankle after a fall. Six weeks of quiet demos, and it never once said anything alarming, because nothing alarming had come through it yet. Confidence in the room grew the way confidence always grows when nothing has gone wrong: quietly, and past the point anyone would have chosen on purpose.

Hand sketched two panel comparison titled The night the skunkworks build spoke first. Left panel, a gauge icon labeled What the model said, caption: Routine, next slot in 5 days. Right panel, a person icon labeled What actually helped, caption: his wife called for care anyway.
One panel is what the model said. The other is what actually kept Marrek alive that night.

Then, on a Thursday night, Marrek Streit opened the app and typed that he had crushing pressure in his chest, spreading down his left arm, that had started twenty minutes before. The pilot build read it. Nothing in its confirmed history looked like this. Almost everything it had ever been checked against was a cold, a rash, a sprain. It answered anyway, fluent as ever: "Routine. Next available telehealth slot in 5 days."

We didn't nearly lose a launch demo. We nearly lost Marrek.

His wife read the message over his shoulder. She didn't wait five days, or five minutes. She called for care. He was in a cath lab within the hour, and came home two days later with a stent and a story he still doesn't fully believe.

Idelle heard about it Friday morning, from a clinician, not a dashboard. Nothing had technically broken. The model hadn't crashed or thrown an error. It had answered a question it had no business answering yet, fluently, the same way it answered every question.

Two years earlier, on a different product, she'd made the opposite mistake: held a genuinely useful scheduling tool back for four extra months, waiting for a certainty it never actually needed, and lost that launch moment to a competitor entirely. She wasn't going to overcorrect into that habit either, not blindly.

So this time she made a narrower call. She pulled the pilot build from anywhere it could reach a patient, kept the summarizer live, and moved the severity flag up next, the one honest step that could actually use the labeled data coordinators were already producing by correcting drafts. The full recommendation went into a locked pilot: it kept running, kept guessing, but nothing it said reached anyone, while the team built real, confirmed outcomes across every symptom category that actually mattered, week by week, until the bar was cleared for real.

Here's what I'd tell myself, standing in that glass-walled room the week the pilot build first demoed clean: a feature that's never been wrong yet hasn't earned anything. It just hasn't met the case that matters.

ORDER, and the sequence that would have caught Marrek's chest pain in time

This isn't a straight tradeoff between two sides of one coin. Thrumwell had three features competing for one launch slot, and the job was ranking all three, not picking a lane. That's what ORDER is built for, not PICK.

Hand sketched labeled parts diagram titled Five checks before a feature goes first. A center gauge icon labeled Sillmark's launch order, with five callouts around it: Outcome, what it protects. Reversibility, hardest to unsay. Dependency, what unblocks it. Evidence, cheap to check. Rank, the pick, defended.
Five checks a feature clears before it earns the right to go first.
OOutcome. What the first three features together need to prove.
Not the biggest headline, and not what plays best in an investor demo. A first set of AI features has one real job: build genuine confirmed-outcome data, a working way to check the model's judgment, and a patient's trust, fast, more than maximizing any single feature's own ambition. Everything else Idelle checked exists to serve that one thing.
Name the outcome before ranking anything. Skip this and "which feature demos best" quietly becomes the real ranking rule.
RReversibility. Which pick is hardest to walk back if it's wrong.
An internal summarizer that disappoints gets quietly reworked, almost nobody outside the team ever hears about it. A severity flag that misfires is visible but recoverable, a coordinator still reviews the queue behind it. A full recommendation that flops is the loudest, hardest miss to walk back, because it was pitched as needing nobody's judgment but its own. Once Sillmark told Marrek his chest pain was routine, that couldn't be unsaid.
This is why order matters, not preference. One miss is a quiet fix. The last one is a story a patient's family tells for years.
Hand sketched two panel comparison titled Which one can Idelle still take back. Left panel, a box icon labeled Summarizer disappoints, caption: quietly reworked, nobody outside the team knows. Right panel, a scale icon labeled Autonomous flag flops, caption: already told a patient he was fine, in writing.
Reversibility isn't a reason to avoid the bold idea. It's a reason to check the evidence before shipping it.
DDependency. What has to already be true before the next one can work.
The summarizer needs an existing, well-understood task with a real "before" number to prove itself against, which it had: eight minutes, every case, every day. The severity flag needs a fast, cheap way to tell right from wrong, which the summarizer builds for free, coordinators correcting drafts is labeled data nobody had to ask for. The full recommendation needs real confirmed outcomes across every major symptom category, and at the time of the near miss, Sillmark had 70, almost none shaped like a real emergency.
This is the step a normal feature roadmap doesn't have. A new report doesn't need a labeled outcome dataset behind it before anyone can build it. A judgment about someone's chest pain does.
Hand sketched flow diagram titled What unblocks what. Five boxes connected by arrows, left to right, the third box highlighted: Summarizer ships. Coordinators confirm severity. Confirmed cases build up. Severity flag ships. Autonomous flag becomes viable.
Every box after the third one only means something once the third one is actually true.
EEvidence. What's cheap to check before committing.
Whether a real, already-measurable baseline exists for the summarizer, versus how much unproven belief the full recommendation rests on. The eight-minute number cost nothing to confirm, a few timed shifts with real coordinators. The recommendation's readiness cost nothing to check either, and what it turned up was 70 thin, lopsided cases, no real way yet to know if the model's judgment could be trusted on the case that actually mattered.
Cheap to check, and it settled the whole question, not a guess about which feature felt more ready.
Confirmed severity-outcome cases logged, 14 weeks of the severity flag running
600 300 0 500-case bar Crosses the bar, week 13 Wk 1 Wk 7 Wk 14
Confirmed outcomes, running totalCrosses the 500-case bar
This climbed from 70 to 512 in the fourteen weeks after the severity flag shipped, entirely as a byproduct of it running. Skip straight to the full recommendation and this line never gets drawn at all.
RRank. The actual sequence, defended.
Summarizer first, severity flag second, full recommendation third. The first two have a real baseline and a small cost if they disappoint, and together they build the exact confirmed-outcome data the third one needs to be trusted at all. The recommendation doesn't get killed, it gets locked in a pilot, checked silently against real outcomes, until it clears a real bar across every symptom category it claims to cover, not just the common ones.
If this rank would be identical no matter what the Outcome in step O was, it was picked by instinct, not judgment. Swap the outcome to "win the loudest investor demo" and the rank flips completely, which is why naming Outcome first matters.

One alternative is worth naming and rejecting directly: shipping the full recommendation first anyway, but routing only its lowest-confidence answers to a coordinator for review. It lost, because the model's own confidence score was itself unvalidated on rare presentations. It was most falsely confident exactly on cases like Marrek's, the ones it had never been shown, so the safety net would have caught almost nothing on the case type that mattered most. The AI-specific failure worth naming is a training-data coverage gap dressed up as a general accuracy problem: the model wasn't slightly wrong about chest pain, it had simply never been shown a confirmed case shaped like it, so it defaulted to the pattern that covered nearly everything it actually had seen, a routine complaint. The guardrail is the locked-pilot requirement: any full recommendation runs silently against real outcomes, logged and scored by symptom category, before it ever tells a patient anything, and it only speaks once it clears a real bar on the categories it claims to cover, not just the common one. And the trade-off is accepted on purpose: a slower path to the flashy "the model just tells you what to do" launch moment, in exchange for a foundation that doesn't fail on a real patient in public.

And if you want to be sure it really works, try it somewhere else

Same five letters, a city permits office instead of a telehealth company, and the honest answer doesn't change.

The City of Ashendell runs Gatehouse, a tool that reads a building-permit application and helps a reviewer work through it. Permits program lead Wynstan Cazares faced the same fork: ship a feature that turns a long, messy application into a clean summary for reviewers, or ship a feature that tells an applicant which permits qualify for automatic approval, no reviewer in the loop, something no tool at Ashendell had ever tried.

Hand sketched quadrant chart titled Sorting Ashendell's three permit features. X axis, how ready is the evidence, from barely any to well tested. Y axis, hard to walk back if wrong, from easy to adjust to already announced. Announce auto-approval at the ribbon cutting sits in the upper left, barely any evidence and hard to walk back. Auto-approval, quiet pilot sits mid-left. Missing-document flag sits mid-right. Summarize applications for reviewers sits lower right, well tested and easy to adjust.
The option in the top-left corner is the one worth saying no to, not the one worth rushing.

Same steps, mapped onto Ashendell. Outcome: protect whether an applicant actually trusts what Gatehouse tells them, not whether the office looks impressive at a ribbon cutting. Reversibility: announcing auto-approval at a public event is far harder to walk back than a quiet internal pilot; the application summarizer can be quietly reworked if reviewers don't like it. Dependency: auto-approval needs real logged outcomes of permits that were wrongly approved or wrongly denied, and Ashendell had only 55 confirmed outcomes on file, almost entirely simple residential fence permits, none for commercial or environmental-review permits; the summarizer needs a well-understood existing task, and every reviewer already reads a full packet by hand, about 35 minutes a case. Evidence: Wynstan could check the 35-minute number cheaply, right away, by timing three reviewers; auto-approval had no such cheap check, only those 55 thin cases. Rank: the same call. Ship the summarizer first, then a missing-document flag that uses reviewers' corrections as labeled data, and hold auto-approval until real outcomes build up across every permit type, at least 400 confirmed cases, not just residential fences.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to Rank, name the pick and why, everything else is support.
Cost: there's no budget to build a real confirmed-outcome set before a deadline. Say so plainly, and use whatever's cheap and real, informal reviewer notes already on file, rather than assuming the data is probably fine.
The model got better, for real: a newer base model turns out to need far less labeled data to hit a reliable bar. Say that too, plainly, and move the ambitious feature up in the queue. The method never says never build it, it says decide from real evidence either way.

Where people run it wrong.
They ship the bold feature because leadership wants a launch story, and skip the evidence check entirely.
They assume a bigger model fixes a data problem, when it's a coverage gap the model has simply never been shown.
They never revisit the shelved feature once trust is built, so it quietly dies instead of shipping with real evidence behind it.

How to use it live. Before picking any feature, ask out loud: "do we actually have enough real, confirmed outcomes to trust this model's judgment, or are we hoping it figures it out?" If the honest answer is "we don't know," that's the whole Dependency check, and it's reason enough to let the humbler feature go first.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits deciding the order of a company's first three AI features?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. Built for ranking real choices by what's hardest to undo, and for checking whether a feature is even ready to compete at all.
2 · THE CAST
Who holds each role in this story, and where do they work?
Tap to flip
ANSWER
Idelle Kohlmann is the product lead for Sillmark at Thrumwell Health. Marrek Streit is the patient whose chest pain an early, unreviewed build of the recommendation feature called routine.
3 · THE OUTCOME
What does a first set of AI features actually need to protect?
Tap to flip
ANSWER
Real confirmed-outcome data, a working way to check the model's judgment, and a patient's trust, fast, more than maximizing any single feature's own ambition.
4 · REVERSIBILITY
Which kind of feature miss is hardest to walk back once it ships?
Tap to flip
ANSWER
A full, autonomous recommendation that flops, since it was pitched as needing nobody's judgment but its own. An internal tool that underperforms can be quietly reworked instead.
5 · THE OLD DECISION
What decision would Idelle take back?
Tap to flip
ANSWER
Letting an eager engineering pod pilot the full recommendation ahead of the planned order, before Sillmark had anywhere near enough confirmed cases across every symptom category to trust its judgment.
6 · THE NUMBER
Fill in the blank: the confirmed-outcome dataset had only ___ cases at the time of the near miss, and manual summary writing took about ___ minutes before the summarizer shipped.
Tap to flip
ANSWER
70 confirmed cases; 8 minutes, cut to under 2 minutes once the summarizer shipped.
7 · THE RANK
State the actual pick, in one line.
Tap to flip
ANSWER
Ship the summarizer first, then the severity flag, then hold the full recommendation until it clears a real bar on confirmed outcomes across every major symptom category, not just the common ones.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs ORDER again on a different product. Which one, and who runs it?
Tap to flip
ANSWER
Gatehouse, the City of Ashendell's permit review tool. Permits program lead Wynstan Cazares runs the same method against an autonomous approval recommendation.

Check yourself Score: 0 / 0

Multiple choice
1. Why did Sillmark's pilot build call Marrek's chest pain "routine"?
  • A. The app's connection dropped mid-message.
  • B. It had been checked against only 70 confirmed cases, almost none shaped like a real emergency.
  • C. Marrek typed his symptoms in the wrong field.
  • D. Thrumwell had turned off the feature's safety settings.
Show hint
Check the Dependency letter and the spark box in Let's learn.
Show answer
B. The gap came from a real, checkable data shortage, not a bug, a typo, or a disabled setting.
Fill in the blank
2. Idelle's plan put the summarizer first, the severity flag second, and the full recommendation ___.
Show hint
Check stage 4 of the walkthrough.
Show answer
Third, last. It needed confirmed outcomes the other two features hadn't built yet.
True or false
3. True or false: Idelle's plan is to cancel the full recommendation feature for good.
  • True
  • False
Show hint
Check the Rank letter and the priority list's sixth bullet.
Show answer
False. It goes into a locked pilot, checked silently against real outcomes, and ships once it clears a real bar across the symptom categories it claims to cover.
Short answer, name the rejected option
4. What alternative did Idelle consider instead of simply holding the full recommendation back, and why did it lose?
Show hint
Check the closing paragraph of the framework recap, right after the Rank letter.
Show answer
Model answer: Shipping the recommendation anyway, but routing only its lowest-confidence answers to a coordinator. It lost because the model's confidence score was itself unvalidated on rare cases, so it would have been most falsely confident exactly on presentations like Marrek's.
Short answer, apply it yourself
5. Think of an AI feature you've used yourself. Was its first real capability something that saved you time on a task you already understood, or something that claimed a new kind of judgment call? What would you have wanted checked first?
Show hint
Check the Dependency and Evidence letters, what has to be true and what's cheap to check.
Show answer
Model answer: A feature that claimed a new judgment call (like flagging a risk or approving something automatically) is worth asking: how many real, confirmed outcomes did the team have before it started acting alone, and did those outcomes cover the case that would matter most, not just the common one.
Short answer, work the number
6. If Thrumwell had 800 confirmed outcomes instead of 70, spread evenly across every symptom category including chest pain, would the same rank still hold? Why or why not?
Show hint
Check the Dependency and Rank letters together.
Show answer
Model answer: not necessarily in the same order, but the process still holds. With a real, varied confirmed-outcome set, the full recommendation might clear the Dependency bar sooner and move up. But the Rank step's real point, that nothing ships past a locked pilot without a genuine evidence check, applies exactly the same way either way.
Before you close the answer
Why this works
Tests whether a candidate will resist shipping the most impressive-sounding AI feature first just because leadership wants a launch moment, and instead protect the slower, less visible thing an AI team actually needs early: real confirmed-outcome data and earned trust.
Follow-up traps
"Isn't the summarizer just the boring, obvious choice?" Response: it's the provable choice, and it's what actually builds the confirmed-outcome data the ambitious feature needs, for free, nothing else does that.

"What if a competitor ships the autonomous recommendation first and wins the headline?" Response: a recommendation built on 70 thin cases isn't a lead, it's a liability with a delay on it. The trade-off is accepted on purpose: slower to the exciting launch, in exchange for a foundation that doesn't fail on a real patient in public.
If pressed
The 500-case bar isn't one number. It's a bar per symptom category, chest pain, breathing trouble, and abdominal pain scored separately, so the feature can go live on categories it's actually proven on before it's proven on all of them, instead of waiting for one single bar to clear across everything at once.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more