Your company has no AI features. Propose a sequence of three, and explain the order.
Thrumwell Health runs virtual urgent care and primary care, and its intake team leans on Sillmark, a set of AI tools built into the app that reads what a patient types or says the moment they reach out. With zero AI features live, three were ready to compete for the first launch slot: a plain summarizer for staff, a severity flag for patients, and a full pre-visit recommendation that would need nobody's judgment but its own. Product lead Idelle Kohlmann had a planned order. Leadership wanted the boldest one moved up, for a launch event two months out, and an eager engineering pod nearly gave it to them.
- Ship the summarizer, then the severity flag, then the full recommendation, in that order, never the reverse.Why: this is the actual answer, not a wish list of nice features.
- Check that a real, measured baseline exists for the summarizer before committing to it.Why: "saves staff time" is only a real claim if there's a number to save time against.
- Treat the severity flag as the feature that turns the summarizer's byproduct into real labeled data.Why: coordinators correcting drafts is free, honest data the third feature will actually need.
- Hold the full recommendation back until it clears a real, confirmed-outcome bar across every major symptom category.Why: a bar measured only in total volume hides the exact gap that nearly hurt a patient.
- Weigh reversibility honestly: a public miss from the full recommendation is far harder to walk back than a quiet internal tool underwhelming.Why: this asymmetry, not raw ambition, is what should decide the order.
- Once the first two features have built real evaluation habits, revisit the full recommendation with actual evidence behind it.Why: the point was never to kill it, just to stop it from going first.
How to answer this, stage by stage
Nobody is grading whether Idelle can name three cool features in five minutes. They're grading whether the order she gives them would survive a patient trusting the boldest one with something serious.
Let's learn
Sillmark is the set of AI tools built into Thrumwell Health's telehealth app. It reads what a patient types or says the moment they reach out, before a clinician ever sees the case.
Before any of it existed, an intake coordinator took every call or chat live, then spent about eight minutes typing up a clean note for the clinician: what hurts, since when, what's already been tried. Thrumwell runs about 410 of these a day.
Three features were ready to compete for the first launch slot. One would read the transcript back as a clean note for staff in under a minute. One would flag how urgent a case sounded, so a coordinator could triage the queue faster. One would skip people entirely: read the intake, decide what the patient should do next, and tell them, on its own.
Here is the turn. Leadership wanted the full recommendation for a launch event, and an early, unreviewed build of it went into a quiet pilot ahead of schedule. For six weeks it worked fine on the cases it saw. Then a patient named Marrek Streit typed that he had crushing pressure in his chest, spreading down his left arm, that had started twenty minutes before. The model answered, calm as anything: "Routine. Next available telehealth slot in 5 days."
At its worst, a confident wrong answer on something this serious doesn't just cost one patient's trust. It costs the whole feature's credibility the moment his story reaches anyone who hears how close it came.
What I would leave alone: the appointment-scheduling logic underneath all three features doesn't need any of this. It's plain arithmetic against a calendar, not a model's judgment call, and it's been right every time anyone has checked it.
The lesson: a first AI feature's job is not to win a launch. It's to prove, cheaply, that a team can tell when their model is right, before they ever hand it something a patient's life depends on.
Now here is the same thing as a story
What you'd actually say sits above. Read this one for the six quiet weeks nobody at Thrumwell was watching what a fluent, confident sentence was risking.
Every Tuesday, before the coordinators' shift started, Idelle Kohlmann walked the intake floor with a spare headset and listened to three or four calls live. Six years in healthcare product work had taught her that a real call tells you more than a dashboard ever will.
The summarizer shipped in March. For most of the spring it was the easy win people pointed to in the all-hands. Coordinators went from typing a note for eight minutes to reading a draft and fixing a line or two, two minutes, sometimes three on a bad connection. Clinicians liked it. Idelle let herself enjoy that for a few weeks, which she almost never did.
Then leadership started asking, gently at first, when the real feature was coming. The one where Sillmark would just tell a patient what to do, nobody in the loop, the thing that would make the investor deck sing. Idelle's plan had it third, behind a severity flag that still needed real confirmed outcomes behind it. An engineering pod, eager and not unreasonable, built a rough version anyway, on their own time at first, to see if it was even possible.
It worked, on the cases it saw. A cold. A rash. A sprained ankle after a fall. Six weeks of quiet demos, and it never once said anything alarming, because nothing alarming had come through it yet. Confidence in the room grew the way confidence always grows when nothing has gone wrong: quietly, and past the point anyone would have chosen on purpose.
Then, on a Thursday night, Marrek Streit opened the app and typed that he had crushing pressure in his chest, spreading down his left arm, that had started twenty minutes before. The pilot build read it. Nothing in its confirmed history looked like this. Almost everything it had ever been checked against was a cold, a rash, a sprain. It answered anyway, fluent as ever: "Routine. Next available telehealth slot in 5 days."
His wife read the message over his shoulder. She didn't wait five days, or five minutes. She called for care. He was in a cath lab within the hour, and came home two days later with a stent and a story he still doesn't fully believe.
Idelle heard about it Friday morning, from a clinician, not a dashboard. Nothing had technically broken. The model hadn't crashed or thrown an error. It had answered a question it had no business answering yet, fluently, the same way it answered every question.
Two years earlier, on a different product, she'd made the opposite mistake: held a genuinely useful scheduling tool back for four extra months, waiting for a certainty it never actually needed, and lost that launch moment to a competitor entirely. She wasn't going to overcorrect into that habit either, not blindly.
So this time she made a narrower call. She pulled the pilot build from anywhere it could reach a patient, kept the summarizer live, and moved the severity flag up next, the one honest step that could actually use the labeled data coordinators were already producing by correcting drafts. The full recommendation went into a locked pilot: it kept running, kept guessing, but nothing it said reached anyone, while the team built real, confirmed outcomes across every symptom category that actually mattered, week by week, until the bar was cleared for real.
Here's what I'd tell myself, standing in that glass-walled room the week the pilot build first demoed clean: a feature that's never been wrong yet hasn't earned anything. It just hasn't met the case that matters.
ORDER, and the sequence that would have caught Marrek's chest pain in time
This isn't a straight tradeoff between two sides of one coin. Thrumwell had three features competing for one launch slot, and the job was ranking all three, not picking a lane. That's what ORDER is built for, not PICK.
One alternative is worth naming and rejecting directly: shipping the full recommendation first anyway, but routing only its lowest-confidence answers to a coordinator for review. It lost, because the model's own confidence score was itself unvalidated on rare presentations. It was most falsely confident exactly on cases like Marrek's, the ones it had never been shown, so the safety net would have caught almost nothing on the case type that mattered most. The AI-specific failure worth naming is a training-data coverage gap dressed up as a general accuracy problem: the model wasn't slightly wrong about chest pain, it had simply never been shown a confirmed case shaped like it, so it defaulted to the pattern that covered nearly everything it actually had seen, a routine complaint. The guardrail is the locked-pilot requirement: any full recommendation runs silently against real outcomes, logged and scored by symptom category, before it ever tells a patient anything, and it only speaks once it clears a real bar on the categories it claims to cover, not just the common one. And the trade-off is accepted on purpose: a slower path to the flashy "the model just tells you what to do" launch moment, in exchange for a foundation that doesn't fail on a real patient in public.
And if you want to be sure it really works, try it somewhere else
Same five letters, a city permits office instead of a telehealth company, and the honest answer doesn't change.
The City of Ashendell runs Gatehouse, a tool that reads a building-permit application and helps a reviewer work through it. Permits program lead Wynstan Cazares faced the same fork: ship a feature that turns a long, messy application into a clean summary for reviewers, or ship a feature that tells an applicant which permits qualify for automatic approval, no reviewer in the loop, something no tool at Ashendell had ever tried.
Same steps, mapped onto Ashendell. Outcome: protect whether an applicant actually trusts what Gatehouse tells them, not whether the office looks impressive at a ribbon cutting. Reversibility: announcing auto-approval at a public event is far harder to walk back than a quiet internal pilot; the application summarizer can be quietly reworked if reviewers don't like it. Dependency: auto-approval needs real logged outcomes of permits that were wrongly approved or wrongly denied, and Ashendell had only 55 confirmed outcomes on file, almost entirely simple residential fence permits, none for commercial or environmental-review permits; the summarizer needs a well-understood existing task, and every reviewer already reads a full packet by hand, about 35 minutes a case. Evidence: Wynstan could check the 35-minute number cheaply, right away, by timing three reviewers; auto-approval had no such cheap check, only those 55 thin cases. Rank: the same call. Ship the summarizer first, then a missing-document flag that uses reviewers' corrections as labeled data, and hold auto-approval until real outcomes build up across every permit type, at least 400 confirmed cases, not just residential fences.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to Rank, name the pick and why, everything else is support.
Cost: there's no budget to build a real confirmed-outcome set before a deadline. Say so plainly, and use whatever's cheap and real, informal reviewer notes already on file, rather than assuming the data is probably fine.
The model got better, for real: a newer base model turns out to need far less labeled data to hit a reliable bar. Say that too, plainly, and move the ambitious feature up in the queue. The method never says never build it, it says decide from real evidence either way.
Where people run it wrong.
They ship the bold feature because leadership wants a launch story, and skip the evidence check entirely.
They assume a bigger model fixes a data problem, when it's a coverage gap the model has simply never been shown.
They never revisit the shelved feature once trust is built, so it quietly dies instead of shipping with real evidence behind it.
How to use it live. Before picking any feature, ask out loud: "do we actually have enough real, confirmed outcomes to trust this model's judgment, or are we hoping it figures it out?" If the honest answer is "we don't know," that's the whole Dependency check, and it's reason enough to let the humbler feature go first.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if a competitor ships the autonomous recommendation first and wins the headline?" Response: a recommendation built on 70 thin cases isn't a lead, it's a liability with a delay on it. The trade-off is accepted on purpose: slower to the exciting launch, in exchange for a foundation that doesn't fail on a real patient in public.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Opportunity identification for AI
- #1 What characteristics make a workflow a good candidate for AI? List five.
- #2 Describe a method for finding AI opportunities inside an existing product without starting from the technology.
- #3 How do you distinguish a problem AI solves from a problem AI merely touches?
- #4 Rank these by AI suitability and justify: expense approval, contract review, invoice matching, hiring decisions.
- #5 Explain why high-volume, low-stakes, tolerant-of-error tasks are the best first targets.
- #6 Your support team handles 8,000 tickets a month. Structure a discovery process to find the AI opportunity.