ConceptIntermediateAI Opportunity & Model Strategy / Opportunity identification for AI / #3
How do you distinguish a problem AI solves from a problem AI merely touches?
PICK · Bicuspid's model was right about who wouldn't show, 62% of the time. Right was never going to be the same thing as solved
Bicuspid builds the scheduling and reminder software a dental practice runs its day on, including a model that scores every booked appointment for no-show risk. Lorand Kirkpatrick owns that model. Oswin Sunderland runs growth, and he just closed a nine-location renewal with Fettle Dental Partners by telling them Bicuspid's AI fixes their no-show problem. The model does something real. It was never built to do that.
The direct answer
A problem AI solves is one where the model's output is the finished thing, an outcome the person can act on directly. A problem AI merely touches is one where the output is a real, useful clue, but the actual fix still sits on someone's desk. Test it with one question: after the model runs, does the person still have to do the hard part themselves? If the answer is yes, you built a signal, not a solution.
Do this, in order
Run every AI feature through one test before you scope or sell it: after it runs, does the person still have to do the hard part alone?Why: a yes means you built a signal, not a fix, no matter how good the model is underneath it.
Separate what the model produces from what the product does with it.Why: a risk score and an action that changes the outcome are two different builds, and only one of them earns the word "solves."
Word the pitch to match the actual deliverable, not the ambition for it.Why: "flags who's at risk" and "fixes no-shows" describe two different products, and the client only hears the one you said out loud.
Build the action layer as its own real feature, with its own launch.Why: it's the piece that turns a correct prediction into a changed outcome, which is the only thing that actually earns "solved."
Treat the cost of honesty as cheap and the cost of overselling as expensive, and act like it.Why: a smaller, honest deal costs some excitement now. A bigger, oversold deal costs the whole relationship the day the gap shows up in production.
Recalibrate the model per location before trusting it anywhere new.Why: a risk score tuned on one clinic's patients can be quietly wrong at another one, and an unchecked flag is worse than no flag at all.
How to answer this, stage by stage
Nobody in the room is grading whether you sound thoughtful about "AI limitations." They're grading whether you can point at the exact seam between a correct model and a finished product.
1
Scope it to one product and one real decision
Say it like this
"Let's ground this in something real. Bicuspid builds scheduling and reminder software for dental practices, and it scores every booked appointment for no-show risk. Lorand owns that model. His growth lead, Oswin, just closed a nine-location renewal by telling the client the AI fixes their no-show problem. It doesn't, not by itself, and Lorand has to work out what it actually does before the next quarterly review."
Why this works
Keeps the interviewer grading a specific decision, not a theory about honesty in sales.
2
Say your structure out loud
Say it like this
"I'll use PICK. Position, the actual line between a problem AI solves and one it only touches. Impact, what breaks for the client when that line gets crossed without anyone noticing. Cost asymmetry, which mistake is cheap and which one is expensive. Kill criteria, the one test that tells you which kind of feature you actually built."
Why this works
Two seconds of structure tells the interviewer a method is coming, not a vibe.
3
Draw the line, before any story
Say it like this
"My position: a problem AI solves is one where the model's output is the finished thing, something the person can act on directly. A problem AI merely touches is one where the output is a clue about the problem, real and useful, but the actual fix is still sitting on someone's desk."
Why this works
This is the direct answer, said early enough that the story can't blur it later.
4
Tie the line to what the model can actually do, not to features in general
Say it like this
"This only matters because a model's output is a probability, not a decision. A risk score of 'this patient is 71% likely to no-show' is real information. It is not a rebooked seat, and it is not a filled slot. Something has to turn that number into an action, or the number is just a fact nobody used."
Why this works
Keeps the answer anchored to model behavior instead of drifting into generic feature-honesty advice.
5
Prove it with the compressed failure
Say it like this
"Here's what happened. Bicuspid's model flags about 51 high-risk appointments a day across Fettle's nine locations, right about 62% of the time. For thirteen weeks, that flag just sat on the front-desk screen. Flagged patients got the exact same reminder text as everyone else. The no-show rate moved from 21% to 19%. Basically noise."
Why this works
A real number the whole argument would fall apart without, not a hypothetical.
6
Name what's lost on each side
Say it like this
"Undersell it, call it a risk flag and nothing more, and you lose a little shine on the sales call. Oversell it, call it a fix, and the client finds the gap themselves, in a quarterly review, in front of their own leadership, and now it's not a feature problem, it's a trust problem."
Why this works
Naming both losses stops the answer from collapsing into "just be honest," which isn't a decision.
7
Name which mistake is cheap and which is expensive
Say it like this
"Scoping it honestly from day one would have meant a smaller first pilot and a few extra weeks explaining the roadmap. Call that cheap. Overselling it cost Bicuspid a six-week emergency sprint, about 240 hours pulled off other work, to build the thing that should have shipped with the pitch, and it nearly cost the renewal outright. That's not a close call. It only points one way."
Why this works
This is the center of PICK: naming which error a team can actually afford.
8
Give the kill test, and close on the fix
Say it like this
"Here's my test: after the feature runs, does the person still have to do the hard part themselves? If a flagged patient gets nothing different, that's a touches feature, still useful, but not a fix. So Bicuspid built the actual fix: flagged patients get an earlier reminder, a one-tap reschedule to a guaranteed slot, and if they don't confirm by 24 hours out, the seat goes to someone on standby automatically. No-show rate at Fettle went from 19% to 9% over the next six weeks. That's what solved looks like."
Why this works
Ends on something countable, not a feeling that it probably worked.
Let's learn
Five steps, and the whole argument sits inside the fourth one: what happens between the model being right and the chair actually filling.
Bicuspid runs the schedule for Fettle Dental Partners, nine dental clinics under one ownership group. Its biggest and busiest location, Kestbridge, is where the pilot started and where most of the model's training history came from. Book an appointment anywhere in that network, and Bicuspid's model quietly scores it: how likely is this specific patient, at this specific time, to not show up.
Before any of this, Fettle's no-show rate sat at 21% network-wide. Roughly one in five booked chairs, empty, about 71 a day across nine locations, and a dental chair that sits empty for thirty minutes can't be filled after the fact, that revenue is just gone. Fettle had been living with that number for years, the ordinary cost of running a busy practice.
Once Bicuspid's model went live, and Oswin sold it as the fix, front desks got a new tile on the dashboard every morning: about 51 names, flagged red, the ones most likely to skip. Nothing else changed for those patients. Same reminder text, same call script, same slot. The rate held at 19%.
The list changed which names got a red flag. It never changed what happened to the person wearing one.
Knowledge spark: what is a no-show risk score, actually?
A number the model attaches to one booked appointment, based on things like the patient's past no-show history, how far out they booked, and whether they've confirmed anything yet. Say it comes out at 71%. That means, out of a hundred appointments that looked just like this one, about seventy-one didn't show up. It's a real pattern. It isn't a seat filled.
Here's the turn. The model wasn't wrong. When it said a patient was high risk, it was right about 62% of the time. Being right was never the problem. The problem was that nobody had built anything that happened differently once the model was right. A correct guess with no action behind it doesn't fill a chair.
The model wasn't wrong about who would skip. It was just never wired to anything that stopped them.
The choice I would take back
Oswin's renewal deck said, plainly, that Bicuspid "fixes no-shows." That sentence made sense in the room where it got written: the risk model had just cleared its accuracy bar, the deal was two weeks from closing, and nobody on the product side had flagged the gap before it went out. It stopped making sense the day Fettle found that gap themselves. I would take that sentence back, sell the risk score for exactly what it is, and hold the word "fixes" until the action loop that earns it actually exists.
What I would leave alone: Fettle's clinical director also sees the raw risk score on an internal staffing dashboard, purely to plan chair capacity for the week ahead. Nobody promised that dashboard would fix anything. It's honest about being a clue, and a clue is exactly what it's for. That one stays exactly as it is.
The lesson: a model's output being correct and a product being finished are two different claims, and it's easy to let the first stand in for the second because they feel like the same kind of good news. They aren't. Only one of them means the customer's actual problem went away.
Now here is the same thing as a story
The short version above is what you'd actually say out loud in the room. Read this one for what it cost to find out the slow way.
Lorand Kirkpatrick has spent four years building scoring models for Bicuspid, the kind that don't do anything flashy, they just quietly rank things: which claim needs review first, which lead is worth a call, which appointment is about to become an empty chair. He is good at the part everyone skips, checking whether a model's confidence number actually means what it claims to mean before anyone ships it.
The no-show model was his best work yet. Trained on two years of Fettle's own booking history, it could look at a Tuesday 7:40am slot, booked by a patient with two prior no-shows and no insurance on file, and say, correctly about 71% of the time, that the chair was about to sit empty.
Oswin Sunderland runs growth at Bicuspid, and he'd been chasing the Fettle renewal for two quarters. Nine locations, a real logo, the kind of deal that gets a mention in the all-hands. When Lorand's model cleared its accuracy bar in March, Oswin built it straight into the pitch: Bicuspid's AI now fixes your no-show problem. Nobody in the room pushed back. The number really was good.
For the first few weeks after signing, it felt like it was working, in the vague way these things always feel like they're working right after a deal closes. Fettle's ops team got a new tile on their dashboard every morning, 51 names, flagged red. People liked looking at it. It felt like control.
Then the weeks kept passing, and the number underneath the tile didn't move. Week four, still 20%. Week eight, still 19%. Lorand noticed it before anyone at Fettle said a word, the way you notice a metric that's gone quiet in exactly the way it shouldn't.
He didn't chase it right away. There was always something else that quarter.
Thirteen quiet weeks, then one sentence in a screen-share, then six weeks to build the thing that should have shipped with the pitch.
The trigger, when it came, was small. Fettle's quarterly business review, week thirteen, a screen-share and a spreadsheet. Their ops director scrolled to the no-show line, paused, and said, mostly to herself: "You said this fixes no-shows. All it did was put a color on a name we already both knew was going to skip."
Nobody argued with her. She wasn't wrong. Lorand sat there doing the math in his head: thirteen weeks, fifty-one names a day, and not one of them had gotten anything different because the model flagged them. Same text. Same call script. Same chair, empty at 7:40.
We didn't lose Fettle nine empty chairs a day. We lost the sentence in the pitch deck that promised them a fix.
I want to say the model was the problem. It wasn't. The model was doing exactly what a risk score does, sitting there, correct, waiting for someone to act on it. Nobody had built the acting-on-it part. Oswin's deck had promised a finished thing. Lorand's model had shipped a clue.
Neither of them was lying about what they were holding. The contract was real. So was the score. They just weren't the same thing.
Months earlier, in the release meeting where the model cleared its bar, someone had asked, almost as an aside, "so what happens differently for a flagged patient?" The honest answer at the time was: nothing yet, that's phase two. It got written down as a note for later, and the deck went out anyway, because the Fettle pitch was two weeks out and the accuracy number was too good not to use.
I would take that back. Not the model. The sentence. I'd hold the word "fixes" until phase two actually existed, and sell phase one for exactly what it was: a flag good enough to build a real fix on top of.
Here's the replay. Six weeks after the QBR, phase two shipped: flagged patients get their reminder a day earlier, a one-tap button to a guaranteed alternate slot, and if they haven't confirmed by 24 hours out, the seat goes to the next name on a standby list automatically, no front-desk phone call required. Same model, same 51 names a day. No-show rate at week fourteen: still 19%. Week sixteen: 15%. Week twenty: 9%.
One version of this story sells "fixes no-shows" in March and spends the next two quarters explaining a flat line to an increasingly unhappy client. The other spends six weeks building the part that was always going to be phase two, and gets to say "solved" honestly, four months later than the deck claimed, but true when it finally landed.
What I'd tell my past self, sitting in that release meeting: a good number is not the same thing as a finished product, and the two feel identical for about thirteen weeks, right up until someone asks what actually changed.
PICK, for the day a right answer gets sold as a finished one
Not a script for staying calm when someone spots the gap. PICK only pays off if it makes you name, before anyone asks, the one test that tells solved from touched.
One path costs a slightly smaller deal today. The other already sent Bicuspid its real bill, in a screen-share, thirteen weeks in.
PPosition. The real distinction.
A problem AI solves is one where the model's output is the finished thing the user needed, an answer or an outcome they can act on directly, no extra step required. A problem AI merely touches is one where the output is a real, useful clue, but the actual fix is still sitting on someone's desk.
This isn't about how good the model is. Bicuspid's risk score was genuinely accurate, right 62% of the time on the flag. Accuracy was never the question. The question was whether anything changed for the patient once the model spoke.
State the position before the story, or it reads like it got reverse-engineered from the QBR complaint.
IImpact. What's lost each way.
Scope it honestly, call it a flag and nothing more, and the pitch loses a little shine, maybe a smaller first deal while the real fix gets built. That's recoverable inside a single sales cycle.
Oversell it, and the client finds the gap themselves, in production, in front of their own leadership, and now every other claim in the deck gets re-read with suspicion too.
Naming both losses is what keeps this from collapsing into "just be honest," which is a mood, not a decision.
CCost asymmetry. The heart of it.
Scoping the pitch honestly from day one costs a smaller first deal and a few extra weeks explaining a roadmap, cheap, and forgotten the moment phase two ships. Overselling it cost Bicuspid a six-week emergency sprint, about 240 unplanned hours pulled off other work, plus a renewal that came within one bad quarter of not renewing at all. Start from the cheap mistake. Only claim "solved" once the thing that earns the word actually exists.
KKill criteria. The one test.
Does the person still have to do the hard part themselves after the feature runs? For the flagged list alone, yes, someone still had to notice, still had to call, still had to hope the call worked. Lorand's team considered one shortcut instead of building the real action loop: pay front-desk staff a small bonus for every flagged patient they personally called. It got rejected. A bonus changes how hard someone tries. It doesn't change what the flagged patient actually receives, and two of Fettle's nine locations barely used it within a month anyway, since a bonus program is only as consistent as the manager running it.
The score was step one all along. Solved needed the other four parts wired to it.
Cost, by the numbers: scoping it honest first versus the emergency fix after overselling it
Cheap, scheduledExpensive, forced
Building the action loop as part of the normal roadmap would have cost about 160 hours over five planned weeks. Building it after the QBR complaint, compressed and pulled off other work, cost about 240 hours, and that number doesn't include the $184,000 renewal that sat at risk the whole time.
The kill line, charted: Fettle's no-show rate, weeks 0 to 20
Flag only, above the kill lineReal loop, cleared the kill line
Fettle's contract renewal was tied to getting the network no-show rate under 12% within two quarters. The flag alone spent thirteen weeks going nowhere near it. The action loop cleared it in six.
The trade worth saying out loud: building the real action loop before promising it meant Bicuspid couldn't say "solved" back in March, when the deal was hot and everyone wanted that exact sentence. It's worth paying, because a risk score with nothing behind it doesn't fill a chair, and a client who hears "fixes" and gets "flags" stops trusting every number you show them next, not just this one.
And if you want to be sure it really works, try it somewhere else
Same four letters, a truck instead of a dental chair, and the missing piece is a part number instead of a phone call.
Ironvent runs dispatch software for HVAC field service companies, including a model that scores every scheduled repair for callback risk: how likely a technician is to have to drive back out because the first visit didn't actually fix the problem. Mid-way through a growth push, someone on Ironvent's team proposes selling that callback score as "AI that eliminates repeat visits."
Same shape as Fettle's flagged list: a correct warning, still leaving the technician to guess.
Position: a problem AI solves is one where the model's output is the finished thing, here, the actual part sitting on the truck before the technician needs it. A problem AI merely touches is one where the output is a warning with nothing behind it, here, a callback-risk score that tells dispatch "this one might come back" while the technician still drives out guessing which part to bring. Impact: sell the score alone as a fix, and technicians keep making blind guesses on high-risk jobs, they just now know, uselessly, that they were about to guess wrong. Cost asymmetry: scoping the pitch as "flags likely callbacks" costs a smaller first rollout, cheap, three depots instead of twelve. Selling "eliminates callbacks" and shipping only a bare score costs a technician a wasted second truck roll on a homeowner's roof in July, expensive and visible to the customer, when the real fix means arriving the first time with the part that should already have been on the truck. Kill criteria: does the technician still have to guess which part to bring after the model runs? Today, yes, on every flagged job. Ironvent ships "flags likely callbacks" today, and holds "eliminates callbacks" until the predicted-failure-to-part-order pipeline actually exists.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: name the test, does the person still do the hard part after the model runs, before anything else.
Cost: no time to build the real action layer before a deadline. Fine, but sell the honest label for the piece that exists, and hold the bigger word for the piece that doesn't.
The model got better, for real: say precision climbs to 95%. Still doesn't solve anything on its own. A more accurate clue is still a clue until something acts on it.
Where people run it wrong.
They let a strong accuracy number get translated into a strong outcome claim, without checking whether anything downstream actually changed.
They do eventually build the action layer, but leave the oversold sentence sitting in every deck until then, so the trust is already spent before the real fix ships.
They treat "showing the user more information" as the same product as "doing the hard part for them," when those are two different builds with two different promises.
How to use it live. If you're ever asked whether a feature "solves" something, buy yourself a second with one plain question: "after it runs, does the person still have to do the hard part themselves?" That question is the whole method, asked out loud instead of stated.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a question asking you to draw a real line between two things that sound almost the same?
Tap to flip
ANSWER
PICK: state the position plainly, name the impact on each side, find which mistake is cheap versus expensive, then give the one test that tells them apart.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Lorand Kirkpatrick, who owns the no-show risk model at Bicuspid, a scheduling and reminder platform for dental practices.
3 · THE POSITION
What's the real line between a problem AI solves and one it only touches?
Tap to flip
ANSWER
Solves means the model's output is the finished thing, an outcome the person can act on directly. Touches means the output is a real, useful clue, but the actual fix is still sitting on someone's desk.
4 · THE GAP
What number proved the flagged list wasn't fixing anything?
Tap to flip
ANSWER
The model correctly flagged high-risk patients 62% of the time, and the no-show rate still only moved from 21% to 19% over thirteen weeks, because nothing changed for the patients it flagged.
5 · THE REVERSAL
What old decision would Lorand take back?
Tap to flip
ANSWER
The renewal pitch's line that Bicuspid "fixes no-shows," written before the action loop that would have earned that word existed. It made sense with a hot deal and a good accuracy number. It stopped making sense the day the client found the gap themselves.
6 · THE NUMBER
Fill in the blank: the flag sat on the screen for ___ weeks doing nothing extra. The real fix dropped the no-show rate from ___ to ___ percent over the next six weeks.
Tap to flip
ANSWER
Thirteen weeks. From 19 percent down to 9 percent.
7 · THE KILL TEST
What's the one test that tells a solved feature from a touched one?
Tap to flip
ANSWER
After the feature runs, does the person still have to do the hard part themselves? For the flagged list alone, yes. Once flagged patients got an earlier reminder, a guaranteed reschedule slot, and an automatic standby backfill, no.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question for a different product. Which one, and what plays the role of the flagged patient list?
Tap to flip
ANSWER
Ironvent's HVAC callback-risk score. The role goes to a technician who knows a job might come back, but still drives out guessing which part to bring, the same shape as a flag with no action behind it.
Check yourself Score: 0 / 0
True or false
1. True or false: since Bicuspid's model correctly flagged high-risk patients 62% of the time, the flagged-list feature had already solved Fettle's no-show problem.
True
False
Show hint
Think about what actually happened to a patient once they got flagged.
Show answer
False. Accuracy of the flag isn't the same as a changed outcome. Nothing acted on the flag for thirteen weeks, and the no-show rate barely moved, from 21% to 19%.
Multiple choice
2. What made the flagged-list feature a "touches" feature rather than a "solves" one?
A. The model wasn't accurate enough to trust.
B. Front-desk staff didn't like the new dashboard tile.
C. Flagged patients received the exact same reminder as everyone else, so nothing about the outcome changed.
D. The model took too long to score each morning's appointments.
Show hint
Check the direct answer's test: does the person still have to do the hard part themselves?
Show answer
C. The score was correct. What made it a "touches" feature was that nothing downstream of the score was different for a flagged patient.
Fill in the blank
3. The renewal Bicuspid nearly lost was worth about $___ a year across Fettle's ___ locations. The emergency sprint to build the real fix took about ___ hours over ___ weeks.
Show hint
Look at the "cost by the numbers" chart and its note, in the PICK recap section.
Show answer
$184,000 a year, across 9 locations. About 240 hours over 6 weeks. The gap between that and the planned 160 hours over 5 weeks is the cost asymmetry the whole no rests on.
Short answer, name the rejected alternative
4. Lorand's team considered a shortcut instead of building the real action loop. What was it, and why did it lose?
Show hint
Look at the K, kill criteria, step of the framework recap.
Show answer
Model answer: Pay front-desk staff a small bonus for every flagged patient they personally called. It got rejected because it changed how hard staff tried, not what the flagged patient actually received, and it was inconsistent across Fettle's nine independently-managed locations.
Short answer, where it wouldn't matter
5. Name a place in Fettle's own rollout where this same scrutiny wasn't needed, and say why.
Show hint
Look at "what I would leave alone" in the Let's learn section.
Show answer
Model answer: The clinical director's internal staffing dashboard, which shows the same risk score purely to plan chair capacity for the week. Nobody promised it would fix anything, so there's no gap between the claim and the delivery.
Short answer, apply it yourself
6. Think of an AI feature you've used that told you something true but didn't actually fix your problem. What would have needed to exist for it to count as solved?
Show hint
Ask whether you still had to do the hard part yourself after the feature ran.
Show answer
Model answer: A spam filter that flags an email "likely phishing" but still makes me decide whether to open it. It would count as solved if it also blocked the link automatically or quarantined the message outright, instead of just labeling it and stepping back.
Before you close the answer
Why this works
Tests whether you understand that "does the model work" and "is the problem solved" are two different questions, and that a probability score needs something built on top of it before it earns the word fix. Most candidates stop at the accuracy number.
Follow-up traps
"Isn't a good flag still valuable, even without the action loop?" Response: yes, and that's exactly why it should be sold as a flag, honestly. Its value never required calling it something it isn't.
"What if building the action loop takes too long to ever promise it?" Response: then sell the flag as the whole product, honestly, and treat the action loop as a second, later sale, not a line in a deck for a feature that doesn't exist yet.
If pressed
The model's calibration was built mostly on Kestbridge's own urban, high-volume history. Two of Fettle's locations are smaller and more rural, where distance to the clinic matters more as a no-show driver than the model had learned to weigh. Turning the action loop on everywhere without a per-location calibration check first would have quietly under-flagged exactly the clinics that needed it most.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.