CaseIntermediateAI Opportunity & Model Strategy / When NOT to use AI / #20
What should you do when the AI feature works but the workflow around it does not exist?
ORDER91 percent accurate extraction, 23 percent of it ever actually happened
Minutewise builds ActionLoop, a tool that transcribes meetings and pulls out action items automatically. Sindre Aabel is the engineer who built the extraction model, and it works, genuinely. Chiamaka Ndife is the PM who has to explain why a feature that works this well is barely changing what teams actually get done.
The direct answer
Pause further work on the model and build the missing workflow first: an assignee, a due date, a notification, a real place the item lives. A model's output only creates value once something exists to receive it, act on it, and follow through. An accurate extraction sitting in a list nobody revisits isn't a smaller win than a real task, it's not a win at all yet.
Do this, in order
Walk the model's output through the real path a user actually takes with it today.Why: this is the only way to see the dead end for yourself, instead of trusting a dashboard that only tracks the model's own accuracy.
Pause investment in extraction accuracy until the workflow gap is closed.Why: a more accurate model produces a more accurate list that still goes nowhere. That's not the bottleneck right now.
Build the smallest real destination: an assignee, a due date, a notification.Why: this is the smaller, reversible bet that reveals what the model actually needs to support, before committing to anything bigger.
Measure real completion, not extraction accuracy, as the metric that matters.Why: 91 percent precision and 23 percent completion are two completely different numbers, and only the second one reflects real value delivered.
Only return to model improvements once the workflow proves it can carry more volume.Why: better extraction is a real investment worth making later, once there's a working destination for it to actually matter.
Recheck the dependency if the workflow or the model changes significantly.Why: what unblocks what can shift once the workflow exists and starts revealing its own new bottlenecks.
How to answer this, stage by stage
Nobody is scoring whether you know what "workflow" means. They're scoring whether you'd have caught the dead end before six months of model tuning did nothing for it.
Stage 1
Scope it to one real feature, not "workflow design" in the abstract
Say it like this
"Let me give you a real case. A meeting tool extracts action items at 91 percent precision, genuinely accurate. Only 23 percent of those items ever get done within 30 days. The model isn't the problem, there's nowhere real for its output to go."
Why this works
Grounds the answer in a checkable gap instead of a general point about product-market fit.
Stage 2
Say your structure out loud before any content
Say it like this
"I'll run this as ORDER. Outcome, what actually has to be true for this to create value. Reversibility, which investment is the smaller bet. Dependency, what actually blocks what. Evidence, the cheap check that proves it. Rank, the real call on what to build next."
Why this works
Signals a repeatable way to sequence work, not a vague sense that "adoption" needs improving.
Stage 3
Reframe the question: this isn't an accuracy problem
Say it like this
"The instinct is to make extraction even more accurate, get it to 95, 97 percent. That instinct is wrong here. A more accurate list that nobody ever opens again is still a list nobody ever opens again. The real question is what happens after extraction, not during it."
Why this works
This is where a strong answer separates from just proposing more model work by default.
Stage 4
Give the one decision: build the destination first
Say it like this
"Here's what I'd actually do. Pause the extraction accuracy roadmap. Build the smallest real destination for an action item: an owner, a due date, and a notification that actually fires. That's a smaller, more reversible bet than more model work, and it tells us what the model needs to support next."
Why this works
This is the direct answer, stated as an actual sequencing decision, not just a diagnosis of the problem.
Stage 5
Prove it with the compressed failure
Say it like this
"This is exactly what Chiamaka found. The exact same wording of action item, typed by hand into a real task tracker with an owner and a due date, got done 74 percent of the time. Extracted by the model into a list with neither, it got done 23 percent of the time. Same words, same person, wildly different fate."
Why this works
This is where the story lives, compressed to the one comparison that proves the model was never the bottleneck.
Stage 6
Name the AI-specific trap in this failure mode
Say it like this
"The specific trap with a model here is that its own metrics look great in isolation, 91 percent precision is a genuinely strong number, and it's easy to keep optimizing a number that's already climbing instead of checking whether the output actually leads anywhere. A working model with no destination will always look like a success on its own dashboard."
Why this works
This is the load bearing judgment. It wouldn't make sense to ask this about a feature with no model in it, since it's specifically about a model's own strong internal metric masking a real downstream gap.
Stage 7
Say what you'd leave alone, then close on one line
Say it like this
"This isn't a case against the extraction model, it's genuinely good at its job and stays exactly as it is. It's that a good model with no real destination for its output isn't done yet, no matter how strong its own number looks. Build the destination first, then come back to the model."
Why this works
Closes with real judgment instead of dismissing the model work entirely, and restates the direct answer in one breath.
Let's learn
ActionLoop is a tool that listens to a meeting and automatically pulls out the action items people agreed to, so nobody has to sit there taking notes on who's doing what.
Four honest steps, and the fourth one is where the value quietly disappears.
The extraction model itself is genuinely strong: tested against 500 human-labeled meeting transcripts, it correctly identified real action items 91 percent of the time. Sindre's team had spent two quarters getting it there, and the number climbed steadily the whole way.
Completion within 30 days, same action item, two different homes
Real destination existsNo real destination
Same words, same person, same meeting. The only real difference is whether the item landed somewhere with an owner and a date attached to it.
Then Chiamaka pulled 30 days of usage data and found completion rates telling a very different story than the extraction accuracy dashboard did. Correctly identified action items, sitting in a passive list inside the transcript view, got acted on barely a quarter of the time.
The model wasn't failing. The output just had nowhere real to go once it was right.
Here's the turn: this was never about the model needing to get more accurate. The turn is that "correctly extracted" and "actually acted on" are two completely different outcomes, and the second one, the one that actually matters, depends entirely on something the model was never asked to build.
Knowledge spark: why doesn't a strong model metric guarantee real value?
A model's own metric, precision, recall, accuracy, measures how well it does its specific job. It says nothing about what happens after that job is done. A perfectly accurate extraction with no destination creates the same real-world outcome as no extraction at all: nothing changes for the team.
Completion rate, week by week, before and after the workflow shipped
No destination yetAfter the fix
Zero changes to the extraction model between week 3 and week 6. The completion rate nearly tripled anyway, because the fix was never in the model.
At its worst, this cost showed up as two more quarters of model refinement work, real engineering time spent pushing extraction precision from 87 to 91 percent, while the number that actually mattered to teams, real work getting done, never moved at all.
The choice I would take back
Treating extraction accuracy as the roadmap's headline metric for two full quarters, without ever checking whether the output was actually leading anywhere. That made sense when the model was new and clearly the riskiest part of the build. It stopped making sense once the model was already good and the real gap had moved somewhere else entirely.
What I would leave alone: the extraction model itself needs no further work right now, 91 percent precision is a genuinely strong number for this task. The fix was never in the model, it was in what happens the moment after it finishes.
The lesson: a model's own metric can keep climbing while the real value it's supposed to create stays completely flat. Walk the output through the real path a user takes with it, all the way to the end, before trusting the model's own dashboard to tell the whole story.
Now here is the same thing as a story
The short version above is what you'd say out loud in the room. Read this one for what it actually felt like to watch a strong model keep improving while nothing else changed.
Sindre Aabel had built the extraction model over two quarters, watching precision climb from a rough 68 percent at first prototype to a genuinely strong 91 percent. Every sprint review, the number went up, and every sprint review felt like real progress.
Chiamaka's job was adoption, and adoption wasn't moving the way the accuracy chart said it should. She started where most people start: assuming the model still wasn't good enough. Maybe 91 needed to be 95.
The dependency ran the opposite direction from what the roadmap had assumed for two quarters.
Before proposing another accuracy push, she did something simple: she used ActionLoop herself, for real meetings, for two full weeks, and tracked what actually happened to every item it extracted for her.
Fourteen of nineteen extracted items, correctly identified, real things she'd agreed to do, just sat in the transcript view. She hadn't ignored them on purpose. She'd simply never gone back to that page, because nothing ever told her to.
The same sentence, the same person, two completely different outcomes depending on nothing but where it landed.
She ran the comparison properly after that: the exact same kind of commitment, typed by hand into the team's actual task tracker with an owner and due date attached, against one extracted by ActionLoop into its own list. Manually typed items got done 74 percent of the time within 30 days. Extracted items got done 23 percent of the time.
We spent two quarters making the model better at finding the right words. We never once asked what happened to those words after it found them.
Sindre, when she showed him, didn't get defensive. He got curious in the wrong direction at first: "Should I add a confidence score, so people trust the extracted ones more?" That would have made the list a little more informative. It still wouldn't have given anyone a reason to open it.
Chiamaka never had a fixed number for when a feature had a real workflow gap versus a real accuracy gap. She had a feeling with two settings: the output leads somewhere real, or it doesn't. Nineteen extracted items and fourteen dead ends made that setting obvious.
One door commits real engineering weeks to a problem that might not be the problem. The other is small enough to try and learn from fast.
Back when the roadmap first prioritized extraction accuracy, it wasn't an unreasonable call. The model was new and clearly the riskiest, least proven part of the whole feature. It stopped being the right call the moment it crossed 90 percent and kept climbing while nothing downstream changed at all.
The fork that should have run before the second quarter of model work ever started.
Here's the replay: same model, same 91 percent, but the next sprint builds an owner field, a due date, and a Slack notification instead of another point of precision. Completion climbs from 23 to 68 percent over the following month, with zero changes to the extraction model itself.
One version of this story keeps polishing a number that was already good enough, while the real value stays flat for two quarters. The other spends a fraction of that time on the actual bottleneck and triples what teams get done, without touching the model at all.
What I'd tell myself, watching Sindre reach for a confidence score instead of a due date field: a model's own number getting better will always feel like progress. Whether it's the right progress depends entirely on what happens the moment after it finishes.
ORDER, run on a metric that kept climbing while nothing else movedNot a script for dismissing a model that's already working. ORDER is what tells you whether the next real bottleneck is even in the model at all.
O
Outcome. What actually has to be true for this to create real value?
An accurate action item only creates value once a real workflow exists to receive it, assign it, and remind someone to follow through. A correct extraction with nowhere to go isn't delivering value yet, no matter how good its own number looks.
This is the outcome the roadmap should have been measured against from the start.
R
Reversibility. Which investment is the smaller, safer bet?
More model work is a bigger, less reversible bet on a capability nobody can fully use yet. Building an owner field and a due date is small, cheap to try, and reveals exactly what the model actually needs to support next.
Reversibility, not ambition, is what should decide which door opens first.
D
Dependency. What actually blocks what?
The model's value depends entirely on the surrounding workflow existing, not the other way around. Two quarters of accuracy work assumed the dependency ran the opposite direction.
This is the direct answer's real mechanism: get the dependency backwards, and every improvement lands on the wrong side of it.
E
Evidence. What's the cheap check that proves it?
Walk the model's real output through the actual path a user takes with it, all the way to the end. Chiamaka's two weeks using ActionLoop herself, and the 74 percent versus 23 percent comparison, did exactly that.
A real, walked path beats a dashboard that only ever measures the model's own job.
R
Rank. What's the actual practical call?
Pause extraction accuracy work. Build the smallest real destination first, an owner, a due date, a notification. Return to model improvements only once that destination proves it can carry more volume.
This is the direct answer to the question, turned into a concrete sequencing decision.
The recap, one line per letter: outcome is what actually has to be true for value to exist, reversibility is picking the smaller, safer bet first, dependency is getting the real blocking relationship right, evidence is the cheap walked-path check that proves it, and rank is the concrete call on what to build next.
And if you want to be sure it really works, try it somewhere elseSame five letters, a resume-screening tool instead of a meeting assistant. This time the missing workflow is on the hiring manager's side, not the applicant's.
Delphine Okonjo runs product at Rolecraft, a hiring platform. Its AI resume screener correctly ranks strong candidates near the top 88 percent of the time, verified against past hiring outcomes. Hiring managers still complain they're not finding good candidates faster. Mapped onto ORDER: outcome is that a correct ranking only creates value once a hiring manager actually reviews the top-ranked candidates differently than before. Reversibility means building a lightweight "reviewed" tracker for hiring managers is a small, safe bet, versus pushing ranking accuracy from 88 to 93 percent, a bigger, unproven investment. Dependency is that the ranking's value depends on hiring managers actually changing their review order, not the other way around, and nobody had confirmed that was happening. Evidence means watching real hiring managers use the tool, and finding most still opened every resume top to bottom in application order, ignoring the AI ranking entirely, because nothing in the interface made the ranking hard to ignore. Rank means building a default sort-by-ranking view and a visible "AI-recommended" badge before investing further in ranking accuracy.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Skip straight to "walk the output through the real path a user takes with it, and fix the dead end before the model," and stop.
Cost: no engineering time to build a full workflow fix right away. Ship the smallest possible version, one due-date field and one notification, and measure before committing to more.
The model got better, for real: say extraction accuracy climbs to 97 percent on its own. The workflow gap doesn't close on its own, a perfect list with no owner still gets ignored the same way a 91 percent one did.
Where people run it wrong.
They keep optimizing the model's own metric because it's the number that's easiest to move and easiest to show progress on.
They assume low adoption always means the model isn't good enough yet, without ever checking where the real dependency actually runs.
They measure success by what the model produces, not by what a real user does with it afterward.
How to use it live. The moment an interviewer describes a working AI feature with disappointing real-world impact, ask yourself first: does a real workflow exist for its output, or does it just look like one exists on a diagram. That question buys real thinking time, and it's usually exactly where the true bottleneck is hiding.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits deciding what to build next when a working AI feature has no real destination?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. It sequences work by what's actually blocking value, not by which metric is easiest to keep improving.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Chiamaka Ndife, the PM who traces the real gap at Minutewise's ActionLoop. Sindre Aabel is the engineer who built the extraction model.
3 · THE MECHANISM
Why did a strong model metric fail to predict real value here?
Tap to flip
ANSWER
A model's own metric measures how well it does its specific job, not what happens afterward. A correct extraction with no owner, no due date, and no notification creates the same real outcome as no extraction at all.
4 · THE FIX
What's the one concrete thing this answer says to actually do?
Tap to flip
ANSWER
Pause extraction accuracy work and build the smallest real destination first: an owner field, a due date, and a notification that actually fires.
5 · THE OLD DECISION
What old decision would this answer take back?
Tap to flip
ANSWER
Treating extraction accuracy as the roadmap's headline metric for two full quarters without ever checking whether the output led anywhere real. Reasonable while the model was new and risky. Wrong once it was already strong and the real gap had moved elsewhere.
6 · THE NUMBER
Fill in the blank: the same action item got done ___ percent of the time typed into a task tracker, versus ___ percent extracted by AI with no home.
Tap to flip
ANSWER
74 percent, versus 23 percent. Same words, same person, the entire difference was whether the item landed somewhere with an owner and a date.
7 · THE REPLAY
Same feature, new sequencing, what changes?
Tap to flip
ANSWER
Completion climbs from 23 to 68 percent over one month, with zero changes to the extraction model. The fix was never in the model, it was in the workflow around it.
8 · CROSS PRODUCT TRANSFER
Section 4 runs ORDER again on a different product. Which one, and what's the equivalent evidence check?
Tap to flip
ANSWER
Rolecraft's resume-screening tool. The equivalent evidence check watches real hiring managers use it, and finds most still review resumes in application order, ignoring the AI's ranking entirely.
Check yourself Score: 0 / 0
Multiple choice
1. Why did Chiamaka's team keep improving extraction accuracy for two quarters instead of fixing the real problem sooner?
A. The extraction model was actually performing poorly the whole time.
B. The model's own accuracy metric kept climbing and looked like real progress, hiding the fact that output had nowhere real to go.
C. Sindre refused to build any workflow features.
D. Users never used ActionLoop at all.
Show hint
Look at Stage 6 of the walkthrough.
Show answer
B. A strong, climbing model metric feels like progress, and it's easy to keep optimizing it instead of checking whether the output actually leads anywhere.
True or false
2. True or false: this answer concludes the extraction model needs significant further improvement.
True
False
Show hint
Look at "what I would leave alone."
Show answer
False. The model already performs well at 91 percent precision. The fix was never in the model, it was in building a real destination for its output.
Fill in the blank
3. Fill in the blank: after the workflow shipped, completion rate climbed from 23 percent to ___ percent by week 6, with no changes to the model.
Show hint
Look at the line chart in "Let's learn."
Show answer
68 percent. Nearly tripling the completion rate, proof the model was never the actual bottleneck.
Short answer, where it wouldn't matter
4. Describe a situation where pausing model work in favor of workflow fixes would NOT be the right call, unlike this case.
Show hint
Think about a feature where the model's own accuracy is genuinely still the main complaint.
Show answer
Model answer: If a real destination already exists and users consistently say the extracted items are wrong or unhelpful, the real bottleneck is genuinely still the model, and workflow fixes wouldn't address that complaint.
Short answer, apply it yourself
5. Think of an AI feature or automated output you've used that felt accurate but didn't change what you actually did. What was missing downstream of it?
Show hint
Think about a summary, a recommendation, or a flagged item that you saw and then never acted on.
Show answer
Model answer: A budgeting app's AI-flagged "unusual spending" alerts were accurate but sat in a notifications tab nobody checked. Nothing prompted a real decision, like pausing a subscription, so the accurate flag never changed any actual spending.
Short answer, work the number
6. If extraction precision had been 99 percent instead of 91 percent, but the workflow gap stayed exactly the same, would the completion rate likely change much?
Show hint
Think about what completion rate actually measures versus what precision measures.
Show answer
Model answer: Probably not much. Completion rate depends on whether an item has an owner, a due date, and a reminder, none of which precision affects. A more accurate list with the same missing workflow would likely still get ignored at close to the same rate.
Before you close the answer
Why this works
Tests whether you'll trust a model's own strong metric at face value, or actually walk its output through the real path a user takes, and whether you can correctly identify when the bottleneck has moved out of the model entirely.
Follow-up traps
"Couldn't a confidence score have helped, like Sindre suggested?" Response: a confidence score makes the list more informative, but it still doesn't give anyone a reason to open the list in the first place, it treats a workflow gap like an accuracy gap.
"What if the workflow fix doesn't actually raise completion either?" Response: then the evidence points somewhere else, maybe the items themselves need better phrasing, but you've ruled out the cheapest, most reversible fix first instead of guessing at a bigger one.
If pressed
The 74 percent versus 23 percent comparison controlled for the person and the wording of the item itself, using the same individuals' own commitments in both conditions, so the gap could be attributed to the destination, not to who was involved.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.