ConceptIntermediateAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #9

Explain the difference between technical feasibility and product feasibility.

FLIPS89 percent right, and nobody left to tell

Say we build a tool that reads vibration and temperature off a building's cooling compressors and warns a field technician before one fails. Thermavane Controls calls theirs CircuitSense. Ingrid Pham-Solberg is the AI PM who proved it worked, then had to explain why proving it worked was not the same as it working.

The direct answer
Technical feasibility asks whether the model can be right, tested against held-out data. Product feasibility asks whether being right changes what a real person does, tested against their actual calendar. CircuitSense's model was 89 percent precise, a technically real result. It still failed as a product, because its warning arrived two to four days before failure while a technician's schedule was booked three weeks out, so being right almost never meant anything got fixed in time. Prove the model on its own terms first. Then prove the product on the person's terms, because those two questions can have completely different answers.
Do this, in order
  1. Test the model against a benchmark, then test the product against a real person's actual day.Why: those are two separate feasibility questions, and passing one says nothing about the other.
  2. Ask how much lead time a person needs to act, not just whether the flag is correct.Why: a correct alert that arrives too late to matter fails at the product layer even after passing the model layer clean.
  3. Watch how often people actually open or act on the tool, not just its accuracy, after launch.Why: a model can hold perfectly steady while the number of people still using it quietly falls to nothing.
  4. Explain why an alert fired and what window someone has, not just that it fired.Why: an unexplained flag teaches nobody what to do differently with the next one.
  5. Pilot with the people who have to act on it, not only the data that proves the model works.Why: technical feasibility gets proven in a notebook. Product feasibility only gets proven on someone's actual desk.
  6. Say plainly when the two line up, like a spell-checker where correct and useful are the same thing.Why: shows judgment about where the distinction actually matters, instead of treating every feature as a special case.

How to answer this, stage by stage

Nobody is scoring whether you know the two phrases exist. They're scoring whether you can point at the exact place a technically sound model stopped mattering to a real person.

Stage 1
Scope it to one real system
Say it like this
"Let's ground this in CircuitSense, which flags at-risk compressors for a field technician. That's the system where being technically right and being useful turned out to be two separate questions."
Why this works
Keeps the answer from becoming a dictionary definition with nothing real underneath it.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as FLIPS. Find the person whose morning this is. Locate the habit they built. Identify the flip, the verb that snaps. Pinpoint the old decision. Show the replay."
Why this works
Signals a repeatable way to reason about the gap, instead of a one-off memory of something going wrong.
Stage 3
Reframe: it isn't "can the model be right," it's "does being right change what someone does"
Say it like this
"This isn't really a question about whether the model clears a technical bar. It's a question of whether a correct answer, delivered on time, actually changes a real person's next move."
Why this works
This is where a strong answer separates from someone who just repeats that the two phrases mean different things.
Stage 4
Give the flip
Say it like this
"Here's the flip: technicians opened every CircuitSense alert the first month, since it was new and it sounded serious. Once opening one usually meant driving out to find nothing schedulable in time, they didn't complain, they just quietly stopped opening the app at all. There was no in-between."
Why this works
Names the real behavior change as a two-setting switch, not a vague warning about "adoption."
Stage 5
Prove it with the compressed evidence
Say it like this
"The model held at 89 percent precision the entire time. App opens fell from 18 out of 20 alerts in month one to 2 out of 20 by month five. Nobody filed a complaint. The dashboard just quietly emptied out."
Why this works
Compresses the whole case into two numbers moving in opposite directions for no visible reason.
Stage 6
Name the AI-specific reasoning and the trade-off
Say it like this
"The honest reason this isn't a generic adoption problem is that the model's own accuracy number gave zero warning it was happening. Technical feasibility was proven and stayed proven the whole time. We accepted a slower rollout, testing the alert's timing against real schedules before ever measuring model accuracy again, in exchange for never mistaking a quiet product failure for a stable one."
Why this works
This is the load-bearing, AI-specific judgment: a model's benchmark score can stay perfect while the product built on top of it silently dies.
Stage 7
Say what wouldn't change
Say it like this
"I wouldn't run this same lead-time check on a low-stakes alert, like 'this air filter is due for a wash,' where there's no scheduling tension and the model's own accuracy is basically the whole story."
Why this works
Shows judgment about where the distinction matters, instead of applying it as blanket caution everywhere.
Stage 8
Close on the one line
Say it like this
"Technical feasibility is whether the model can be right. Product feasibility is whether being right, on time, changes what somebody actually does. CircuitSense proved the first one and never checked the second."
Why this works
Restates the direct answer in one breath and leaves the interviewer with the exact line that carries the whole answer.

Let's learn

Say we build a tool that reads vibration and temperature sensors on a building's cooling compressors and warns a technician before one actually fails.

Before CircuitSense, Thermavane's technicians serviced every compressor on a fixed 90-day schedule regardless of its real condition, about 40 minutes a unit, catching a real problem mostly by luck. With CircuitSense, the model reads live sensor data and flags a specific unit as at risk, with the flag correct 89 percent of the time against Thermavane's own held-out test data.

Hand sketched icon list titled FLIPS the five letters. Five rows: Find the person whose morning this is. Locate the habit what they stopped doing. Identify the flip the verb that snaps, shown in a different color. Pinpoint the old decision what only made sense before. Show the replay same day, new design.
The five letters, held up as one page. Identify the flip is the step a purely technical review never reaches.

Here's the turn: the 89 percent precision was real and it held steady for months. What nobody had separately tested was lead time: the flag usually arrived two to four days before a likely failure, while a technician's real schedule was booked three weeks out. Being right almost never lined up with being able to do anything about it.

Hand sketched labeled parts diagram titled What technical feasibility never asked. A gauge icon at the center labeled 89 percent Precise, with four labeled callouts around it: How much lead time a fix needs, Who has to act on it, What happens when it's wrong, Whether the app fits their day.
Four questions a benchmark score answers nothing about.

At its worst, a compressor in an occupied building genuinely fails during a week nobody opened CircuitSense at all, and it turns out the flag had been sitting there, correct, for days.

The model held steady. What people did with it did not.
100% 50% 0 89% Model precision 22% Acted on in time
One bar measures the model. The other measures whether being right ever reached a wrench.
The model was never wrong about the compressor. It was wrong about what a technician could still do with two days' notice.
The choice I would take back CircuitSense's alert shipped as a single line: unit flagged, at risk. It never said how much time was left or what evidence the model was reading. That made sense when the team was racing to prove the model worked at all. It stopped making sense the moment a technician acted on a flag, found nothing schedulable in time, and had no way to understand why or what to expect from the next one.

What I would leave alone: for a low-stakes alert, like a filter due for a routine wash, the model's own accuracy really is the whole story. There is no scheduling tension to separately test.

The lesson: a model can be exactly as good as its benchmark says, and the product built on top of it can still fail completely, for a reason the benchmark was never built to catch.

Now here is the same thing as a story

The short version above is what you'd say defending a rollout in a planning review. Read this one for what it felt like the slow, quiet months before anyone thought to check.

For eight months, CircuitSense's dashboard looked healthy. Then one afternoon, Ingrid looked at a graph she hadn't opened in a while.

Hand sketched flow diagram titled How a CircuitSense alert reaches a technician, third step emphasized. Five steps left to right: Sensor reads vibration. Model scores risk. Push alert sent, shown in a different color. Technician opens app. Unit gets scheduled.
The third step is where a technically sound model quietly stopped reaching anyone.

Ingrid had shipped CircuitSense after months of model validation, precision, recall, calibration, all clean on held-out sensor data. Technicians opened every alert eagerly at first, curious what the new tool would catch.

Hand sketched comparison titled Small move, big snap. Left panel, a gauge icon labeled Model score, caption quietly, steadily 89 percent precise. Right panel, a question mark box icon labeled App opened each week, caption fine, fine, fine, then never, shown in a different color.
One number held steady. The other one quietly went to nothing, with no single bad day to point at.

There was no complaint, no support ticket, no single bad incident. A technician would drive out on a flag, find a two-day warning window against a three-week backlog, and learn, quietly, that opening the app rarely changed anything he could actually do that week.

Knowledge spark: why would a technically sound model still fail as a product? A model is only tested against the data you show it. CircuitSense's benchmark measured whether the flag matched a real failure. It never measured whether the flag arrived with enough lead time for anyone to act on it, because that question lives in a technician's calendar, not in the sensor data.

Over five months, the habit thinned in three quiet beats: eighteen of twenty alerts opened in month one, then about half by month three, then two of twenty by month five. Nobody flagged it, because nothing broke loudly enough to notice.

Hand sketched comparison titled The two blocks. Left panel, a document icon labeled Month 1, caption 18 of 20 alerts opened. Right panel, a document icon labeled Month 5, caption 2 of 20 alerts opened.
Same tool, same accuracy, and the number of people still listening had nearly disappeared.

The real question was never whether CircuitSense was 89 percent accurate. It was whether being 89 percent accurate had ever been tested against what a technician could realistically still do about it.

Hand sketched metaphor scene titled Switch, not dial. Left, a gauge icon labeled Assumed, caption a dial, technicians open it a bit less each week. Right, a vending machine icon labeled Actual, caption a switch, they open it or they never open it again.
One full-page image to carry the whole answer: adoption did not fade. It flipped, one technician at a time.

When CircuitSense's alert format was first designed, someone said, "keep it simple, just flag the unit," and it sounded reasonable, since at the time the whole team was focused on proving the model could be right at all.

Rerun the same five months with an alert that names the lead time and the sensor reasoning behind every flag: technicians who see "four days, rising vibration on bearing two" instead of a bare warning re-engage, opens hold near 80 percent by month five, and a genuinely at-risk unit gets caught with real time to act.

What I'd tell myself, watching that graph go quiet: 89 percent was never the number that mattered. The number that mattered was two out of twenty, and it took eight months to think to look.

FLIPS, the difference that finally explained itselfNot a script for distrusting every accurate model. FLIPS is what tells you exactly which half of "it works" was never actually tested.

F
Find the person. Whose morning is this?
Ingrid Pham-Solberg, the AI PM who validated CircuitSense's model and owns its rollout to Thermavane's field technicians.
A specific person with a specific dashboard, not an abstract "usage declined."
L
Locate the habit. What did they stop doing because it worked?
Technicians stopped opening every CircuitSense alert the moment it arrived, once they learned a flag usually came with too little lead time to act on.
The habit formed quietly and reasonably. Nobody decided to stop caring on purpose.
I
Identify the flip. What verb snaps?
Opening every alert versus never opening the app at all. No middle setting once a technician learned what the warning window actually meant for his own schedule.
This is the hardest step, and the one that separates technical feasibility from product feasibility: the model's accuracy never moved, only the person's behavior did.
P
Pinpoint the old decision. Which choice only made sense before?
Shipping the alert as a bare flag, with no lead time and no reasoning attached, to keep the first release simple while the model was still being proven.
Reasonable while proving the model came first. Wrong once the model was proven and nobody revisited the alert itself.
S
Show the replay. Same trigger, new design.
The next alert design states the lead time and the sensor reasoning. Technicians re-engage, opens climb back toward 80 percent, and a real at-risk unit gets caught with days to spare instead of none.
A countable result: opens near 80 percent instead of 2 out of 20, and a caught failure instead of a missed one.

The recap, one line per letter: find the person is Ingrid, who owns CircuitSense's rollout, locate the habit is opening every alert on arrival, identify the flip is opening every alert versus opening none, pinpoint the old decision is the bare, unexplained flag shipped to keep the first release simple, and show the replay is an alert with lead time and reasoning that brings technicians back.

And if you want to be sure it really works, try it somewhere elseSame five letters, a fishing cooperative instead of a machine shop. The flip changes families entirely, the missing second test doesn't.

Tomasz Wibowo runs product at Saltmere Fisheries Cooperative, where StockGauge estimates fish counts from sonar and net-haul photos so biologists can set the next season's catch limits. Its technical benchmark, built on clear-water test images, hit 93 percent accuracy. Mapped onto FLIPS: find the person is Tomasz, who signs off on StockGauge's estimates. Locate the habit is biologists spot-checking roughly one haul in ten by hand, trusting the rest. Identify the flip here is a verification flip, not abandonment: once biologists noticed StockGauge's accuracy quietly worsened in murky, sediment-heavy water the benchmark never included, they stopped spot-checking and started recounting every single haul by hand, since a spot-check could no longer tell them which hauls to trust. Pinpoint the old decision is building the technical benchmark entirely from clear-water images, because that was the easiest data to collect first. Show the replay adds a turbidity reading to every estimate, so biologists spot-check only the hauls flagged as murky, instead of recounting everything.

Hand sketched comparison titled Spot check the count, or check every haul. Left panel, a document icon labeled Before, caption biologist spot-checks 1 in 10 hauls. Right panel, a document icon labeled After, caption biologist recounts every single haul.
A different flip entirely: not abandonment, but a spot-check that stopped being enough once conditions the benchmark never saw showed up.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "technical feasibility is can it be right, product feasibility is does being right change what someone does," and stop.
Cost: there's no time to test the product layer before a launch deadline. Say so honestly, and ship the model with a clearly marked "not yet tested against real schedules" label, rather than letting silence imply both questions were answered.
The model gets better, for real: if CircuitSense's lead time improves to two full weeks, that's still worth re-testing against a technician's real backlog, not assumed safe just because the accuracy number went up.

Where people run it wrong.
They treat a strong benchmark score as proof the whole product works, not just the model inside it.
They watch model accuracy after launch and never separately watch whether people are still acting on it.
They let a quiet decline in usage hide behind a healthy-looking accuracy dashboard for months.

How to use it live. The moment someone says a model is technically feasible, ask yourself: feasible to be right about what, and does being right on time change anything a real person can actually do. Name both, and the distinction explains itself.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Abandonment flip: technicians didn't complain about CircuitSense, they just quietly stopped opening it once a flag rarely arrived with enough lead time to act on.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Ingrid Pham-Solberg, the AI PM at Thermavane Controls, who proved CircuitSense's model worked and later found technicians had quietly stopped using it.
3 · THE HABIT
What did technicians stop doing, over five months, because it stopped paying off?
Tap to flip
ANSWER
They stopped opening CircuitSense's alerts at all, going from 18 of 20 opened in month one to 2 of 20 by month five.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Opening every alert on arrival, versus never opening the app at all. No middle setting once a technician learned the warning window rarely matched his real schedule.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Shipping every alert as a bare flag, with no lead time and no reasoning attached, to keep the first release simple while the model was still being proven.
6 · THE NUMBER
Fill in the blank: the model held steady at 89 percent precision, while only ___ percent of flags were actually acted on in time to matter.
Tap to flip
ANSWER
22 percent.
7 · THE REPLAY
Same five months, alerts now show lead time and reasoning. What changes?
Tap to flip
ANSWER
Technicians re-engage, opens climb back toward 80 percent, and a genuinely at-risk compressor gets scheduled with real days to spare instead of getting missed entirely.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Saltmere Fisheries Cooperative's StockGauge. The flip is verification: biologists stopped spot-checking and started recounting every haul once murky water conditions the benchmark never tested started showing up.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: technical feasibility asks whether the model can be ___, while product feasibility asks whether being right changes what a real ___ does.
Show hint
Look at the direct answer.
Show answer
Right; person. A model can pass every technical test and still fail as a product if being right never reaches someone in time to act.
Multiple choice
2. Why did CircuitSense's 89 percent precision fail to predict that technicians would stop using it?
  • A. The model's accuracy actually dropped, but nobody noticed the drop.
  • B. The benchmark only measured whether the flag was correct, never whether its lead time gave a technician enough time to act.
  • C. Technicians were never trained on how to use the app.
  • D. The sensors themselves were faulty.
Show hint
Look at the knowledge spark about why a technically sound model can still fail as a product.
Show answer
B. Lead time versus a real schedule is a product-layer question a model benchmark was never built to answer.
True or false
3. True or false: technicians filed a support ticket complaining that CircuitSense's warnings arrived too late.
  • True
  • False
Show hint
Look at "locate the habit" in the FLIPS recap.
Show answer
False. Nobody complained. They simply stopped opening the app, which is why the decline showed up only when Ingrid happened to check a graph months later.
Short answer, where it wouldn't matter
4. Name a CircuitSense alert type where technical feasibility and product feasibility would already be the same question, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A routine "filter due for a wash" alert. With no scheduling tension involved, the model's own accuracy is basically the whole story, so there's no separate product-feasibility question to test.
Short answer, apply it yourself
5. Think of a tool you stopped using even though it was technically working fine. What did it fail to fit into, in your actual day?
Show hint
Think about the gap between a tool being correct and a tool being something you could act on.
Show answer
Model answer: A budgeting app that correctly flagged overspending after the money was already spent. Being right came too late to change any decision, so checking it stopped feeling worth the time.
Short answer, work the number
6. If CircuitSense's lead time grew from two to four days up to ten to fourteen days, would you expect the 22 percent acted-on-in-time figure to improve on its own?
Show hint
Think about what was actually limiting that 22 percent figure in the first place.
Show answer
Model answer: Likely yes, since a technician's three-week backlog could actually absorb a ten to fourteen day warning, unlike a two to four day one. This is exactly the product-layer test that never ran alongside the model's own accuracy check.
Before you close the answer
Why this works
Tests whether you'll treat a strong benchmark score as proof the whole product works, or ask separately whether being right ever reaches someone in time to matter.
Follow-up traps
"Isn't this just a generic change-management problem?" Response: no, because the model's own accuracy stayed perfectly healthy the entire time, which is exactly what let the real failure hide behind a dashboard that looked fine.

"Couldn't you have caught this by asking technicians directly?" Response: you could, and now that's a standing check, but the deeper fix is testing lead time against real schedules before launch, not waiting to find out from a quiet usage graph months later.
If pressed
The redesigned alert also logs which lead-time bucket, under 3 days, 3 to 7, over 7, produced the highest rate of an alert actually turning into a scheduled fix, so the product-feasibility question gets its own ongoing metric instead of being tested once and forgotten.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more