ConceptIntermediateAI Opportunity & Model Strategy / Build vs buy vs fine-tune decisions / #3

Explain the strategic risk of building your core differentiator on a third-party model.

FLIPSthe risk never showed up as the model getting worse, it showed up as a pharmacist quietly trusting it more

Talwick Health sells software that pharmacy chains use to check for drug interactions. InteractIQ is the feature meant to be Talwick's actual edge, an explainer that tells a pharmacist why two drugs are flagged and how serious it is, built entirely on top of one outside vendor's model. Zanele Mbatha is the AI PM who owns that decision, and Deacon Whitlock is the pharmacist at Millrose Pharmacy whose one near miss showed the whole company what they'd actually built their edge on.

The direct answer
The risk is that your differentiator's actual behavior can change on a day you don't choose, because the vendor owns the model version, not you. A silent update can shift tone, drop a caveat, or soften a warning with no announcement, and your users will recalibrate how much they trust it before your own team even notices something moved.
Do this, in order
  1. Pin the exact model version your differentiator runs on, and upgrade only after your own re-validation.Why: this is the actual fix. Without it, the vendor's release calendar is quietly your release calendar too.
  2. Keep a visible changelog of every underlying model change, even ones the vendor calls minor.Why: without one, a real behavior shift looks exactly like an ordinary day, until someone gets hurt by it.
  3. Run your own eval set against every candidate upgrade before it goes live.Why: it catches a regression in a forty-minute test, not live, in front of a patient.
  4. Design a real fallback for the moment trust breaks, not just a warning label.Why: once someone stops checking, they need somewhere to land besides checking literally everything again.
  5. Watch for the update that reads as an improvement, not just the one that reads as a bug.Why: a model getting more fluent or more confident-sounding is still a change, and it erodes a checking habit just as fast as a regression does.
  6. Never treat "the vendor has been reliable" as a permanent fact.Why: it's a fact about last quarter, not a guarantee about their next release.

How to answer this, stage by stage

Nobody is scoring whether you can name "vendor lock-in" as a risk. They're scoring whether you can show exactly how that risk reaches a real person's hands.

Stage 1
Scope it to one real feature
Say it like this
"Let me ground this in one real case. Talwick Health sells pharmacy software. InteractIQ is supposed to be Talwick's actual differentiator, it explains why two drugs are flagged, not just that they are. It's built entirely on one vendor's model. Zanele Mbatha, the AI PM, owns that call."
Why this works
Keeps the answer from becoming a general lecture on vendor risk with no real feature behind it.
Stage 2
Say your structure out loud
Say it like this
"I'll run this as FLIPS. Find the person, whose morning is this. Locate the habit, what he stopped doing because it worked. Identify the flip, the verb that snaps. Pinpoint the old decision, which choice only made sense before. Show the replay, same bad day, fixed design."
Why this works
Signals a repeatable way of finding the real risk, instead of a list of abstract worries about vendors.
Stage 3
Reframe the question
Say it like this
"This isn't really 'can you trust a vendor.' It's 'what happens to a pharmacist's own checking habit the day the vendor changes something and nobody tells him.' That's where the actual risk lives."
Why this works
This is where a strong answer separates from a generic "don't depend on vendors" talking point.
Stage 4
Give the one decision
Say it like this
"The risk is that your differentiator's behavior can change on a day you don't choose. The vendor owns the model version, not you. When they push an update, the tone, the caveats, even what counts as a serious flag can shift with zero announcement, and your users recalibrate trust before your team even knows something moved."
Why this works
This is the direct answer, stated as the actual mechanism of the risk, not just the word "risk" repeated with more syllables.
Stage 5
Prove it with the compressed evidence
Say it like this
"Deacon used to cross-check about a third of InteractIQ's flags against the paper monograph himself. Over five months that fell to zero, because the tool had been right every time. Then the vendor pushed an update nobody at Talwick tested first, and a rare combination came through with the dosage caveat quietly missing. Deacon caught it because the patient happened to mention a supplement out loud, not because the screen told him to look twice."
Why this works
This is where the story lives, compressed to the moment the risk stopped being theoretical.
Stage 6
Name the AI-specific reasoning and the trade-off
Say it like this
"The honest reason this is specifically an AI risk, not a normal vendor risk, is that a model's outputs aren't a fixed spec, they're a live behavior that shifts every time the vendor retrains or swaps a version, and there's no contract clause that freezes a model's judgment the way one freezes an API's response format. We accepted faster time-to-market building on someone else's model, in exchange for never fully controlling the day our own differentiator's behavior changes."
Why this works
Names the load-bearing AI-specific judgment, model behavior isn't a frozen spec, and the trade-off actually accepted, speed for control.
Stage 7
Say what wouldn't apply, then close
Say it like this
"This risk doesn't apply the same way to InteractIQ's plain lookup piece, the part that just displays the drug's approved dosage range from a static reference table. That never touches the vendor's model, so it can't drift on someone else's release schedule. For the explainer piece, pin the version, log every change, and test upgrades before they reach anyone's shift."
Why this works
Closes with real judgment about where the risk stops, and restates the decision in one breath.

Let's learn

What happens the first time a product's actual behavior changes, quietly, on a day its own maker didn't choose?

InteractIQ is a feature inside Talwick Health's pharmacy software. It doesn't just flag two drugs that interact, it explains why, in plain language, and says how serious the flag is. Talwick built the whole thing on top of one outside vendor's language model.

Hand sketched icon list titled F L I P S, the five letters. Five rows. F, find the person, whose morning is this. L, locate the habit, what he stopped doing. I, identify the flip, the verb that snaps. P, pinpoint the old decision. S, show the replay.
The five questions that find the actual risk, in order.

In InteractIQ's first months, pharmacists using it still cross-checked a meaningful share of its flags against the paper monograph book kept behind the counter, the way they always had. The tool was reliable, and reliable for long enough that checking it started to feel like a formality rather than a real safeguard.

Share of InteractIQ's flagged interactions Deacon cross-checked against the paper monograph, by month
40% 20% 0 near miss, month 5 Month 1: 34% Month 3: 9%
Share of flags cross-checked by handReached zero
The habit didn't stop in one day. It thinned for five months, and the vendor's silent update landed on exactly the month it finally hit zero.

Here's the turn: the actual danger was never that InteractIQ would eventually get something wrong. Models get things wrong sometimes, that's expected. The danger was that Deacon had no way to know the vendor's model had changed at all, so the change in his own checking habit and the change in the tool's real behavior arrived from two completely different clocks, one that Talwick set and one that Talwick didn't.

Deacon didn't stop checking because the tool got worse. He stopped because it had been right for so long that checking felt like an insult to it.
Hand sketched comparison titled Small move, big snap. Left panel, a gauge icon labeled vendor version, stable, caption quietly patched every few months, nothing felt different. Right panel, a scale icon labeled behavior snaps, caption one release, tone and caveats changed overnight.
For months, nothing about the update schedule was visible from the pharmacy counter. Then one release changed what actually came out the other end.
Knowledge spark: why can't you just freeze a vendor's model in place forever? You can, and pinning a specific version is exactly the fix here. The risk isn't that pinning is impossible, it's that many teams build directly against a vendor's default or "latest" endpoint because it's the easiest way to ship, without realizing that choice quietly hands the vendor a say in when your product's behavior changes.
Hand sketched labeled parts diagram titled What changed the week the vendor pushed a silent update. A document icon at the center labeled InteractIQ output, with four labeled callouts around it: Still cites the right guideline. Severity wording softened. Dropped the dosage caveat line. No version note anywhere.
Three of the four things InteractIQ produced still looked completely normal. The fourth one was the line that actually mattered.

At its worst, this cost is a genuine patient safety event that reaches someone before it reaches a dashboard, because nobody at Talwick was watching for a change they never scheduled.

The choice I would take back Building InteractIQ against the vendor's live, unpinned endpoint instead of a fixed model version. That made sense when speed to launch mattered most and the vendor's updates were assumed to be strict improvements. It stopped making sense the moment "improvement," by the vendor's own definition, turned out to mean quieter caveats.

What I would leave alone: InteractIQ's plain dosage lookup, the part that just displays a drug's approved range from a static reference table, never touches the vendor's model at all. It can't drift on someone else's release schedule, so it doesn't need pinning or a changelog.

The lesson: a vendor's model isn't a fixed spec you bought once. It's a live behavior on someone else's calendar, and the moment you build your actual differentiator on it without pinning the version, their release notes become your product roadmap, whether you read them or not.

Now here is the same thing as a story

The short version above is what you'd say out loud in the room. Read this one for what it actually felt like the week the checking finally stopped.

For five months, Deacon Whitlock trusted a screen a little more each week. Then, on one Tuesday shift, he almost didn't catch the one time it mattered.

Deacon had filled prescriptions at Millrose Pharmacy for six years before InteractIQ ever arrived, and he was good at the part of the job that never shows up on a dashboard: he could glance at two drug names and feel, before he even reached for the monograph book, which combinations were worth a second look.

When InteractIQ launched, it gave him something new, not just a flag, but a plain explanation of why. For the first few months, he still opened the paper monograph for about a third of the flags, just to see if the explanation held up. It always did.

Hand sketched flow diagram titled Deacon's year with InteractIQ, with the near miss emphasized. Five steps left to right: Competent alone. The good months. Habit thinning. The near miss. The replay.
Nothing about the middle three steps looked dramatic from the outside. That was exactly the problem.

So he checked less. A third became a fifth, then barely one in ten. By month five, he wasn't opening the monograph book at all anymore. Nobody told him to stop. The tool simply kept being right, week after week, until checking it felt less like diligence and more like doubting a colleague who'd never once let him down.

Then came a Tuesday. A patient mentioned, almost in passing, that she'd started taking a common herbal supplement for sleep. InteractIQ's flag for her new prescription came back with a note about a mild, manageable interaction, phrased more gently than Deacon remembered it phrasing things like this before. Something about the wording made him pause. He pulled the monograph book for the first time in weeks.

He didn't catch it because the screen told him to look again. He caught it because the screen sounded a little too sure of itself.

The combination was more serious than "mild and manageable." Talwick's own retroactive eval, built only after this near miss, found that on a specific class of rare supplement interactions, the flag's severity language had softened noticeably after a vendor update three weeks earlier, an update nobody at Talwick had tested first, because nobody at Talwick had been told it was coming.

Hand sketched metaphor scene titled Switch, not dial. Left, a gauge icon labeled graduated trust, caption what Talwick assumed Deacon had. Right, a scale icon labeled two positions, caption trust it, or stop trusting it.
Talwick had designed InteractIQ as if trust were a dial pharmacists would adjust a little at a time. It was never a dial. It had two settings, and Deacon had quietly landed on the wrong one.

Building InteractIQ against the vendor's live endpoint wasn't an unreasonable call back when the feature first shipped. Speed to market mattered, and the vendor's updates had, so far, only ever seemed to make things better. It stopped being reasonable the moment "better," in the vendor's own release notes, quietly included "warmer, gentler phrasing on caution language," a change nobody at Talwick had asked for and nobody had tested against a real eval set first.

Here's the replay: with the model version pinned and a changelog in place, that same vendor update would have shown up first against Talwick's own eval set, a set built specifically around rare, high-severity combinations. The severity-language regression would have failed that test in about forty minutes, on a Tuesday afternoon in an engineer's queue, months before it ever reached Deacon's shift.

One version of this story ends with a pharmacist's gut feeling catching what a silent update almost let through. The other ends with an eval suite catching the exact same regression before it ever reaches a real patient, on a clock nobody has to feel lucky about.

What I'd tell myself, watching Deacon reach for that monograph book out of nothing but instinct: we built our actual differentiator on a calendar we didn't control, and called it done the day it shipped, instead of the day we could prove it would still be the same tool next month.

The five steps, on a differentiator that changed without anyone deciding it shouldNot a script for distrusting every vendor model. FLIPS is what shows exactly where an unpinned dependency turns into a real patient-facing risk.

F
Find the person. Whose morning is this?
Deacon Whitlock, six years at Millrose Pharmacy, good enough to feel which drug combinations deserved a second look before he ever reached for the book.
A specific pharmacist, not "clinical staff," is what makes the later flip land.
L
Locate the habit. What did he stop doing because it worked?
Cross-checking InteractIQ's flags against the paper monograph, from about a third of flags down to none, over five months of the tool simply being right.
The habit thinning slowly, not vanishing all at once, is what makes the later snap believable.
I
Identify the flip. What verb snaps, with no middle setting?
Deacon went from spot-checking flags to trusting the flag list completely. There was no in-between state where he checked "a little more carefully."
This is the hard step. The flip here is over-trust, and it's the one people forget an improvement can also cause.
P
Pinpoint the old decision. Which choice only made sense before?
Building InteractIQ against the vendor's live, unpinned endpoint, with no changelog, because the vendor's updates had only ever looked like improvements so far.
Naming a real product decision, not "the vendor changed something," keeps this an answer Talwick could have controlled.
S
Show the replay. Same update, fixed design.
With a pinned version and an eval set, the same severity-language regression fails a forty-minute internal test, months before it reaches a real shift.
A replay that ends in a real, countable moment, not "the team is more careful now."

The recap, one line per letter: find the person is Deacon, six years in; locate the habit is the monograph checks thinning to zero; identify the flip is over-trust, checking sometimes to checking never, right as the tool seemed most reliable; pinpoint the old decision is the unpinned vendor endpoint with no changelog; and show the replay is the same regression caught in a forty-minute eval instead of a live near miss.

And if you want to be sure it really works, try it somewhere elseSame five letters, a different flip family, a legal-translation reviewer instead of a pharmacy tool. This time nobody trusted too much, they just built a private workaround around the gap.

Soraya Lindqvist is the AI PM at Verity Language Partners, whose actual differentiator is a tool that reviews machine-translated legal contracts and flags clauses that likely lost precision in translation, built on a single third-party model. Mapped onto FLIPS: find the person is Piotr, a senior legal translator with a personal stake in never letting a bad clause through. Locate the habit is different from Deacon's, Piotr never stopped checking, he checks everything, because the tool has no memory of which clauses it flagged last time. Identify the flip here is the workaround flip, not over-trust: Piotr starts keeping his own private spreadsheet of every correction he's made, because the tool forgets everything the moment the vendor updates it, and it's the only way he can tell whether the same mistake keeps recurring. Pinpoint the old decision is the same root cause as Deacon's story wearing a different shape, Verity's differentiator holds no run history and no memory of past corrections, an absent-state decision that made sense when the tool first shipped and nobody expected reviewers to compare version to version by hand. Show the replay: once Verity added a real correction log the tool could reference, Piotr's private spreadsheet became unnecessary, and a genuine, visible pattern of recurring mistakes surfaced for the first time.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "a differentiator built on a vendor's live model can change behavior on their calendar, not yours, pin the version and log it," and stop.
Cost: no time or budget to build a full eval suite before launch. Say so honestly, and commit to at minimum pinning the model version and manually spot-checking any vendor-flagged upgrade before accepting it.
The model got better, for real: say the vendor's update genuinely improved accuracy overall. The risk still runs, a real improvement on average can still hide a regression on the rare, high-stakes case nobody's average accounts for.

Where people run it wrong.
They treat "the vendor is reliable" as a permanent fact instead of a fact about last quarter.
They watch for the vendor update that reads as a bug, and miss the one that reads as an improvement.
They build the eval set only after a near miss, instead of before the first real update ships.

How to use it live. The moment an interviewer asks about the risk of depending on a vendor's model, ask yourself first: who is the person whose trust in this tool would quietly change the day its behavior shifted, and would they even know it had? That question alone buys real thinking time, and it's usually exactly where the honest risk lives.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Over-trust: checks sometimes, then stops checking at all, especially right after the tool starts seeming more reliable, not less. Fires on good news, which is why people miss it.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Deacon Whitlock, a six-year pharmacist at Millrose Pharmacy. Zanele Mbatha is the AI PM at Talwick Health who owns InteractIQ, Talwick's differentiator built on a vendor's model.
3 · THE HABIT
What did Deacon stop doing because InteractIQ worked?
Tap to flip
ANSWER
Cross-checking flagged interactions against the paper monograph book, from about a third of flags down to zero over five months.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Checking every flag by hand, versus trusting the flag list completely. There was no middle setting where Deacon checked "a bit more carefully."
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Building InteractIQ against the vendor's live, unpinned endpoint with no changelog. Reasonable when speed mattered and updates seemed like pure improvements. Wrong once "improvement" quietly meant softer caution language.
6 · THE NUMBER
Fill in the blank: Deacon's monograph cross-check rate fell from ___ percent in month one to ___ percent by month five, the same month as the near miss.
Tap to flip
ANSWER
34 percent, then 0 percent. The vendor's silent update landed in the exact month the habit had fully thinned away.
7 · THE REPLAY
Same vendor update, new design. What changes?
Tap to flip
ANSWER
With the model version pinned and an eval set watching for exactly this, the severity-language regression fails a forty-minute internal test, months before it reaches a real pharmacy shift.
8 · CROSS PRODUCT TRANSFER
Section 4 runs this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Verity Language Partners' contract reviewer. The workaround flip, a translator builds a private spreadsheet of corrections because the tool holds no memory across the vendor's updates.

Check yourself Score: 0 / 0

Multiple choice
1. What actually caused Deacon's near miss?
  • A. InteractIQ's underlying model got measurably less accurate overall.
  • B. A silent vendor update softened severity language, and Deacon had already stopped cross-checking flags by hand.
  • C. Deacon was undertrained on how to read InteractIQ's flags.
  • D. Millrose Pharmacy had disabled the interaction-checking feature entirely.
Show hint
Look at the line chart and the labeled-parts diagram together.
Show answer
B. The model wasn't broadly worse, one specific line softened after an untested update, and Deacon's own checking habit had already thinned to zero by that same month.
True or false
2. True or false: the vendor's update that caused the near miss was described in its own release notes as a bug fix.
  • True
  • False
Show hint
Look at what "warmer, gentler phrasing" was called in the story.
Show answer
False. The vendor's own notes framed it as an improvement, warmer and gentler phrasing, which is exactly why nobody at Talwick flagged it as risky before it shipped.
Fill in the blank
3. Fill in the blank: Deacon's rate of cross-checking flagged interactions by hand fell from 34 percent in month one to ___ percent by month five.
Show hint
Look at the line chart in "Let's learn."
Show answer
0 percent. Zero, the same month the vendor's silent update reached production and nearly went unnoticed.
Short answer, where it wouldn't matter
4. Name a part of InteractIQ where this exact risk would NOT apply, and say why not.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The plain dosage lookup, which just displays a drug's approved range from a static reference table. It never touches the vendor's model, so it can't drift on someone else's release schedule.
Short answer, apply it yourself
5. Think of a tool you use that's built on someone else's model. What habit of yours would change if its behavior shifted overnight, and would you even notice it had?
Show hint
Think about a habit that formed because the tool has been consistently right, not consistently wrong.
Show answer
Model answer: A developer who stopped reading an AI coding assistant's suggested code line by line, after months of it being reliable, realized they'd have no way to notice a silent quality drop until a bug reached production.
Short answer, work the number
6. If Deacon's cross-check rate had still been at 20 percent instead of 0 percent when the update landed, would the near miss still have happened?
Show hint
Think about what a one-in-five chance of catching a specific flag actually means for one specific patient.
Show answer
Model answer: Maybe not this exact case, but the risk itself doesn't go away at 20 percent, it just becomes a smaller chance of catching any one specific error, not a system that reliably catches this class of error at all.
Before you close the answer
Why this works
Tests whether you understand that a vendor-model dependency isn't a one-time integration risk, it's an ongoing behavioral risk that changes on the vendor's schedule, and whether you can trace exactly how that reaches a real person's trust.
Follow-up traps
"Isn't pinning a model version just going to make you fall behind on real improvements?" Response: no, pinning doesn't mean never upgrading, it means upgrading on your own tested schedule instead of the vendor's untested one, which is strictly safer at the same eventual destination.

"Couldn't you catch this with better user training instead?" Response: training doesn't fix a silent change nobody was told about, Deacon was well trained, the actual gap was that nobody at Talwick knew the model had changed at all.
If pressed
The vendor's contract technically did allow requesting a pinned, dedicated model version, at a higher tier Talwick hadn't purchased. The strategic risk wasn't just technical, it was a cost decision made before anyone had priced out what an unpinned dependency could actually cost in a near miss.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more