CaseAdvancedAI Opportunity & Model Strategy / Model selection from a PM lens / #14

How would you decide between an open-weights model and a proprietary API for a healthcare product?

PICKthe phone in Marisol's hand was never the risk, the tier her employer clicked was

Marisol Pena dictates a visit note into her phone after every home health stop for Solace Home Health, and NoteBridge turns it into a structured clinical note before she's back in the car. The question underneath the question: which model gets to hear her dictate the actual words a patient's daughter said about her mother's confusion this morning?

The direct answer
Put the step that touches raw patient narrative, names, addresses, anything identifiable, on a self-hosted open-weights model, inside your own network, so it never leaves. Keep the proprietary API only for whatever runs after identifiers are already stripped out. The split isn't about which model writes better sentences. It's about which mistake you can actually take back.
Do this, in order
  1. Send raw dictation and identifiers only to a self-hosted, open-weights model inside your own network.Why: data that never leaves can't be logged, retained, or exposed by someone else's tier mistake.
  2. Check the actual contract tier before trusting any vendor's "HIPAA-ready" claim.Why: a standard API tier and an enterprise BAA tier can look identical in a demo and mean completely different things for your data.
  3. Reserve the proprietary API for steps that only ever see de-identified text.Why: once identifiers are stripped, the asymmetry that favors self-hosting disappears.
  4. Set a kill criterion: revisit if a vendor signs an audited, zero-retention BAA at a real price.Why: the pick is a decision made with today's evidence, not a permanent rule.
  5. Give every note a visible flag showing which model touched it and why.Why: an auditor's first question is always "show me the data-flow map," and a flag answers it instantly.

How to answer this, stage by stage

This isn't a question about which model writes better prose. It's a question about which mistake you're willing to eat.

Stage 1
Scope it to one real system
Say it like this
"Let's ground this in one product. NoteBridge turns a home health nurse's spoken visit note into a structured clinical record. I'll answer for the model choice behind that specific step."
Why this works
Keeps "open-weights versus proprietary" from turning into a generic architecture debate with nobody's data actually at stake.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as PICK. Position, my pick, stated first. Impact, who feels each kind of error. Cost asymmetry, which one is cheap and which one is hidden and expensive. Kill criteria, what evidence would change my mind."
Why this works
Signals you can commit to a pick, not just list pros and cons for both sides.
Stage 3
Position: state the pick before the reasoning
Say it like this
"My pick: self-hosted, open-weights, for any step that touches raw patient narrative or identifiers. Proprietary API only for whatever runs after identifiers are stripped."
Why this works
This is the direct answer, stated plainly, before any hedge about "it depends."
Stage 4
Reframe: it isn't quality versus quality, it's which error you can undo
Say it like this
"Both models can write a clean clinical note. The real question is what happens on the day one of them makes a mistake. A clunky sentence gets fixed in thirty seconds. A patient's name logged on a non-BAA endpoint doesn't come back."
Why this works
Separates a candidate who's thought about model quality from one who's thought about what a mistake actually costs.
Stage 5
Prove it with the compressed failure
Say it like this
"At Solace, NoteBridge launched on a proprietary vendor's standard API tier, not their enterprise BAA tier, because the demo looked identical and it shipped four months faster. A compliance audit found raw dictation, patient names included, had been flowing through a non-BAA-covered endpoint the whole time."
Why this works
Compresses the whole cost asymmetry into the one failure that a clean quality comparison never would have caught.
Stage 6
Name the kill criteria and the trade-off
Say it like this
"I'd revisit the whole pick the moment a proprietary vendor offers a real, audited, zero-retention BAA at a reasonable price. Until then, I'm accepting slightly less fluent phrasing from a self-hosted model, in exchange for never having to explain to a regulator where a patient's words actually went."
Why this works
Names the load-bearing AI-specific judgment and states the trade-off plainly, instead of pretending both paths are free.
Stage 7
Close on the one line
Say it like this
"So: raw and identifiable stays open-weights, in-house. De-identified can go to the proprietary API. The split follows the data, not the demo."
Why this works
Restates the direct answer in one breath, exactly what a live interview rewards.

Let's learn

NoteBridge listens to a home health nurse's spoken visit note and turns it into a structured clinical record a physician can review.

Before NoteBridge, Marisol typed every note by hand in her car after each visit, about 30 minutes a stop.

With NoteBridge, she reviews and edits a drafted note instead of writing one from nothing, about 8 minutes a stop, across roughly 9 visits a day.

Here's the turn: the risk in this system was never how well the model wrote a sentence. It was which tier of which vendor's contract Marisol's dictation was quietly routed through, a decision made once, in an integration meeting, that nobody revisited for four months.

Hand sketched comparison diagram titled The asymmetry, drawn. Left panel, a document icon labeled Cheap and visible, caption a clunky sentence, a 30 second edit. Right panel, a scale icon in a different color labeled Hidden and expensive, caption PHI logged on a non-BAA endpoint, a breach review.
One error costs Marisol thirty seconds. The other costs Solace a compliance review and a very hard conversation with patients.

At its worst, choosing a proprietary API by demo quality alone, without checking the actual contract tier, exposes patient narrative on an endpoint nobody at the company meant to use for real medical dictation.

The choice I would take back Solace's integration team launched NoteBridge on the vendor's standard API tier, not the enterprise BAA tier, because the sales demo behaved identically and switching tiers would have delayed launch by six weeks. That made sense when nobody expected "just summarization" to matter for a contract's fine print. It stopped making sense the moment that fine print turned out to govern where a patient's actual words went.
Hand sketched icon list titled Signs a task needs to stay open weights, on prem. Four rows: a document icon, raw audio or transcript with a patient's name attached. A question mark box icon, a field a regulator would ask for a data flow map on. A box icon, a vendor's standard tier with no signed BAA. A person icon, anything a patient could ask you to delete on request.
Four quick checks that decide the split faster than any architecture debate would.

What I would leave alone: the downstream grammar and formatting pass, which only ever runs after identifiers are already stripped from the note, can stay on the proprietary API. The asymmetry that favors self-hosting doesn't apply once there's nothing identifiable left to expose.

The lesson: a demo tells you how a model writes. It tells you nothing about which contract tier your integration actually landed on. Go check the tier before you ever compare the writing.

Hand sketched decision tree titled Does this step touch identified patient data. Root, a new NoteBridge step is proposed. Branches: yes raw dictation or identifiers leads to self hosted open weights our own VPC. No already de identified leads to proprietary API is fine. Unsure mixed fields leads to strip identifiers first then decide.
One question decides the whole architecture. Everything else follows from the answer.

Now here is the same thing as a story

The short version above is what you'd say in a vendor-selection review. Read this one for the week the audit actually landed.

Marisol Pena carried her phone into every home the way she'd carried a clipboard for the six years before it, dictating each visit's story the moment she stepped back to her car.

NoteBridge shipped fast: six weeks from pitch to every field nurse's phone, on the vendor's standard API tier, because the sales team's demo notes looked exactly like the enterprise tier's, word for word.

Hand sketched metaphor scene titled Rented or owned. Left, a box icon labeled Rented, caption a sealed box pay per call data leaves the building. Right, a person icon in a different color labeled Owned, caption weights you host yourself nothing leaves.
Two ways to get the same sentence written. Only one of them tells you exactly where the sentence went.

For four months, nothing looked wrong. Notes came back clean, physicians signed off, nurses saved twenty-two minutes a visit.

Knowledge spark: what's the actual difference between a standard tier and a BAA tier? A standard API tier is built for anyone, and its default terms often allow the vendor to log or retain requests for abuse monitoring or model improvement. A BAA, business associate agreement, tier is a specific, signed, audited contract that legally binds the vendor to HIPAA's rules on that data. Two tiers, same model, same sentence quality, completely different legal reality underneath.

A routine security review, the kind scheduled a year in advance, asked for a data-flow map of every system touching patient information. NoteBridge's map led straight to the vendor's standard tier, the one without a signed BAA.

The model never made a single mistake. The contract Solace clicked into, four months earlier, was the mistake.

The fix took three weeks: stand up an open-weights model inside Solace's own network for the raw dictation step, and keep the proprietary API only for the formatting pass that runs after every identifier is already stripped out.

Hand sketched timeline titled The hybrid rollout. Four milestones: Week 1, audit finds the standard tier gap, this one emphasized. Weeks 2 to 4, stand up open weights model in our VPC. Week 5, raw dictation cut over. Week 6, proprietary kept only for de identified polish.
The fix didn't throw out the proprietary API. It just moved the one step that actually mattered.

When the original integration was pitched, someone said, "their demo output is identical to the enterprise tier, let's ship on the cheaper one and upgrade later if we need to." Reasonable, at the time. Nobody had yet asked what "later" would cost.

Cost per note, standard proprietary tier vs self-hosted open-weights
12¢ 0 Proprietary, standard tier 11¢ Self-hosted, open-weights
Nearly three times the per-note cost buys one thing the cheaper tier never did: proof of where the data went.
Cost per note as visit volume grows, the kill line
24¢ 12¢ 0 proprietary, flat at 4¢ crossover, ~40k notes/mo self-hosted, falling with volume 10k/mo 60k/mo
Solace runs about 35,000 notes a month, close enough to the crossover that self-hosting already pays for itself in cost alone, before counting what it buys in control.

What I'd tell myself, reading that audit finding: the phone in Marisol's hand was never the risk. The tier her employer clicked, four months earlier, in a meeting about launch speed, was.

PICK, held against one contract tierNot a rule against proprietary models. A test for exactly which step earns the caution.

P
Position. The pick, stated first.
Self-hosted, open-weights, for raw dictation and identifiers. Proprietary API only for de-identified steps.
This is the direct answer, and the hardest step: committing before laying out every argument.
I
Impact. Who feels each kind of error?
A clunky sentence costs Marisol thirty seconds at her next visit. A non-BAA-tier exposure costs Solace a compliance review and every patient on that endpoint a disclosure letter.
Naming both sides in real units, not "user experience" versus "risk," makes the asymmetry checkable.
C
Cost asymmetry. Which one is hidden and expensive?
Phrasing errors are visible immediately and cheap to fix. A contract-tier gap is invisible for months and, once tripped, can't be undone, the data already left.
This is the heart of PICK: optimize against the error you can't take back, not the one you notice first.
K
Kill criteria. What would change the pick?
A proprietary vendor offering a genuinely audited, zero-retention BAA at a cost close to the standard tier would be worth revisiting the whole architecture for.
Naming the kill criteria is what separates a real decision from stubbornness.

The recap, one line per letter: position is self-hosted for raw data, proprietary for de-identified data, impact is thirty seconds versus a compliance review, cost asymmetry is a hidden, unrecoverable exposure against a cheap, visible edit, kill criteria is a real audited zero-retention BAA at a fair price.

And if you want to be sure it really works, try it somewhere elseSame four letters, a small newsroom instead of a home health agency. A different asymmetry this time: a source's safety instead of a patient's record.

Petra Voss reports on local government for the Millbrook Courier, and her paper uses a transcription tool on recordings of public meetings, some of which include off-the-record remarks from anonymous sources who assume a reporter's recorder, not a cloud vendor, is listening. Mapped onto PICK: position is running any recording with an unnamed source's voice through a self-hosted open-weights model on the paper's own machine, keeping a proprietary API only for already-public, on-the-record city council footage. Impact: a clunky transcript of public footage costs an editor a few minutes of cleanup; a source's identity surfacing in a vendor's retained logs could cost that person their job or worse. Cost asymmetry: the visible cost is a slower, less polished transcript for sensitive recordings; the hidden cost is a subpoena reaching a vendor's servers instead of stopping at the newsroom's door. Kill criteria: a vendor offering a contractually enforced no-subpoena-compliance-without-notice clause, verified by the paper's own lawyer, not just a marketing page.

Hand sketched labeled parts diagram titled The hybrid architecture. A document icon at the center labeled NoteBridge, with four labeled callouts around it: Raw dictation on prem model, De identified text proprietary polish, Audit log every hop, Patient consent flag.
The same split, in different words, protects a patient's record and a source's identity for the same underlying reason.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "raw and identifiable stays self-hosted, de-identified can go anywhere," and stop.
Cost: self-hosting is genuinely too expensive at your current volume. Say so honestly, and shrink the pick to only the single highest-risk field, not all of it.
The proprietary model got better, for real: if a vendor ships a real, audited, zero-retention healthcare tier, that's exactly the kind of evidence the kill criteria was built to act on, not ignore out of habit.

Where people run it wrong.
They pick a vendor based on a demo's writing quality, without ever reading which contract tier they're actually integrating against.
They treat "open-weights" as automatically safer, when a badly-configured self-hosted model with no access controls can leak data just as easily.
They apply the same caution to every field in a record, instead of finding the one that actually carries identifiers.

How to use it live. When an interviewer asks open-weights versus proprietary, ask yourself: which specific field in this record, if exposed, can't be taken back? Build the split around that field, not around the model's reputation.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits an open-weights versus proprietary decision?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. It forces a committed pick instead of a list of pros and cons for both sides.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Marisol Pena, a home health nurse at Solace Home Health, who dictates a visit note into her phone after every stop.
3 · THE HABIT
What did Marisol stop doing once NoteBridge worked?
Tap to flip
ANSWER
Writing every visit note from nothing in her car. She started editing a drafted note instead, cutting the time from 30 minutes to about 8.
4 · THE ASYMMETRY
What's the cheap error, and what's the hidden one?
Tap to flip
ANSWER
Cheap: a clunky sentence, fixed in 30 seconds. Hidden: raw patient dictation flowing through a non-BAA-covered vendor tier for four months.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Launching NoteBridge on the vendor's standard API tier instead of the enterprise BAA tier, because the demo output looked identical and switching would have delayed launch.
6 · THE NUMBER
Fill in the blank: self-hosting costs about ___ cents a note versus ___ cents for the proprietary standard tier.
Tap to flip
ANSWER
11 cents self-hosted, versus 4 cents proprietary, nearly three times the cost for data that never leaves the network.
7 · THE REPLAY
Same audit, hybrid architecture already in place. What changes?
Tap to flip
ANSWER
The data-flow map shows raw dictation never left Solace's own network, and the audit closes in an afternoon instead of triggering a four-month retroactive contract review.
8 · CROSS PRODUCT TRANSFER
Section 4 runs this again for a different product. Which one, and what's the equivalent of a patient record?
Tap to flip
ANSWER
The Millbrook Courier's meeting-transcription tool. The equivalent of a patient record is an anonymous source's identity inside an off-the-record recording.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: the compliance audit found raw dictation had been flowing through a non-BAA endpoint for about ___ months.
Show hint
Look at the story, right after "for four months, nothing looked wrong."
Show answer
Four months. Long enough that nobody thought to double-check a decision made once, in a launch-speed meeting.
Multiple choice
2. Why does the proprietary API stay in the architecture at all, instead of being replaced entirely?
  • A. It's required by Solace's existing vendor contract.
  • B. Once identifiers are stripped, the cost asymmetry that favors self-hosting no longer applies.
  • C. Open-weights models can't produce clean formatting.
  • D. It's cheaper in every case, regardless of what data it touches.
Show hint
Look at "what I would leave alone."
Show answer
B. The asymmetry between a cheap, visible error and a hidden, expensive one only exists while identifiers are present.
True or false
3. True or false: this answer treats "open-weights" as automatically safer than "proprietary," regardless of how either is configured.
  • True
  • False
Show hint
Look at "where people run it wrong."
Show answer
False. A badly-configured self-hosted model with weak access controls can leak data just as easily. The split is about which field carries identifiers, not the vendor's label.
Short answer, where it wouldn't matter
4. Name a step in NoteBridge where this caution genuinely doesn't apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The grammar and formatting pass that runs after identifiers are already stripped from the note. There's nothing left to expose at that point.
Short answer, apply it yourself
5. Think of an app on your phone that sends something personal to a server. What's the one field in it you'd want to keep on-device if you could?
Show hint
Think about what part of what you type or say is the part you'd least want stored somewhere you don't control.
Show answer
Model answer: A journaling app's actual entry text, versus its metadata like word count or mood tag, which carries far less risk if it ever leaked.
Short answer, work the number
6. If Solace's visit volume doubled to 70,000 notes a month, would self-hosting still make financial sense on its own, without counting the compliance benefit?
Show hint
Look at the crossover point on the cost-per-note line chart.
Show answer
Model answer: Yes, even more so. The crossover sits around 40,000 notes a month, and cost per note keeps falling with volume, so 70,000 notes puts self-hosting further ahead on cost alone.
Before you close the answer
Why this works
Tests whether you evaluate a vendor by its demo, or by its actual contract terms and data-flow reality. Most candidates only compare output quality.
Follow-up traps
"Couldn't you just buy the enterprise BAA tier from the same proprietary vendor instead of self-hosting?" Response: yes, and that's a real option, but it still leaves the model provider as a party who holds the data, which some governance requirements and some patients' own comfort won't accept, regardless of the contract.

"Isn't self-hosting just moving the risk to your own security team instead of removing it?" Response: it does shift the burden, which is exactly why it needs real access controls and an audit log of its own, not an assumption that "on-prem" alone equals "safe."
If pressed
The self-hosted model runs on hardware inside Solace's own compliance boundary, with its GPU costs amortized against total visit volume, which is exactly why the cost per note falls as volume grows, unlike the flat, linear pricing of the proprietary API.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more