ConceptAdvancedResponsible AI & Advanced Practice / Compliance and legal partnership / #15
How do sector regulations change the AI PM's job in healthcare or finance?
LEAD the product is Meridian Health's AI intake triage assistant, used at the front desk of a primary care clinic
What happens when "just describe it in your own words" turns out to have a catch nobody built on purpose? Meridian Health's triage assistant reads a patient's intake note and scores how urgently they need to be seen. Zanele Dube is the nurse practitioner who leads the clinical product team behind it.
The direct answer
In a regulated sector, your north star metric can't be the model's own accuracy, because accuracy on a shrinking, sanitized slice of real cases tells you nothing. Watch the leading indicator instead: how often intake notes are getting shorter and cleaner over time, without a real drop in how complicated patients' actual symptoms are. That drift shows up weeks before a missed-escalation ever does, and it's the one number a healthcare or finance AI PM has to own that a consumer-app PM never needs to.
Do this, in order
Track clean-note rate against real case complexity, not against volume.Why: a rising clean-note rate with no matching drop in real complexity means people are learning to feed the model, not that patients got simpler.
Stop letting confidence rise from brevity alone.Why: a short, tidy note isn't the same thing as a genuinely clear case, and treating them the same is exactly what gets gamed.
Audit a sample of "clean" notes against the real encounter, on a schedule, not just after an incident.Why: by the time a missed-escalation happens, the drift has usually been building for weeks already.
Tell staff plainly what the model actually rewards.Why: an onboarding promise of "just talk normally" with no follow-up leaves people to discover the model's real incentives on their own, informally.
Leave simple, genuinely clear cases alone.Why: not every short note is a red flag, and treating all of them as suspicious would slow down the easy cases for no reason.
How to answer this, stage by stage
Seven stages. The question sounds like a policy essay. The job is to make it one real number, fast.
Stage 1
Scope it to one real product
Say it like this
"I'll answer this for a triage assistant at a primary care clinic, one that scores how urgently a patient needs to be seen from their intake note."
Why this works
Grounds "how does regulation change the job" in one measurable product instead of a policy debate.
Stage 2
Say your structure out loud
Say it like this
"I'll use LEAD. Link to the real outcome, find the early signal, name how it gets abused, and say what you'd do at each threshold."
Why this works
Signals a metric-first answer instead of a list of regulations to name-check.
Stage 3
Reframe the question
Say it like this
"The real change isn't a longer list of rules to follow. It's that the metric you'd normally trust, the model's own accuracy, stops being trustworthy, because the input feeding it quietly changes shape under pressure."
Why this works
Moves past "there are more rules in healthcare" into the actual mechanism being tested.
Stage 4
Give the one decision
Say it like this
"Watch clean-note rate against real case complexity. If notes are getting shorter and tidier without patients actually getting simpler, that's the leading sign something's drifting, weeks before a bad outcome shows up."
Why this works
Matches the direct answer exactly. This is the one sentence an interviewer should walk away remembering.
Stage 5
Prove it with a real case
Say it like this
"Here's what it looks like. A patient describes chest pain that comes and goes, with a strange feeling down her arm. The front desk types it up as 'chest pain,' the form's easiest bucket, and the model scores it moderate. She waits six hours instead of getting flagged same-day."
Why this works
A compressed, real case, four sentences, showing exactly how the drift causes harm.
Stage 6
Say what you'd leave alone
Say it like this
"A short note for a genuinely simple case, a prescription refill, a sprained ankle, is completely fine. I'm not asking anyone to write longer notes for the sake of it."
Why this works
Shows judgment instead of treating every short note as automatically suspicious.
Stage 7
Close on one line
Say it like this
"In a regulated sector, the AI PM's real job is owning the metric that would have looked fine right up until the morning it wasn't, not the metric that only tells you after."
Why this works
Restates the decision in one breath, closing on the point instead of trailing off into policy talk.
Let's learn
Say a clinic gives its front desk a tool that reads a patient's intake note and scores how urgently they need care, so the sickest patients get seen first instead of by arrival order.
At launch, the clinic told every staff member the same reassuring line: describe the patient's symptoms in your own words, no special format needed, the tool understands normal language. For months, that promise held up fine. Notes stayed messy, real, and full of the odd detail a patient actually says.
Knowledge spark: what's a leading indicator?
A number that moves before the thing you actually care about does. In this story, the clinic's real concern is a patient who needed urgent care and didn't get flagged for it. The leading indicator is something that drifts weeks earlier, quietly, before that ever happens.
Then, without anyone deciding it, front-desk staff started noticing that short, tidy notes seemed to get faster, more confident scores. Nobody was told to write that way. It was simply the cheapest adjustment available, and people take the cheap adjustment.
Clean-note rate, five months, before the missed-escalation rate ever moved
Nothing about real patients got simpler over these five months. The notes describing them just got shorter.
At its worst: a patient's intermittent chest pain and arm tingling gets typed up as "chest pain" to fit the fastest form field, scores as a moderate case, and she waits six hours before an on-call physician happens to flag it as worth an urgent look.
The decision I would take back
We told every new staff member, on day one, that the tool understands plain, everyday language, no special format required. That was true and helpful at launch, when a short note and a genuinely simple case were still the same thing. It stopped being true once staff informally learned that shorter, tidier notes scored faster and more confidently, with nothing in the tool ever correcting that impression.
What I would leave alone: a short note for a real prescription refill or a sprained ankle is exactly as good as it looks. The fix isn't to make every note longer. It's to stop treating brevity itself as a sign of a clear case.
The model didn't get worse at reading symptoms. The symptoms it was being shown got quietly edited before they ever reached it.
The lesson: a friendly onboarding promise with no follow-up leaves people to discover, on their own, exactly what the model rewards, and they will learn to feed it that, whether or not it's the honest version of the patient in front of them.
Now here is the same thing as a story
The short version above is what you'd say defending this metric to a clinical director. Read this one for how Zanele actually found the drift.
The triage desk at Meridian's clinic gets a new intake note every few minutes during the morning rush, and for a long while, those notes stayed exactly as messy as real patients actually are.
There was no single bad morning that started the change. It built up slowly, the way a habit does when nobody's watching for it. A staff member would type a full, rambling symptom description, notice the score come back a little slower and a little less confident, and quietly learn that a shorter, cleaner version scored faster next time. Nobody taught this. Everybody discovered it separately, at their own pace, over months.
Five steps, and the second one quietly stopped reflecting the first one honestly.
Zanele found it during her own quarterly model-health review, the kind of routine check that usually turns up nothing. This time, clean-note rate had nearly doubled over five months, with no matching change in how complicated patients' actual visits were.
The third branch is the one nobody designed. It just turned out to be the cheapest path available.
She pulled the note behind that six-hour wait and compared it to the actual recorded encounter. The patient had described pain that came and went, and a strange feeling down her arm. The note on file said two words: chest pain.
One of these four parts had been quietly getting smaller for months, and nobody had a number on it until now.
She thought about the two clocks running underneath the whole system.
One clock rings after the damage is done. The other one had been ringing for months.
With the redesign, the model no longer lets confidence rise from note length or tidiness alone. It calibrates against real complexity signals, multiple symptoms, hedging words like "comes and goes," and flags an ambiguous presentation for a higher tier regardless of how short the note reads. And staff now see, plainly, that a tidy note doesn't score faster just for being tidy.
The old design rewarded a clean note. The new one asks whether the case behind it was actually clean.
I told every new hire the tool understood plain language because it was true and it made onboarding easy. It took my own quiet, routine number-check, not a single dramatic failure, to see that a true promise with no upkeep quietly turns into a false one as people learn what a system actually rewards.
LEAD, in one screenNot a dashboard metric. LEAD is what forces you to find the number that moves before the one you actually care about.
L
Link. The real outcome.
A patient who genuinely needs urgent care actually gets flagged for it, not the model's raw accuracy on whatever notes happen to reach it.
Not the model's score. The thing the business, and the patient, actually needs.
E
Early signal. The hardest step.
Clean-note rate against real case complexity, which drifted for five months before missed-escalation rate ever moved.
Accuracy on a shrinking, sanitized slice of cases tells you nothing. This number tells you everything, weeks earlier.
A
Abuse. How it gets gamed.
Staff could hit a great clean-note number by templating notes without really listening, satisfying the metric while patients get worse care.
Every metric has a cheap way to be hit without doing the real work behind it.
D
Decision. What you'd actually do.
If clean-note rate rises without a matching drop in real complexity, audit a sample of "clean" notes against the actual encounter that week, not after an incident.
A metric nobody acts on is a decoration. This one has a stated trigger and a stated response.
The form can look perfectly filled in while the actual person behind it walks right past the thing that mattered.
Missed-escalation rate: before the drift, at discovery, and after the fix
The lagging number went up and came back down. The clean-note rate saw the whole thing coming, five months in advance.
The recap, one line per letter: link is a patient who needs urgent care actually getting flagged for it, early signal is clean-note rate drifting five months ahead of the outcome, abuse is staff templating notes to hit the number without truly listening, and decision is auditing a sample the moment the drift shows up, not after an incident forces the question.
And if you want to be sure it really works, try it somewhere elseSame four letters, a trading desk instead of a clinic. This time the input isn't a symptom, it's a trade, and the leading indicator lives in a review queue instead of a note.
Ridgeline Trading runs an AI copilot that flags potentially manipulative trades for a compliance team to review. Henrik Solberg manages that surveillance model.
Mapped onto LEAD: link is a manipulative trade actually getting caught and escalated to a regulator, not the model's raw flag-accuracy on whatever gets reviewed. Early signal is same-day review completion rate on flagged trades, which fell from 95 percent to 30 percent over three months as flag volume grew faster than headcount, long before any regulatory inquiry ever landed. Abuse is analysts clearing a backlog by batch-approving clusters of similar-looking flags, hitting a healthy-looking queue-length number while barely reading half of them. Decision is that once same-day completion drops below a set floor, new flags above a risk threshold get a second reviewer automatically, rather than waiting for a compliance backlog to become a genuine inquiry.
A different desk, a different input, and the same shape of gap: a queue that looked handled without actually being read.
Swap the trigger and it still runs.
Speed: an interviewer caps you at a minute. Say "the leading indicator is clean-note rate against real complexity, not raw accuracy," and stop there.
Cost: there's no budget this quarter for a full complexity-calibration model. Start by simply flagging any note under a certain length for a light human glance, a cheap floor while the real fix gets built.
The model gets better, for real: if the underlying triage model's overall accuracy improves, the pre-editing problem doesn't go away. A more accurate model reading a quietly sanitized note is still reading the wrong input.
Where people run it wrong.
They track the lagging outcome alone and call it "watching quality," when it only tells them after the fact.
They assume a rising "good-looking" metric like clean-note rate is progress, instead of checking what's actually driving it.
They blame the model for getting worse, when the real change was in what the model was ever shown.
How to use it live. When this question comes up, ask yourself what the honest input to the model actually looks like under real time pressure, not in a demo. That's usually where the regulated-sector answer actually lives.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Pre-editing flip: staff learn the tool's failure shape and start sanitizing the input, stripping the messy real detail before the model ever sees it.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Zanele Dube, a nurse practitioner who leads the clinical product team behind Meridian Health's triage assistant.
3 · THE HABIT
What did staff stop doing because it worked?
Tap to flip
ANSWER
They stopped writing full, messy, real symptom descriptions, learning informally that shorter, tidier notes scored faster.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Feeding the model the real, messy symptom description versus a cleaned-up version tailored to score faster. Once staff learned the pattern, there was no drifting back.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Promising every new hire the tool understood plain language, with nothing to correct the impression once staff discovered brevity scored faster.
6 · THE NUMBER
Fill in the blank: clean-note rate rose from 40 percent to ___ percent over five months.
Tap to flip
ANSWER
78 percent. Missed-escalation rate, the lagging outcome, only started climbing well after that.
7 · THE REPLAY
Same chest-pain-and-arm-tingling patient, redesigned model. What changes?
Tap to flip
ANSWER
The hedging language and multiple symptoms flag the case as ambiguous regardless of note length, and she's seen same-day instead of waiting six hours.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the early signal there?
Tap to flip
ANSWER
Ridgeline Trading's surveillance copilot. There, the early signal is same-day review completion rate on flagged trades, falling from 95 to 30 percent.
Check yourself Score: 0 / 0
Short answer, name the reversal
1. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Promising every new hire the tool understood plain language, no format needed. True at launch, when a short note and a genuinely simple case were still the same thing.
Multiple choice
2. Why is clean-note rate a better metric here than the model's raw accuracy?
A. Clean-note rate is easier to calculate.
B. Raw accuracy on a quietly sanitized slice of notes tells you nothing, while clean-note rate reveals the sanitizing itself, weeks early.
C. Accuracy can't be measured in a clinical setting at all.
D. Clean-note rate is required by healthcare regulation.
Show hint
Look at the L and E steps in the LEAD recap.
Show answer
B. LEAD's whole point is finding the number that moves first. Accuracy on a shrinking, cleaned slice of input is measuring the wrong thing entirely.
True or false
3. True or false: this answer recommends requiring longer notes for every patient, regardless of how simple the visit is.
True
False
Show hint
Look at "what I would leave alone."
Show answer
False. A short note for a genuinely simple case, like a prescription refill, is left exactly alone. The fix targets brevity used to game confidence, not brevity itself.
Fill in the blank
4. Fill in the blank: the missed-escalation rate rose from 1.2 percent to ___ percent before the fix shipped.
Show hint
Look at the bar chart comparing before, at discovery, and after.
Show answer
3.1 percent. After the redesign, it returned to about 1.0 percent, slightly better than where it started.
Short answer, apply it yourself
5. Pick a product you use yourself. What's one habit it built in you that you'd stop doing if it got a little worse?
Show hint
Think about a habit like spot-checking, double-typing, or trusting a suggestion without reading it fully.
Show answer
Model answer: Most people can name a moment they stopped double-checking an autocomplete or a recommendation once it had been right enough times in a row.
Short answer, the number question
6. If clean-note rate had only risen to 55 percent instead of 78 percent, would the missed-escalation spike still be as severe? Why or why not?
Show hint
Look at how the two charts relate to each other across the same five months.
Show answer
Model answer: Likely less severe, since fewer notes would have been sanitized overall, but the same underlying mechanism, brevity read as confidence, would still exist and would still need fixing.
Before you close the answer
Why this works
Tests whether you understand that sector regulation changes which metric you're accountable for, not just how much paperwork you file, and whether you can find a leading indicator instead of settling for a lagging one.
Follow-up traps
"Isn't this just a training problem, tell staff to write better notes?" Response: training helps briefly, but the underlying incentive, shorter notes score faster, stays in place and will win out again under time pressure unless the model itself stops rewarding brevity.
"How do you know clean-note rate isn't just reflecting genuinely simpler cases?" Response: by comparing it against an independent complexity signal, like vitals and symptom count, that doesn't move with note length, which is exactly what caught the gap here.
If pressed
Meridian's real fix runs a separate, smaller model just to estimate case complexity from structured vitals alone, so the triage score's confidence can be checked against a signal the front desk can't accidentally edit.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.