CaseIntermediateAI Opportunity & Model Strategy / Model selection from a PM lens / #15

Describe how you would evaluate a model on your own domain-specific terminology.

LEAD92 percent on a generic benchmark, 61 percent on the words Tobias actually writes

Say we build a tool that reads a physical therapist's spoken shorthand and drafts a home exercise program a patient takes home on a clipboard. The vendor's own accuracy number for that tool was 92 percent. Almost none of that 92 percent was measured on the words a physical therapist actually uses.

The direct answer
Don't trust a vendor's general accuracy score for a domain you didn't build the eval for. Pull real notes from your own clinic, build a small golden set of the specific terms that matter, precaution flags, rep and set shorthand, joint-angle abbreviations, and track the rate a person still has to correct those specific terms by hand. That correction rate moves weeks before overall trust or satisfaction would ever show a problem.
Do this, in order
  1. Build a golden set from your own real notes, not the vendor's generic benchmark.Why: a domain's specific shorthand almost never shows up in a general-purpose eval set.
  2. Track the manual-correction rate on your highest-stakes terms specifically.Why: this is the leading signal, it drifts for weeks before any overall satisfaction score would move.
  3. Set a different confidence threshold for different note types.Why: a routine strength program and a joint-replacement precaution flag don't carry the same cost if the model gets it wrong.
  4. Watch for the eval being gamed by easy cases.Why: a golden set stacked with common phrasing can look great while the rare, critical terms still fail quietly.
  5. Refresh the golden set whenever the model or the clinic's own shorthand changes.Why: a domain eval that's never updated ages into the same blind spot it was built to catch.

How to answer this, stage by stage

The interviewer isn't testing whether you know what an eval is. They're testing whether you'd trust a vendor's number for a domain that number was never built for.

Stage 1
Scope it to one real system
Say it like this
"Let's ground this in one product. It reads a physical therapist's dictated notes and drafts a home exercise program. I'll answer for how I'd evaluate that model on the clinic's own terminology."
Why this works
Stops "evaluate the model" from becoming an abstract lecture on eval methodology.
Stage 2
Say the structure out loud
Say it like this
"I'll run this as LEAD. Link, the outcome that actually matters. Early signal, the thing that moves first. Abuse, how the metric gets gamed. Decision, what I'd actually do at each threshold."
Why this works
Signals a repeatable way to think about domain evaluation, not a one-off checklist.
Stage 3
Reframe: the vendor's score measured the wrong words
Say it like this
"92 percent sounds like a strong number, until you ask what it was measured against. If it's a general medical benchmark, it likely never saw a single joint-replacement precaution flag, which is exactly the term this clinic can't afford to get wrong."
Why this works
This is where a strong candidate separates from one who just repeats the vendor's marketing number.
Stage 4
Give the one decision: build the golden set
Say it like this
"I'd pull a few hundred real notes from the clinic's own history, hand-label the precaution flags, rep and set counts, and joint-angle shorthand, and score the model against exactly that, not the vendor's benchmark."
Why this works
This is the direct answer, and it's checkable: you can point to the exact golden set you'd build.
Stage 5
Prove it with the compressed failure
Say it like this
"When Anchor Point actually ran the vendor's model against a domain golden set, the 92 percent generic score dropped to 61 percent on precaution flags specifically. A new hire had already caught one wrong precaution line on a total hip patient's home program before anyone built the eval at all."
Why this works
Compresses the whole argument into the one gap between a vendor's number and the clinic's own reality.
Stage 6
Name the AI-specific reasoning, the trade-off, and close
Say it like this
"A model's accuracy is only ever accuracy on whatever it was tested against. Testing it against your own domain costs time nobody budgeted for, usually a few weeks. I'd take that cost every time over a precaution flag reaching a patient wrong. So: build the domain golden set, gate high-stakes note types behind it, and let routine notes ship faster."
Why this works
Names the load-bearing AI-specific judgment and states plainly what's being traded for it.

Let's learn

The tool listens to a physical therapist dictate a home exercise program and drafts it into a printed sheet the patient takes home.

Before it, Tobias Lindqvist wrote every home program from nothing, about 12 minutes a patient, across roughly 14 patients a day.

Once the vendor's 92 percent accuracy score convinced the clinic to trust the tool, Tobias started skimming and lightly editing drafts instead of rewriting them, about 3 minutes a patient.

Here's the turn: the 92 percent was never wrong, exactly. It was measured against a general medical benchmark that had almost no PT-specific precaution language in it at all. The number stayed excellent while the thing that actually mattered, whether a joint-replacement patient's home program listed the right movement limits, quietly got worse.

Hand sketched icon list titled Terms that break a generic medical eval. Four rows: a document icon, AROM and PROM active versus passive range of motion. A gauge icon, PNF a specific neuromuscular stretch cue. A scale icon, closed kinetic chain versus open kinetic chain phrasing. A question mark box icon, joint replacement precaution flags flexion and adduction limits.
None of these four show up reliably in a benchmark built for general medical language.

At its worst, trusting a vendor's generic score means a patient recovering from a hip replacement takes home written instructions that quietly violate their own surgeon's precautions.

The choice I would take back Anchor Point accepted the vendor's 92 percent accuracy figure as proof the tool was ready across every note type. That made sense when nobody had separated "accuracy in general" from "accuracy on the twelve terms a wrong answer could actually hurt someone." It stopped making sense the moment a new hire asked why a precaution line didn't match the surgeon's actual restrictions.

What I would leave alone: the tool's handling of routine strength-and-flexibility phrasing, sets, reps, simple stretches, was already accurate and didn't need its own special golden set. The caution belongs on the terms where a mistake is expensive, not on every term equally.

The lesson: a vendor's accuracy score tells you how the model did on their test. It tells you nothing about your domain until you've built a test out of your own words.

Hand sketched flow diagram titled Building a domain eval start to ship. Five steps left to right, the middle one emphasized: Pull real PT notes. Build a precaution term golden set. Score the model against it. Set a per term threshold. Gate auto drafts by confidence.
The vendor's benchmark never appears in this pipeline at all. It was never built for this clinic's words.

Now here is the same thing as a story

The short version above is what you'd say in a vendor-evaluation meeting. Read this one for the day a new hire actually asked the question nobody else had.

Tobias Lindqvist had run Anchor Point's outpatient floor for six years, and he could write a home exercise program faster by hand than most therapists could type one.

The tool arrived with a 92 percent accuracy claim on its cover sheet, and it earned trust fast: drafts came back clean, patients left with printed sheets that looked exactly like what Tobias would have written himself, most of the time.

Knowledge spark: why would a 92 percent score hide a 61 percent problem? An accuracy score is only ever a score on whatever it was tested against. A vendor's benchmark built from general medical text can score high while never once testing a domain's specific shorthand, precaution flags, joint-angle abbreviations, rep and set notation, because those terms were never in the test to begin with.

A new hire, three weeks in, printed a home program for a patient recovering from total hip replacement surgery and paused. The draft told the patient to perform a stretch that crossed the leg past the body's midline, a movement standard post-op hip precautions explicitly rule out. She asked Tobias why nobody had caught this before.

The 92 percent was real. It just never once touched the twelve words this clinic actually depends on.

Tobias didn't have an answer, because nobody had ever tested the tool against Anchor Point's own precaution language. The clinic pulled 240 real notes, hand-labeled every precaution flag, rep and set count, and joint-angle term, and ran the model against exactly that.

Hand sketched metaphor scene titled How it gets gamed. Left, a document icon labeled Looks done, caption checkbox ticked on the vendor's generic benchmark. Right, a person icon in a different color labeled Walks away, caption Tobias still rewrites the precaution line by hand.
A benchmark can look finished on paper while a person is still quietly doing the real work by hand.

The domain score came back at 61 percent on precaution flags specifically. The clinic's overall satisfaction survey, taken the same month, still read fine, because most notes were routine and the model handled those well. The survey was the lagging signal. It hadn't moved yet, and it wouldn't have for weeks.

Hand sketched comparison diagram titled Two signals, one leads. Left panel, a gauge icon labeled Leading signal, caption manual correction rate on precaution terms. Right panel, a scale icon in a different color labeled Lagging signal, caption overall therapist trust survey score.
One of these two signals would have caught the problem weeks before the other ever noticed.

When the tool was first adopted, someone in the vendor-selection meeting said, "92 percent is well above what we'd expect from a human, let's move forward." Reasonable, since nobody at the table knew to ask what that 92 percent actually covered.

Manual correction rate on precaution terms, twelve weeks
50% 25% 0 week 12, still climbing satisfaction survey, flat Wk 1 Wk 12
The precaution-term correction rate had been climbing for twelve weeks before the incident that finally got noticed.
Precaution-flag errors reaching a printed home program, per month
16 8 0 14 Before the gate 1 After the gate
The lagging outcome the correction rate was quietly predicting for twelve weeks: how many precaution mistakes actually reached a patient's hands.

The real question was never whether the model was good overall. It was whether it was good on the twelve specific terms this clinic could not afford to get wrong.

What I'd tell myself, hearing the new hire's question: a vendor's benchmark tells you how a model did on somebody else's words. It tells you nothing about yours until you check.

LEAD, held against one clinic's shorthandNot a lecture on eval design. The specific golden set you'd actually build before trusting a draft.

Hand sketched labeled parts diagram titled What's in a PT golden set. A document icon at the center labeled PT Golden Set, with four labeled callouts around it: Precaution flags, Rep and set notation, Joint angle shorthand, Home program phrasing.
A domain eval isn't one big test. It's four specific categories, weighted by how much a mistake in each one actually costs.
L
Link. The outcome that actually matters.
Therapists sign a drafted home program without rewriting it, freeing time for patients, without a precaution flag ever being wrong.
Not the model's score. The actual outcome the score is supposed to stand in for.
E
Early signal. The thing that moves first.
The manual-correction rate on precaution-flag terms specifically, which climbed for twelve weeks before anything else looked wrong.
This is the hard step: finding the number that moves before the obvious one does.
A
Abuse. How the metric gets gamed.
A golden set stacked with routine, easy phrasing would score high while still missing the rare, high-stakes precaution language entirely.
Every metric has a way to be satisfied without doing the real work.
D
Decision. What you'd actually do at each threshold.
Below a set accuracy bar on precaution terms specifically, every draft with a flag gets required review. Above it, routine programs auto-ship with a weekly spot-check.
A metric with no attached decision is a dashboard nobody acts on.

The recap, one line per letter: link is trusting a draft without rewriting it, early signal is the precaution-term correction rate, abuse is a golden set stuffed with easy phrasing, decision is gating high-stakes note types behind a real threshold.

And if you want to be sure it really works, try it somewhere elseSame four letters, a port authority instead of a physical therapy clinic. A different kind of hidden vocabulary this time: customs and cargo-manifest shorthand.

Farrukh Rashidov manages cargo-manifest review for Amberlyn Port Authority, using a model that reads shipping documents and flags manifests needing a customs hold. Mapped onto LEAD: link is customs officers trusting a flag enough to act on it without re-reading the full manifest; early signal is the manual override rate on a specific set of hazardous-materials classification codes, which drift for weeks before the port's overall clearance-time metric would ever move; abuse is a golden set built only from routine consumer-goods manifests, missing the rare hazardous codes that actually matter; decision is requiring a second reviewer on any flag touching those specific codes, while routine manifests clear on the model's word alone.

Hand sketched decision tree titled Does this auto draft need therapist review. Root, a home program draft is ready to sign. Branches: has a joint replacement precaution flag leads to require therapist review always. Routine strength and flexibility program leads to auto ship spot check weekly. Confidence below threshold on any term leads to flag for review before it reaches a patient.
The same three-branch logic sorts a home exercise program and a cargo manifest by the same real question: how much does a wrong answer here actually cost.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "build a golden set from your own words, track correction rate on the highest-stakes terms specifically," and stop.
Cost: there's no budget to hand-label hundreds of notes. Say so honestly, and start with the twenty highest-stakes terms only, not the whole vocabulary.
The model got better, for real: if a new model version genuinely improves, the domain golden set still earns its cost, because it tells you exactly which specific terms actually improved, not just an average that moved.

Where people run it wrong.
They treat a vendor's published accuracy number as proof the model works for their specific domain.
They build a golden set only from the easiest, most common examples, missing the rare terms a mistake would actually hurt.
They track only lagging signals like overall satisfaction, missing the specific correction rate that would have warned them weeks earlier.

How to use it live. When an interviewer asks how you'd evaluate a model on domain terms, ask yourself: which specific words, if the model gets them wrong, actually cost someone something? Build the eval around those words first.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits evaluating a model on domain-specific terminology?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. It finds the signal that moves before an overall satisfaction score ever would.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Tobias Lindqvist, who has run Anchor Point Physical Therapy's outpatient floor for six years.
3 · THE HABIT
What did Tobias stop doing once the tool's 92 percent score earned trust?
Tap to flip
ANSWER
Writing every home program from nothing. He started skimming and lightly editing drafts instead, cutting his time from 12 minutes to about 3.
4 · THE EARLY SIGNAL
What's the leading signal in this story?
Tap to flip
ANSWER
The manual-correction rate on precaution-flag terms specifically, which climbed for twelve weeks before anything else looked wrong.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Accepting the vendor's 92 percent generic accuracy score as proof the tool was ready across every note type, instead of testing it against the clinic's own terms first.
6 · THE NUMBER
Fill in the blank: the vendor's general score was ___ percent, but the domain score on precaution flags was ___ percent.
Tap to flip
ANSWER
92 percent general, 61 percent on precaution-flag terms specifically, from a 240-note golden set built from the clinic's own history.
7 · THE REPLAY
Same new-hire moment, domain golden set and threshold gating already in place. What changes?
Tap to flip
ANSWER
The precaution-flagged draft is caught by the required-review gate before it's ever printed, instead of reaching a patient's home program by accident.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which one, and what's the equivalent hidden vocabulary?
Tap to flip
ANSWER
Amberlyn Port Authority's cargo-manifest review tool. The equivalent hidden vocabulary is rare hazardous-materials classification codes.

Check yourself Score: 0 / 0

Short answer, name the gap
1. Why did a 92 percent accuracy score fail to catch the precaution-flag problem?
Show hint
Look at the knowledge spark about why a 92 percent score can hide a 61 percent problem.
Show answer
Model answer: The 92 percent came from a general medical benchmark that never tested PT-specific precaution language, so it measured the wrong words entirely.
Multiple choice
2. Why is the manual-correction rate on precaution terms a stronger signal than the overall satisfaction survey?
  • A. It's easier to collect than a survey.
  • B. It moves weeks before the survey would, because most notes are routine and don't affect overall satisfaction.
  • C. Surveys are always less accurate than any numeric metric.
  • D. Therapists don't fill out satisfaction surveys honestly.
Show hint
Look at the line chart comparing the two signals over twelve weeks.
Show answer
B. The correction rate climbed for twelve straight weeks while the satisfaction survey stayed flat, because routine notes, most of the volume, were still handled well.
True or false
3. True or false: this answer recommends building a domain golden set for every single term the tool might ever produce.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. The caution is targeted at high-stakes terms, like precaution flags. Routine strength-and-flexibility phrasing was already accurate and didn't need its own special eval.
Fill in the blank
4. Fill in the blank: the clinic built its golden set from ___ real notes pulled from its own history.
Show hint
Look at the story, right after the new hire's question.
Show answer
240 notes. Enough to hand-label every precaution flag, rep and set count, and joint-angle term the clinic actually uses.
Short answer, apply it yourself
5. Think of a domain you know well, cooking, a sport, a hobby. What's one piece of its specific shorthand a general-purpose model would likely get wrong?
Show hint
Think of a term or abbreviation an outsider would need explained to them.
Show answer
Model answer: A baking recipe's "fold, don't stir" instruction. A generic model might treat both as interchangeable mixing verbs, when the difference decides whether a batter collapses.
Short answer, work the number
6. If the domain score on precaution terms had come back at 85 percent instead of 61, would the same required-review gate still make sense?
Show hint
Think about what's actually at stake if even a small percentage of precaution flags are wrong.
Show answer
Model answer: Yes, likely still yes. Even at 85 percent, 15 percent of precaution flags would be wrong, and a wrong post-surgical precaution is expensive enough that the review gate is still worth keeping.
Before you close the answer
Why this works
Tests whether you'd trust a vendor's number at face value, or ask what it was actually measured against. Most candidates never ask the second question.
Follow-up traps
"Isn't building your own golden set just extra work the vendor should have done?" Response: the vendor's benchmark is built to sell across every customer, not to know your clinic's specific precaution language, so the extra work is unavoidable if you want a real answer.

"What if you don't have enough historical notes to build a real golden set?" Response: start with the twenty highest-stakes terms and grow the set over time, a small, focused golden set beats no domain testing at all.
If pressed
The golden set held out a separate slice of notes from a second clinic location, since a domain eval built entirely from one location's own shorthand can quietly overfit to one therapist's personal phrasing habits.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more