ConceptIntermediateModel Fluency & the AI PM Role / AI PM vs traditional PM vs technical PM / #3
Explain why an AI PM often owns the evaluation set while a traditional PM would not own a test plan.
LEAD · an AI posture and form-check coach that grades a squat differently depending on the room it is filmed in
Formshift is Kinetra Labs' app for people who lift at home: point a phone at a set of squats and it tells you, in a few seconds, whether your knees are caving in or your depth is off. Idelette Draeger owns the question of when a Formshift verdict is actually good enough to trust. Three releases in, her own eval set was already telling a story nobody was reading by slice.
The direct answer
An AI PM owns the eval set because the bar for good enough to trust has to move every time real bodies, angles, and rooms shift, and somebody has to own deciding where that bar sits right now. A traditional test plan checks fixed rules that do not change, so a QA team can run it and hand back a pass or a fail. You cannot hand off a decision that has to keep moving to a team whose whole job is checking a decision that already stood still.
Do this, in order
Own the eval set yourself, do not hand it to QA as a fixed pass or fail check.Why: the bar for good enough to trust has to move as real use shifts, and a checklist cannot move itself.
Score every slice on its own, not one blended number.Why: a real gain in the easy conditions can hide a real loss in the one slice that actually matters.
Treat the eval set's own slice scores as your earliest warning.Why: a hard slice sliding down shows up inside the eval set weeks before a single user complains.
Set a floor no slice can fall under, no matter how good the average looks.Why: an average that keeps climbing gives everyone permission to stop checking underneath it.
Feed every real production miss back into the eval set.Why: a set that never grows keeps testing yesterday's users while today's have already moved on.
Gate only the slice that is actually broken.Why: holding back a whole release to fix one corner throws away real, honest gains everywhere else.
How to answer this, stage by stage
Nobody is grading whether you know the words eval set and test plan. They are grading whether you can say, plainly, why one of them needs an owner and the other one just needs a runner.
1
Scope it to one real product and one owner
Say it like this
"Let's ground this in one product. Formshift is Kinetra Labs' app that watches a workout video and tells you if your form's off. Idelette Draeger owns what good enough to trust actually means for it. That's the seat I'll answer this from."
Why this works
A named owner keeps the whole answer about a real decision instead of a definition of two vocabulary words.
2
Say your structure out loud
Say it like this
"I'll run this as LEAD. Link, what a passing score is actually supposed to protect. Early signal, what tells you it's slipping before anyone complains. Abuse, how the number gets gamed. Decision, what I actually get to do about it that a test plan never lets anyone do."
Why this works
Two seconds of structure tells the interviewer you have a method, not just an opinion about ownership.
3
Reframe the question in one breath
Say it like this
"A test plan checks fixed rules: does the button work, does the charge go through. Those rules don't move, so QA can run them and hand back pass or fail. An eval set checks whether a model that gives a different answer every time is still good enough to trust right now, on real bodies, in real light. That bar has to move as the world moves, and somebody has to own moving it."
Why this works
This is the whole answer in miniature. Skip it and everything after sounds like a preference for one job title over another.
4
Give the decision, committed
Say it like this
"So here's what I'd actually do. I'd own the eval set myself, not hand it to QA as a fixed regression. I'd score it by slice, not one blended number. And I'd hold a floor: no camera angle, no lighting, no clothing type ships under 75 percent, no matter how good the average looks."
Why this works
This is the direct answer, said out loud, before a single number shows up to distract from it.
5
Prove it with the real case, numbers first
Say it like this
"Here's the case that made this real. Formshift's blended squat accuracy went 82, 87, 92 percent across three releases, a clean win on paper. But the side-on, dim-light, hoodie slice, the exact setup a lot of evening lifters actually use, went the other way: 78, 68, 55. Same eval set, same three releases, opposite direction, and the blended number never once showed it."
Why this works
A real number with a real direction beats any amount of talk about ownership in the abstract.
6
Name the abuse before the interviewer does
Say it like this
"Three ways that number lies to you. The team can quietly tune to the same 1,180 clips it reruns every release instead of the real world. The set can start out easy, because a well-lit, front-on clip is faster to label than a blurry side angle. And it can just go stale, since it keeps testing the old distribution while real filming habits keep moving."
Why this works
Naming the failure mode yourself is stronger than waiting for a follow-up question to expose it.
7
Say what you would not touch, and what it costs
Say it like this
"I wouldn't touch the front-on, bright, fitted-clothing slice. It's held at 94 percent through all three releases and it's most of our users. And fixing the hard slice isn't free: the fallback check we added there runs 2.4 seconds slower, on purpose, only for that slice, because slow and honest beats fast and wrong on a knee."
Why this works
Naming a real trade-off and a place you would leave alone shows judgment instead of blanket caution.
8
Close on the one line
Say it like this
"So: a test plan checks a fixed rule someone else can run. An eval set decides what good enough means today, and that decision has to live with whoever owns whether the product is actually safe to trust, not whoever's asked to run the same 1,180 clips one more time."
Why this works
Leaves the room with the decision, not just a well-told story about a squat app.
Let's learn
Formshift is Kinetra Labs' app for people who lift at home. Point a phone at a set of squats and it tells you, in a few seconds, whether your knees are caving in or your depth is off, the same two things a trainer would watch for.
A test plan checks a rule that already stood still. An eval set checks a bar that has to keep moving.
Before Formshift, a home lifter had two real options. Pay a trainer sixty dollars a session to watch a set in person. Or film it on a phone propped against a shoe box and guess, which is hard to do about your own knee from an angle you can barely see. Most people picked guessing, or gave up watching their form at all.
With Formshift, the same lifter gets a verdict in under three seconds, on nearly every set they film, for free. The average user films four to five sets a week and gets specific feedback on each one, not a trainer's once-a-month glance.
Five steps. The fourth one, how sure the pose model actually is about where a knee sits, is the step this whole answer turns on.
Knowledge spark: what is a slice?
Formshift's eval set is 1,180 clips, hand scored by certified trainers for real form errors. A slice is one narrow piece of it, say, side angle, dim lamp light, loose clothing. The blended score averages every slice together. A slice score checks one on its own.
Idelette's team retrained the model three times in five months. Formshift's blended squat accuracy, how often its knee-cave call matched a trainer's own label across the whole 1,180-clip set, climbed release over release: 82 percent, then 87, then 92. On paper, that is three clean wins in a row.
The eval set's own hardest slice, charted across three releases
Blended squat accuracy, all 1,180 clipsHardest slice: side-on, dim light, loose clothing
The blended number climbed for three releases straight. The slice a lot of evening lifters actually film in had already fallen through the 75 percent floor nobody had drawn yet.
The eval set was never wrong about the average. It was wrong about which average anyone was reading.
Here is the turn. Split the same eval set by camera angle, lighting, and clothing, and the side-on, dim-light, loose-clothing slice, the exact setup a lot of evening lifters actually use, moved the other way: 78 percent, then 68, then 55. Same eval set. Same three releases. Opposite direction. And the one number leadership watched, the blended 92, never once showed it.
What it costs at its worst: a caved-in knee that never gets flagged is not a small annoyance, it is a way to actually get hurt, repeated for weeks by a tool someone trusted specifically because it had never been wrong before. And even without an injury, the same slice of users just quietly trained less. Average completed sessions a week for that segment fell from 4.2 to 2.6 after the third release shipped.
Average completed sessions per week, hardest-slice users, before vs after v3.3 shipped
Before v3.3 shippedAfter v3.3 shipped
This did not show up for six weeks, not until someone finally split the retention numbers by filming setup. The eval set's own hard slice had already told the story three releases earlier.
The choice I would take back
Eighteen months earlier, when Formshift did one exercise and almost everyone filmed it in a bright room facing the camera, the launch review handed eval-set ownership to QA: rerun the same clip set every release, ship if the blended score holds or climbs. That was a sensible call for one exercise and one filming habit. It stopped being sensible the moment Formshift grew to seven exercises and real users started filming in the dark, from the side, in a hoodie, because nobody had standing authority to add a clip, move the bar, or hold a release over one slice. QA's job was to run the plan, not redesign it.
What I would leave alone: the front-on, bright, fitted-clothing slice. It held at 94 percent through all three releases, and it covers most of Formshift's users. Gating that slice too would just slow down everyone the retrain never actually hurt.
The lesson: a test plan checks a rule that does not move. An eval set decides what good enough means today, for bodies and rooms that keep moving, and that decision has to sit with whoever owns whether the product is actually safe to trust, not whoever is asked to run the same 1,180 clips one more time.
Now here is the same thing as a story
The short version above is what you actually say in the room. Read this one when you want to feel why a new hire's one question mattered more than three straight releases of good news.
Idelette Draeger has run product for Formshift for two years, and before that she spent six months building its very first eval set herself, filming her own squats in her apartment in six different setups so the labelers would have something real to score against.
Formshift launched doing one thing: squats, filmed front-on, in daylight, because that is how the earliest users happened to train. For the first year, that was basically the whole business. QA reran the same 1,180 clips every release, blended accuracy climbed steadily, and Idelette signed off without a second look, because the number had never once lied to her.
Nobody at Kinetra decided any of this was happening. It just happened, the way a season does.
Kinetra kept growing the exercise list, deadlifts, lunges, push-ups, plank holds, overhead presses, hip hinges, and its users kept growing with it. A lot of newer lifters, people like Belmara Krenwick, an engineer who trains most evenings after work, started filming later in the day, propping a phone low near the couch instead of a tripod by a window, wearing whatever she had worn to the office instead of fitted gym clothes.
Three releases into that shift, a new applied scientist, three weeks into the job, asked a question in a routine eval review that nobody could answer: why did the 1,180-clip eval set only have about forty clips filmed after seven at night, when close to half of Formshift's daily sessions happen after eight? Nobody had an answer, because nobody had ever asked.
Idelette pulled the eval set apart by slice for the first time in months. The blended number was exactly what everyone remembered, climbing, 82, 87, 92. Split by camera angle, lighting, and clothing, the side-on, dim-light, loose-clothing slice, Belmara's exact setup, told a different story: 78, 68, 55.
We did not build a worse model. We built a model that was quietly worse at seeing the one thing it was supposed to catch, for exactly the users who needed it most.
Belmara never filed a complaint. She had no reason to; Formshift kept saying clear. She just trained less. Her weekly sessions slipped from about four to just over two and a half over six weeks, and she told herself it was just a busy season at work, not a tool that had quietly stopped seeing her knees.
The decision Idelette would take back happened in a launch review eighteen months earlier, back when Formshift did one exercise. Someone asked whether QA should own re-running the eval set every release as a straight pass or fail. The room said yes, reasonably: one exercise, one filming habit, a fixed check was fast and it worked. Nobody ever came back to ask whether that was still true once there were seven exercises and users training in the dark.
Run the same eighteen months again, with the eval set owned by product from day one instead of handed to QA as a fixed regression. Every slice gets its own floor, 75 percent, checked separately from the blended number. The side-on, dim-light slice trips that floor the same week the third release ships, not six weeks later in a cohort review nobody was running yet. Formshift adds a heavier, slower pose check just for that slice, 2.4 seconds slower, and pulls in real clips of users like Belmara to retrain against. Three weeks after that, the slice is back to 81 percent, and Belmara's segment is back to almost four sessions a week.
What I would tell myself, back in that first launch review: a checklist that never changes is fine for a button that never changes either. It was never going to be fine for anything that has to see a knee bend in the dark.
LEAD: the four things only an eval set's owner gets to decide
Not a way to make an eval set sound more official than a checklist. LEAD is what forces you to say which slice is the one worth losing sleep over, and what you would actually do about it.
Four checks, run on one eval set. Skip any one of them and a real regression can climb three releases in a row without anyone catching it.
LLink. What is the real outcome?
Not whether Formshift's blended score is climbing. The real outcome is whether a lifter who sees a clear verdict can actually trust it not to miss the one thing that could hurt them, and whether they keep training with the app at all.
Grading the model against its own reused clips and grading the product against a real knee, in a real room, are two different questions. Only one of them protects the person on the mat.
EEarly signal. What moves first?
The eval set's own hardest slice is the early signal, not a separate metric bolted on. Side-on, dim-light, loose-clothing accuracy fell 78, 68, 55 across the same three releases the blended score climbed 82, 87, 92. That is not noise sitting inside a good number, it is the real story, three releases and six weeks before a single retention chart showed it.
This is the answer to the question, in one line. The eval set does not just measure the model. Split correctly, it warns you before the product does.
The eval set's hard slice sits top left on purpose. It was measured, and it was early, weeks before anything else on this page.
AAbuse. How does it get gamed?
Three ways this number lies. A team can quietly tune to the same 1,180 clips it reruns every release instead of the real world outside them. The set can start out easy, because a well-lit, front-on clip is faster and cheaper to label than a blurry side angle, so the hard cases are underrepresented from day one. And it can go stale as usage moves, testing yesterday's filming habits while today's users have already moved on.
Every metric can be hit without doing the real work. An eval set that never changes is the cheapest way to hit this one.
A model can clear the same 1,180 clips every single release. That was never proof it could see Belmara's living room.
DDecision. What actually changes?
Owning the eval set means three things a QA-run test plan structurally cannot do. Move the bar: hold a hard floor, 75 percent, per slice, no matter how good the average looks. Add new failure cases: every real miss production finds gets its actual clip fed back into the next eval cycle. Gate a launch by slice, not wholesale: route the failing slice to a slower, heavier check while everyone else ships on time.
A metric nobody acts on is decoration. This is what actually ships, what actually holds, and who gets to say so.
Three things worth saying plainly, since this is where the real judgment sits. Idelette's team considered two other fixes and rejected both. Leaving the eval set with QA as a fixed regression, which was the actual arrangement for eighteen months and the one that let this slide. And simply raising the blended pass bar, say to 96 percent, instead of scoring by slice, which would not have caught this either, since the blended score was already climbing past 90 while the real slice sat at 55. The AI-specific failure underneath both is a keypoint problem, not a bug: side angle, dim light, and loose fabric all make it harder for the pose model to see exactly where a knee sits, so its confidence drops in a way a blended average was never built to isolate. The guardrail is the per-slice floor itself, fed by real production misses, not a smarter blended threshold. And the fix is not free: the heavier check on that one slice costs 2.4 seconds of extra wait, on purpose, only there, because slow and honest beats fast and wrong on somebody's knee.
And if you want to be sure it really works, try it somewhere else
Same four letters, a drone instead of a phone, and this time a missed call means a farmer skips a field walk they actually needed to take.
Fieldglass Aerial builds LeafScope, which reads drone photos of a field and flags early blight before it spreads. Amoret Isibor owns LeafScope's eval set the same way Idelette owns Formshift's: not a QA checklist, a living judgment call about what good enough to trust means this season.
Different sensor, same shape of mistake. The blended number climbed while one real slice quietly went the other way.
LeafScope's own eval set is 2,400 labeled drone images across four crops, three growth stages, and both sunny and overcast light, since cloud cover changes how a leaf's color reads from four hundred feet up. Blended blight-detection accuracy climbed for two straight model versions, 88 to 93 percent. Split by slice, overcast days on late-stage soybean, when the canopy is thickest and the light is flattest, fell from 81 to 62 percent over the same two versions, because the model had far fewer late-stage, overcast clips to learn from than early-stage, sunny ones, the conditions most of Fieldglass Aerial's corn customers actually fly in.
Run the same four letters. Link: not the offline detection score, but whether a farmer who sees no blight detected can skip walking that field themselves without missing a real outbreak. Early signal: the eval set's own overcast, late-stage slice, falling while the blended score climbed, the same shape as Formshift's. Abuse: the easy slice, sunny, early-stage corn, is what most of the labeled set was built from first, because it is the biggest customer base and the clearest images to hand-label, so the hard slice stayed thin. Decision: Amoret set a slice floor at 75 percent, the same number Idelette landed on independently, added real farmer-flagged misses from overcast fields to the next eval cycle, and routed any overcast, late-stage flag through an agronomist's review instead of auto-clearing it, while every other slice ships as scored.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: own the eval set, score it by slice, hold a floor no slice can fall under, and feed real misses back in, because a test plan checks a rule that never moves and this one has to.
Cost: no budget this quarter to grow the eval set. Ship with the thin slice honestly labeled as thin, and grow it with every real miss instead of waiting to launch until it is bigger.
The model got better, for real: say the hardest slice climbs to 90 percent on its own next release. The floor does not go away, because 90 percent still is not the 94 the easy slice has held for a year, and slices keep drifting even when they look healthy today.
Where people run it wrong.
They watch the blended score because a segmented one feels like admitting the launch might not be a clean win.
They treat a fixed eval set as proof of quality forever, instead of a snapshot of whatever the world looked like the day it was built.
They fix the whole model when only one slice is actually broken, and lose real, honest gains everywhere else.
How to use it live. Ask yourself one question before trusting any climbing benchmark: what does this look like split by the slice that was always going to be hardest, not averaged across the slice that was always going to be easy? That question alone is usually the exact distinction a LEAD question is listening for.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
LEAD: link to the real outcome, find the early signal that moves first, name how it gets gamed as abuse, decide what actually changes. Built for questions about who owns a moving bar, not a fixed checklist.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Idelette Draeger, who owns Formshift's eval set at Kinetra Labs, and has to decide whether three releases of a climbing blended score are actually a win.
3 · THE HABIT
What did the team stop doing because the blended number kept looking fine?
Tap to flip
ANSWER
They stopped pulling the eval set apart by slice, camera angle, lighting, clothing, and just watched the one blended squat-accuracy number climb release after release.
4 · THE DIVERGENCE
What two numbers pulled apart here, and which directions did they move?
Tap to flip
ANSWER
Blended squat accuracy, 82 to 87 to 92 percent, up. The side-on, dim-light, loose-clothing slice, 78 to 68 to 55 percent, down. Same eval set, same three releases, opposite direction.
5 · THE OLD DECISION
What decision would Idelette take back?
Tap to flip
ANSWER
Handing eval-set ownership to QA eighteen months earlier as a fixed regression: rerun the same 1,180 clips every release, ship if the blended score holds or climbs, with nobody authorized to move the bar or add new clips.
6 · THE NUMBER
Fill in the blank: Formshift's eval set has ___ labeled clips. The hardest slice fell to ___ percent while the blended score climbed to ___ percent.
Tap to flip
ANSWER
1,180 labeled clips. The hardest slice fell to 55 percent while the blended score climbed to 92 percent, on the same three releases.
7 · THE REPLAY
Same eighteen months, eval set owned by product from day one, what changes?
Tap to flip
ANSWER
The failing slice trips a 75 percent floor the week it ships, not six weeks later. A heavier check recovers it to 81 percent in three weeks, and Belmara's segment climbs back to almost four sessions a week.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs LEAD again on a different product. Which one, and what's the equivalent hard slice?
Tap to flip
ANSWER
LeafScope, Fieldglass Aerial's crop-disease drone tool, owned by Amoret Isibor. The equivalent hard slice is overcast days on late-stage soybean, which fell from 81 to 62 percent while the blended blight-detection score climbed.
Check yourself Score: 0 / 0
Multiple choice
1. Why did Idelette refuse to treat Formshift's eval set like Kinetra's subscription checkout test plan?
A. Because the eval set was too expensive for QA to run every release.
B. Because a checkout test plan checks a fixed rule that never moves, while the eval set has to be redrawn and re-weighted as real bodies, angles, and lighting shift.
C. Because QA was not certified to review fitness content.
D. Because the eval set only mattered for Formshift's marketing claims.
Show hint
Look at Stage 3 in the walkthrough, the reframe.
Show answer
B. A checkout flow's rules do not change, so a fixed test plan run by QA is the right tool. A pose model's real-world accuracy does change, so the bar for good enough has to be owned by someone who can move it.
True or false
2. True or false: once Formshift's blended squat accuracy hit 92 percent, that was proof v3.3 was ready to ship to every user.
True
False
Show hint
Check the Abuse step in the LEAD recap.
Show answer
False. The blended 92 percent blended a real, honest gain in the easy slices with a real loss in the side-on, dim-light, loose-clothing slice, which had fallen to 55 percent on the same release.
Fill in the blank
3. Formshift's eval set has ___ labeled clips. Across three releases, the blended squat accuracy climbed to ___ percent while the hardest slice fell to ___ percent.
Show hint
Check the line chart titled "The eval set's own hardest slice, charted across three releases."
Show answer
1,180 labeled clips, 92 percent, 55 percent. The gap between those two numbers, on the exact same eval set, is the whole reason this needed an owner and not just a runner.
Short answer, name the rejected alternative
4. Besides leaving the eval set with QA, what other fix did Idelette's team consider and reject, and why did it lose?
Show hint
Look at the "three things worth saying plainly" paragraph near the end of the LEAD recap.
Show answer
Model answer: Simply raising the blended pass bar, say to 96 percent, instead of scoring by slice. It lost because the blended score was already climbing past 90 while the real hard slice sat at 55, so a stricter blended bar would not have isolated the one slice that actually mattered.
Short answer, apply it yourself
5. Think of an AI product you use or have built. What real-world condition, body type, accent, lighting, device, connection speed, could split its eval set into an easy slice and a hard one nobody is watching separately?
Show hint
Think about who the eval set was probably built from first, and who joined the product later.
Show answer
Model answer: A voice assistant's speech-to-text eval set built mostly from clear, native-accent recordings in quiet rooms. The hard slice, strong regional accents over a phone's built-in mic with background noise, could sit well below the blended accuracy for months without anyone splitting it out.
Short answer, work the number
6. Say a new Formshift exercise, plank hold filmed in a bathroom mirror, comes into the eval set at 70 percent accuracy. Does it ship broadly under the rule Idelette set, and what has to happen first if not?
Show hint
Check what the 75 percent floor actually gates, and what the fallback looks like for a failing slice.
Show answer
No, it does not ship broadly. At 70 percent it falls under the 75 percent floor, so it routes to the slower fallback check the same way the side-on, dim-light slice did, and it needs real production clips fed back into the eval set and a fix before it clears 75 on its own.
Before you close the answer
Why this works
Tests whether you understand that a model's grade is a moving judgment call, not a fixed check, and that someone has to own moving it. A candidate who treats the eval set like a bigger test plan has not actually located the difference the question is asking about.
Follow-up traps
"Why not just have QA run the segmented breakdown instead of moving ownership to product?" Response: someone still has to decide the floor, decide when a real production miss earns a new clip, and decide when a failing slice is bad enough to hold a launch, and that is a judgment call about what good enough to trust means, not a thing you can write into a checklist in advance.
"Isn't holding every release until every slice clears the same bar the safer move?" Response: no, that would gate the 70 percent of users whose slice never had a problem, for a fix only one or two slices actually needed, and it throws away real, shipped gains everywhere else.
If pressed
The 75 percent floor is not a ship-time-only check. It gets recomputed nightly against the last 30 days of real, flagged production clips, so a slice that passed at launch can still trip the floor later if real usage drifts underneath it, without waiting for the next model release to find out.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.