Describe how to run a model review meeting that produces decisions rather than admiration.
Loomcastle sells Truce, a model that untangles scheduling conflicts across everyone's calendars on its own, so an eleven-person meeting spread across four time zones gets a slot without twenty reply-all emails. Perilune Strade, a senior PM, has run Truce's weekly review meeting, Truce Review, for about fourteen months. This is the meeting where the VP of Product finally asked the one question nobody in the room could answer.
- Open the meeting on a named decision, not a demo.Why: this is the actual reversal; skip it and the room defaults to whatever looks best that week, with nothing forcing a call.
- Bring a deliberately hard case, chosen on purpose, every single time.Why: a review built only on wins never sees where Truce actually breaks, so a real ship risk drifts for months with no evidence against it.
- Close every meeting with a written decision and a named owner.Why: a decision that only lives in the room's mood evaporates the moment everyone leaves; a decision with an owner and a date can be checked later.
- Track a target rate for the hard-case segment, not just the average.Why: a 98 percent aggregate accuracy hides a segment failing one in three, and an average never flags a ship risk on its own.
- Keep the demo-only meeting, but call it what it is.Why: sales and leadership do need a highlight reel sometimes; the problem was never demos existing, it was letting one quietly become the only review anyone had.
- Review the decision log itself on a schedule, even in weeks nothing's on fire.Why: a meeting can run for months producing zero decisions and still feel productive, because nobody's watching the log, only the room's mood.
How to answer this, stage by stage
Nobody is grading whether you can describe a tidy meeting agenda. They are grading whether you'd notice the moment a real review turns into a highlight reel, and catch it before a live ship call sits untouched for months.
Let's learn
What happens when the meeting built to make a call on a model quietly turns into a meeting that just watches it do well?
Say a company sells calendar software. Inside it sits a model called Truce, which looks at everyone's calendars for a meeting request and works out a slot that actually fits, moving or splitting things where it has to. Every week, about 40,000 of these conflicts get resolved automatically across Loomcastle's customers. Truce's review meeting, Truce Review, started the same week the model launched: thirty minutes, every Thursday, in a small room. For its first five months, roughly twenty-two meetings, it opened with the worst miss since last time and closed with a written decision twenty times out of twenty-two.
Now, fourteen months in, Truce is genuinely better than it was: overall accuracy has climbed from 91 percent to 98 percent, one small release at a time. But that number hides a second one. On simple conflicts, two to four people, one or two time zones, Truce is right 98 times out of 100. On complex conflicts, eight or more people, three or more time zones, it's right 68 times out of 100. It fails on roughly one in three of the cases it was actually bought to handle.
Here's the turn. Those failed complex cases are not the real problem. The real problem is what the room does with its opening ten minutes once good wins get easy to find.
At its worst, a genuinely live risk sits untouched behind a room full of applause. Corvasta Bank, one of Loomcastle's larger accounts, ran an eleven-person leadership offsite across four time zones. Truce double-booked their CFO's board call against the offsite itself, a slot that would have started at 2am in Singapore for one attendee. Someone caught it the day before and fixed it by hand. Corvasta nearly walked. Nobody in Thursday's meeting connected that near miss to the version-four decision sitting in the queue, because the meeting that week opened, as it always did, with a clean demo of Truce handling a three-person same-timezone request.
What I would leave alone: the monthly Truce Showcase, a separate meeting sales runs for prospects, doesn't need this fix at all. Its whole job is to look good, and nobody in that room expects a decision to come out of it. The split only earns its keep on the meeting whose real job is deciding something.
The lesson: a review meeting doesn't hold its shape on its own just because it started as a decision-forum. Somebody has to name a real decision, on purpose, every single time, and that costs something real: hunting a genuine hard case before each meeting takes someone thirty or forty minutes of real digging, instead of grabbing whatever a Slack thread already praised. That's the trade worth accepting on purpose, slower prep, for a decision that's actually informed by where Truce fails.
Now here is the same thing as a story
The short version above is what you'd actually say in the room. Read this one when you want to feel exactly what five quiet months cost, meeting by meeting.
Perilune Strade can read a Truce failure log and tell, inside a minute, whether it's a real pattern or a fluke. She's run Truce Review since the model first shipped, fourteen months now, the same thirty-minute slot every Thursday.
For the first five months, that habit was simple and it worked. Every Thursday she'd go looking, on purpose, for the worst miss Truce had made since the last meeting, and she'd open with it. Nobody enjoyed those first ten minutes. But the room argued, decided something, and moved on. Twenty of the first twenty-two meetings closed with an actual line in the decision doc: ship this, hold that, kill this idea.
Then Truce kept getting better. Not in one leap, in the ordinary way software gets better: a release here, a fix there, 91 percent climbing toward 98 over more than a year. And good wins started arriving on their own, without her having to dig for anything. A clean save on a nine-person cross-office scheduling mess. A customer email thanking the team by name. Opening with one of those worked just as well to fill the room's attention, and it took a tenth of the effort.
So, without ever deciding to, she stopped hunting for the hard case.
By month nine, that was just how Thursdays worked. Whoever had the best recent save got the floor first, the room nodded along, and the meeting closed on time with nothing written down. Nobody voted on this. It was simply the easiest way to fill thirty minutes once good news kept arriving before anyone had to go looking for the bad kind.
Meanwhile, a version-four candidate had been technically ready to ship for five months. Nobody had said yes. Nobody had said no. It just sat in the queue, mentioned occasionally, decided on never.
Then Corvasta Bank happened. Their leadership offsite, eleven people, four time zones, and Truce double-booked the CFO's board call against the offsite's own opening session, a slot that would have opened at 2am for their Singapore office. An account manager caught it the night before and fixed it by hand. Corvasta's team was furious, and for about a week the contract genuinely looked shaky. It got smoothed over. Nobody in Thursday's meeting ever put it on the board, because that week's meeting opened, like every week's meeting, with a clean win: a three-person same-timezone request Truce had resolved in under a second.
Two weeks after that, Wisteria Duda, Loomcastle's VP of Product, sat in on Truce Review for the first time in months. She usually trusted the team enough to skip it. She watched the whole thirty minutes: a good demo, warm nods, a few jokes, a clean close. Afterward, in the hallway, she asked Perilune one plain question. "So what did we just decide about v4?"
Perilune didn't have an answer. Not a vague one, not a half one. Nothing.
We considered the obvious fix first: just tell Perilune to end every meeting by forcing a vote, whatever the room happened to be looking at that day. We rejected that. A decision forced on top of a highlight reel just becomes a rubber stamp on whatever demo got shown, ship it, because the last ten minutes felt good, not because anyone weighed the failure mode that actually mattered.
Here's the decision I'd take back instead, and it isn't Perilune's, and it isn't "make her hunt harder." It goes back to the week Truce Review started, when the team built the whole meeting's shape around one goal: keep a skeptical leadership team sold on continuing to fund a model that was new and unproven. Opening on the best case, with no decision required, was exactly the right design for that goal. It made complete sense while buy-in was the actual risk. Nobody ever redesigned the meeting once buy-in stopped being the question and an actual ship call started sitting on the table instead.
Run the same kind of story again, five months later, with Truce Review redesigned in between. New rule, first thing, every single Thursday: one named decision goes on the whiteboard before anyone opens a laptop. "Does version four ship to enterprise this cycle." Evidence gets pulled for and against, including a nine-person, three-timezone case that version four handles correctly and version three never did. The meeting doesn't end until someone writes down a call and a name next to it.
Same trigger, a version-four candidate sitting ready. This time the meeting closes with: ship v4 to 20 percent of enterprise accounts starting Monday, Perilune owns the rollout, revisit the numbers in three weeks. Written down. Owned. Dated.
That's the whole difference. One design hands the room a highlight reel. The other hands it a decision it can't leave without making.
What I'd tell my past self, the one who set that meeting up back when Truce was new and needed defending: a review meeting doesn't keep its own shape just because it started right. Somebody has to decide, on purpose, that a real question belongs on the table every single week, before the week a VP has to ask it for you.
Five letters, and the one Perilune had stopped asking
Not a trick to sound structured. FLIPS is what makes you notice that Wisteria's plain question was the whole audit, five months before anyone ran one on purpose.
The AI-specific failure worth naming plainly is silent segment-level miscalibration: Truce's aggregate accuracy genuinely improves while one real segment, complex multi-timezone conflicts, keeps failing about a third of the time, invisible behind a good-news average and a meeting that never pulls that segment on purpose. The guardrail is a target rate for that segment, checked against a labeled set of past complex cases, on a schedule, not a person's sense of whether the room "felt fine" that week. There's a real trade-off, accepted on purpose: naming a decision and hunting a hard case for it costs someone real prep time, something like thirty to forty minutes instead of grabbing whatever a Slack thread already praised, in exchange for a decision that's actually informed by where Truce fails, not just where it shines.
And if you want to be sure it really works, try it somewhere else
Same five letters, a completely different flip. This time nobody's chasing a highlight reel on purpose. The room that could say yes or no just isn't in the building anymore.
Greymantle Health runs Greytrace, a model that flags anomalies on chest X-rays for a second read before a radiologist signs off. Alderyth Ruzicka, the chief radiologist, owned Greytrace's weekly model review meeting when it launched. Two years in, she'd delegated running it to a rotating pair of second-year residents, freeing up her own Thursdays for something that felt more urgent.
F · Alderyth Ruzicka, chief radiologist, who owned Greytrace's model review at Greymantle Health when it launched.
L · She stopped requiring the meeting to name a real decision before it started, once early Greytrace versions were informational only and nothing needed a true go or hold.
I · The delegation flip, a different shape from Perilune's. Old setting: the senior radiologist runs the meeting and holds ship authority in the same room, in the same breath. New setting: rotating residents run the meeting, present flagged wins, the room applauds, but nobody present can actually authorize a version rollout, so the real call silently waits for someone senior who isn't there. Nothing in between: either the person naming the decision has the authority to make it, or the meeting produces agreement with nowhere to go.
P · Delegating the meeting's facilitation to residents without keeping the naming of the decision with someone who had authority to make one, because early on nothing in that meeting ever needed a real yes or no.
S · Alderyth keeps decision-naming with herself even as residents keep running the logistics. The meeting opens on "does Greytrace v2.3 roll out hospital-wide," and the residents' agenda now has to include a trauma-complex case pulled on purpose, not just the week's cleanest catch.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: a review meeting that opens on a demo instead of a decision will always let a real ship risk sit untouched, no matter how good the demo is.
Cost: no budget to hunt a fresh hard case every week. Reuse last quarter's worst real misses as a rotating pool, it costs nothing extra, it's just deciding to keep them instead of letting them vanish once the meeting moves on.
The model got better, for real: say Greytrace's benchmark jumps on the next release. Doesn't matter, maybe it matters more. A model that's mostly excellent is exactly the one nobody thinks to keep checking on its worst ten percent.
Where people run it wrong.
They treat a good demo as proof nothing needs deciding this week, instead of asking what decision is actually due.
They let "the room agreed" stand in for "someone with authority said yes," especially once a meeting gets delegated to whoever's free that day.
They wait for a bad outcome to force the question, when the whole point of naming a decision up front is that you don't need one to show up first.
How to use it live. When an interviewer asks how you'd run a model review meeting, ask yourself one thing before answering: if someone pulled last month's meeting notes right now, could they find a written decision with a name next to it? If the honest answer is "probably just some nice examples," that's the whole question, answered.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the hard case you pull is just a one-off, not a real pattern?" Response: that's exactly what the target rate and the eval set are for, checking one bad case against a labeled set of past complex cases before treating it as the whole story.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Working with ML engineers and researchers
- #1 How do you write a requirement for a team whose output is a probability distribution?
- #2 An engineer says the model cannot do that. What questions do you ask before accepting it?
- #3 Describe how you would run a planning session when effort estimates are genuinely unknowable.
- #4 What does a healthy PM-to-research relationship look like when research timelines are open-ended?
- #5 How do you keep a research team connected to user problems without constraining their exploration?
- #6 Your ML team wants three months to improve accuracy by two points. How do you evaluate that ask?