CaseAdvancedModel Fluency & the AI PM Role / Working with ML engineers and researchers / #18

Describe how to run a model review meeting that produces decisions rather than admiration.

FLIPS · a routine version upgrade to Truce, the scheduling-conflict resolver behind Loomcastle's calendar product

Loomcastle sells Truce, a model that untangles scheduling conflicts across everyone's calendars on its own, so an eleven-person meeting spread across four time zones gets a slot without twenty reply-all emails. Perilune Strade, a senior PM, has run Truce's weekly review meeting, Truce Review, for about fourteen months. This is the meeting where the VP of Product finally asked the one question nobody in the room could answer.

The direct answer
Open every review with one named decision, not a demo, something like "does version four ship to enterprise this cycle." Pull only the evidence that bears on that decision, good and bad, including a hard case chosen on purpose to test it. Close the meeting with an actual decision and a named owner, never just applause.
Do this, in order
  1. Open the meeting on a named decision, not a demo.Why: this is the actual reversal; skip it and the room defaults to whatever looks best that week, with nothing forcing a call.
  2. Bring a deliberately hard case, chosen on purpose, every single time.Why: a review built only on wins never sees where Truce actually breaks, so a real ship risk drifts for months with no evidence against it.
  3. Close every meeting with a written decision and a named owner.Why: a decision that only lives in the room's mood evaporates the moment everyone leaves; a decision with an owner and a date can be checked later.
  4. Track a target rate for the hard-case segment, not just the average.Why: a 98 percent aggregate accuracy hides a segment failing one in three, and an average never flags a ship risk on its own.
  5. Keep the demo-only meeting, but call it what it is.Why: sales and leadership do need a highlight reel sometimes; the problem was never demos existing, it was letting one quietly become the only review anyone had.
  6. Review the decision log itself on a schedule, even in weeks nothing's on fire.Why: a meeting can run for months producing zero decisions and still feel productive, because nobody's watching the log, only the room's mood.

How to answer this, stage by stage

Nobody is grading whether you can describe a tidy meeting agenda. They are grading whether you'd notice the moment a real review turns into a highlight reel, and catch it before a live ship call sits untouched for months.

1
Scope it to one product, one meeting, one person
Say it like this
"Let's ground this. Loomcastle sells Truce, a model that auto-resolves scheduling conflicts across calendars. Perilune Strade runs Truce's weekly review meeting. This is the meeting that quietly stopped deciding anything, for about five months, while a live ship call sat untouched."
Why this works
Naming the product, the model, and the meeting stops the answer from staying a vague statement about running good meetings.
2
Say your structure out loud
Say it like this
"I'll run this as FLIPS. Find the person whose habit is at risk. Locate what she used to do that worked. Identify the exact flip, the two-setting switch. Pinpoint the old decision that only made sense before. Show the replay with a different meeting design."
Why this works
Two sentences of structure buy you a plan before any story starts, and tell the interviewer you're not just improvising.
3
Reframe what's actually being tested
Say it like this
"This question isn't really asking me to describe a meeting agenda. It's asking whether I'll notice the moment a recurring review stops forcing a decision and starts just admiring good outputs, and whether I'd catch that before a real ship call sits untouched for months."
Why this works
Compresses the whole answer into one breath, before a single detail can bury it.
4
Give the one decision
Say it like this
"Here's what I'd actually do. Every review opens with one named decision on the board, something like 'does version four ship to enterprise this cycle.' Only evidence that bears on that decision gets airtime, including a hard case somebody had to go dig up on purpose. It closes with a decision and an owner, not applause."
Why this works
This is the direct answer, said plainly, before the story arrives to earn it.
5
Prove it with a compressed failure
Say it like this
"Here's what happens without it. Truce keeps getting better release over release, so good demos get easy to find. Perilune stops hunting for the hard case and just opens with whatever looks best. Five months and about twenty meetings later, the decision log has nothing in it, and a live v4 ship call is still sitting there untouched, while v3 is quietly failing one in three of its hardest cases in production."
Why this works
Four sentences carry a whole incident that a full retelling would take a page to earn.
6
Name the number you'd track
Say it like this
"I'd watch hard-case accuracy on its own, never folded into the average. Truce's overall number looks fine at 98 percent, but that's mostly simple two-to-four-person, same-timezone conflicts. Complex, multi-timezone cases sit at 68 percent, and that's the number a good-news meeting will never show you on its own."
Why this works
Shows you'd measure something the room's mood can't tell you.
7
Say what you'd leave alone
Say it like this
"I wouldn't touch the monthly Truce Showcase, the one sales runs for prospects. That meeting's whole job is to look good, and nobody's expecting a decision out of it. The problem was never demos existing. It was letting a demo meeting quietly become the only review meeting anyone had."
Why this works
Shows judgment instead of blanket suspicion of every demo.
8
Close on the one line
Say it like this
"So: a review meeting doesn't stay a decision-forum on its own just because it started as one. Somebody has to name a real decision on purpose, every single time, or the room just gets better and better at admiring whatever's already working."
Why this works
Leaves the interviewer with the actual decision, not just a well-told story about one meeting.

Let's learn

What happens when the meeting built to make a call on a model quietly turns into a meeting that just watches it do well?

Say a company sells calendar software. Inside it sits a model called Truce, which looks at everyone's calendars for a meeting request and works out a slot that actually fits, moving or splitting things where it has to. Every week, about 40,000 of these conflicts get resolved automatically across Loomcastle's customers. Truce's review meeting, Truce Review, started the same week the model launched: thirty minutes, every Thursday, in a small room. For its first five months, roughly twenty-two meetings, it opened with the worst miss since last time and closed with a written decision twenty times out of twenty-two.

Hand sketched two panel comparison titled The I step, in one picture. Left panel a gauge icon labeled The small move, caption Truce's accuracy climbs from 91 percent to 98 percent over 14 months, release by release. Right panel a question mark box icon in red-orange labeled The big snap, caption the meeting jumps from hunting a hard case every week to grabbing whatever looks best, no middle setting.
The model moved a little at a time. What the meeting did with its opening ten minutes moved all at once.

Now, fourteen months in, Truce is genuinely better than it was: overall accuracy has climbed from 91 percent to 98 percent, one small release at a time. But that number hides a second one. On simple conflicts, two to four people, one or two time zones, Truce is right 98 times out of 100. On complex conflicts, eight or more people, three or more time zones, it's right 68 times out of 100. It fails on roughly one in three of the cases it was actually bought to handle.

Truce v3 accuracy, by how hard the conflict is
100% 50% 0 98% Simple, 2 to 4 people 68% Complex, 8+ people, 3+ zones
Simple conflictsComplex conflicts
Complex conflicts are only 12 percent of Truce's weekly volume, but they generate 55 percent of the support tickets that reach a human.

Here's the turn. Those failed complex cases are not the real problem. The real problem is what the room does with its opening ten minutes once good wins get easy to find.

We did not lose a decision. We lost the habit of asking for one.

At its worst, a genuinely live risk sits untouched behind a room full of applause. Corvasta Bank, one of Loomcastle's larger accounts, ran an eleven-person leadership offsite across four time zones. Truce double-booked their CFO's board call against the offsite itself, a slot that would have started at 2am in Singapore for one attendee. Someone caught it the day before and fixed it by hand. Corvasta nearly walked. Nobody in Thursday's meeting connected that near miss to the version-four decision sitting in the queue, because the meeting that week opened, as it always did, with a clean demo of Truce handling a three-person same-timezone request.

The choice I would take back When Truce Review started, the team built its whole shape around one goal: keep a skeptical leadership team sold on funding a model that was new and unproven. Opening with the best case, no decision required, worked perfectly for that goal. It made complete sense while buy-in was the actual risk. Nobody ever redesigned the meeting once buy-in stopped being the question and a real ship call started sitting on the table instead.
Knowledge spark: what's a target rate? A cut-off you set on purpose for one specific slice of cases, checked on a schedule, not the model's overall score. "Complex conflicts fail under 10 percent of the time, checked weekly" is a target rate. "98 percent accurate" without saying on what is just a number that hides its worst segment.

What I would leave alone: the monthly Truce Showcase, a separate meeting sales runs for prospects, doesn't need this fix at all. Its whole job is to look good, and nobody in that room expects a decision to come out of it. The split only earns its keep on the meeting whose real job is deciding something.

The lesson: a review meeting doesn't hold its shape on its own just because it started as a decision-forum. Somebody has to name a real decision, on purpose, every single time, and that costs something real: hunting a genuine hard case before each meeting takes someone thirty or forty minutes of real digging, instead of grabbing whatever a Slack thread already praised. That's the trade worth accepting on purpose, slower prep, for a decision that's actually informed by where Truce fails.

Now here is the same thing as a story

The short version above is what you'd actually say in the room. Read this one when you want to feel exactly what five quiet months cost, meeting by meeting.

Perilune Strade can read a Truce failure log and tell, inside a minute, whether it's a real pattern or a fluke. She's run Truce Review since the model first shipped, fourteen months now, the same thirty-minute slot every Thursday.

For the first five months, that habit was simple and it worked. Every Thursday she'd go looking, on purpose, for the worst miss Truce had made since the last meeting, and she'd open with it. Nobody enjoyed those first ten minutes. But the room argued, decided something, and moved on. Twenty of the first twenty-two meetings closed with an actual line in the decision doc: ship this, hold that, kill this idea.

Then Truce kept getting better. Not in one leap, in the ordinary way software gets better: a release here, a fix there, 91 percent climbing toward 98 over more than a year. And good wins started arriving on their own, without her having to dig for anything. A clean save on a nine-person cross-office scheduling mess. A customer email thanking the team by name. Opening with one of those worked just as well to fill the room's attention, and it took a tenth of the effort.

So, without ever deciding to, she stopped hunting for the hard case.

Hand sketched horizontal timeline titled Perilune's Thursday, thinning into a demo. Four milestones: The good months, caption opens on the worst miss, closes with a decision. Habit thinning, caption wins arrive on their own, easier to grab, this milestone emphasized. No one decided to stop, caption 5 months, 20 meetings, nothing logged. Wisteria's question, caption what did we just decide about v4.
Nobody announced the change. It thinned out over months, the way most habits actually go.

By month nine, that was just how Thursdays worked. Whoever had the best recent save got the floor first, the room nodded along, and the meeting closed on time with nothing written down. Nobody voted on this. It was simply the easiest way to fill thirty minutes once good news kept arriving before anyone had to go looking for the bad kind.

Meanwhile, a version-four candidate had been technically ready to ship for five months. Nobody had said yes. Nobody had said no. It just sat in the queue, mentioned occasionally, decided on never.

Then Corvasta Bank happened. Their leadership offsite, eleven people, four time zones, and Truce double-booked the CFO's board call against the offsite's own opening session, a slot that would have opened at 2am for their Singapore office. An account manager caught it the night before and fixed it by hand. Corvasta's team was furious, and for about a week the contract genuinely looked shaky. It got smoothed over. Nobody in Thursday's meeting ever put it on the board, because that week's meeting opened, like every week's meeting, with a clean win: a three-person same-timezone request Truce had resolved in under a second.

Two weeks after that, Wisteria Duda, Loomcastle's VP of Product, sat in on Truce Review for the first time in months. She usually trusted the team enough to skip it. She watched the whole thirty minutes: a good demo, warm nods, a few jokes, a clean close. Afterward, in the hallway, she asked Perilune one plain question. "So what did we just decide about v4?"

Perilune didn't have an answer. Not a vague one, not a half one. Nothing.

We considered the obvious fix first: just tell Perilune to end every meeting by forcing a vote, whatever the room happened to be looking at that day. We rejected that. A decision forced on top of a highlight reel just becomes a rubber stamp on whatever demo got shown, ship it, because the last ten minutes felt good, not because anyone weighed the failure mode that actually mattered.

Here's the decision I'd take back instead, and it isn't Perilune's, and it isn't "make her hunt harder." It goes back to the week Truce Review started, when the team built the whole meeting's shape around one goal: keep a skeptical leadership team sold on continuing to fund a model that was new and unproven. Opening on the best case, with no decision required, was exactly the right design for that goal. It made complete sense while buy-in was the actual risk. Nobody ever redesigned the meeting once buy-in stopped being the question and an actual ship call started sitting on the table instead.

Run the same kind of story again, five months later, with Truce Review redesigned in between. New rule, first thing, every single Thursday: one named decision goes on the whiteboard before anyone opens a laptop. "Does version four ship to enterprise this cycle." Evidence gets pulled for and against, including a nine-person, three-timezone case that version four handles correctly and version three never did. The meeting doesn't end until someone writes down a call and a name next to it.

Same trigger, a version-four candidate sitting ready. This time the meeting closes with: ship v4 to 20 percent of enterprise accounts starting Monday, Perilune owns the rollout, revisit the numbers in three weeks. Written down. Owned. Dated.

That's the whole difference. One design hands the room a highlight reel. The other hands it a decision it can't leave without making.

What I'd tell my past self, the one who set that meeting up back when Truce was new and needed defending: a review meeting doesn't keep its own shape just because it started right. Somebody has to decide, on purpose, that a real question belongs on the table every single week, before the week a VP has to ask it for you.

Five letters, and the one Perilune had stopped asking

Not a trick to sound structured. FLIPS is what makes you notice that Wisteria's plain question was the whole audit, five months before anyone ran one on purpose.

Hand sketched numbered list titled FLIPS, five questions before the next meeting. Five rows: F, find the person, whose meeting habit is this. L, locate the habit, what did she stop hunting for. I, identify the flip, what verb snaps, this row in red-orange. P, pinpoint the old decision, what made sense before. S, show the replay, same Thursday, new design.
Four setup and payoff letters, and one hard question sitting in the middle of all of them.
FFind the person. Whose habit is this?
Perilune Strade, the senior PM who has run Truce Review, Loomcastle's weekly model review meeting, for about fourteen months.
The flip belongs to whoever actually shapes what the room looks at first, not whoever happens to sit in on it that week.
LLocate the habit. What did she stop doing?
She stopped deliberately hunting for a hard case to open the meeting with, once genuine wins started arriving on their own as Truce got better, month after month.
That habit cost real digging while Truce was rough. It quietly stopped mattering the moment good news got easy to find.
IIdentify the flip. What verb snaps?
Old setting: the meeting opens on a hard case someone went and dug up on purpose, forcing the room to look at where Truce actually struggles before anyone gets to feel good. New setting: the meeting opens on whichever recent case looks best, chosen for how good it looks, and closes with applause and nothing written down. Nothing in between: there is no version of the meeting that opens on a highlight reel and still ends in a decision, because nothing on the agenda ever names one.
This is the answer to the question in one line. A review meeting doesn't keep forcing decisions just because it started as one; the room quietly optimizes for feeling good the moment good outputs get cheap.
PPinpoint the old decision. Which choice made sense before?
Building Truce Review's entire shape, at launch, around keeping skeptical leadership sold on funding the model, opening on the best case with no named decision required.
"Force a vote at the end, whatever the room saw that day" would be a new dial bolted onto the same broken input. Naming the decision at the start, before any evidence gets shown, is the reversal actually taken back.
SShow the replay. Same Thursday, better ending?
Same version-four candidate, ready and waiting. The meeting opens on "does v4 ship to enterprise this cycle," pulls a hard case v4 gets right that v3 never did, and closes with a written call.
The replay ends in a count: ship to 20 percent of enterprise accounts, starting Monday, owned by Perilune, revisited in three weeks. Not "much better," a date and a name.
Hand sketched two panel comparison titled Truce Review, written decisions logged. Left panel a document icon labeled First 20 meetings, caption 20 of 22 close with a written decision on the board. Right panel a document icon in red-orange labeled Last 20 meetings, caption 0 close with a decision, v4's ship call waits 5 months.
Same room, same thirty minutes, same person running it. What changed is what the meeting was actually for.
Hand sketched full page metaphor titled What the room assumed, and what was true. Left panel a gauge icon labeled DIAL, caption what we assumed, a review meeting stays useful a little at a time. Right panel a switch icon in red-orange labeled SWITCH, caption what was true, it opens on a decision, or it opens on a demo, nothing between.
The whole answer, in one picture. Nobody designed a dial. Everybody got a switch, and for months nobody had to throw it themselves.
Share of Truce Review meetings closing with a written decision, month by month
100% 50% 0 95% 55% 0%, Wisteria asks why Mo. 1 Mo. 7 Mo. 13
Meetings closing with a written decision, that month
Nobody had a chart watching this on purpose. It took a VP sitting in for one Thursday to make the shape of the drop impossible to ignore.

The AI-specific failure worth naming plainly is silent segment-level miscalibration: Truce's aggregate accuracy genuinely improves while one real segment, complex multi-timezone conflicts, keeps failing about a third of the time, invisible behind a good-news average and a meeting that never pulls that segment on purpose. The guardrail is a target rate for that segment, checked against a labeled set of past complex cases, on a schedule, not a person's sense of whether the room "felt fine" that week. There's a real trade-off, accepted on purpose: naming a decision and hunting a hard case for it costs someone real prep time, something like thirty to forty minutes instead of grabbing whatever a Slack thread already praised, in exchange for a decision that's actually informed by where Truce fails, not just where it shines.

And if you want to be sure it really works, try it somewhere else

Same five letters, a completely different flip. This time nobody's chasing a highlight reel on purpose. The room that could say yes or no just isn't in the building anymore.

Greymantle Health runs Greytrace, a model that flags anomalies on chest X-rays for a second read before a radiologist signs off. Alderyth Ruzicka, the chief radiologist, owned Greytrace's weekly model review meeting when it launched. Two years in, she'd delegated running it to a rotating pair of second-year residents, freeing up her own Thursdays for something that felt more urgent.

Hand sketched decision tree titled Where a Greytrace case goes now. Root box: a flagged X-ray reaches review. Four branches: clean catch, easy to present, leading to residents present it, room applauds. Hard trauma case, messy, leading to no one senior in the room to decide. After the fix, named decision first, leading to Alderyth names the call, hard case included. Rollout question, leading to Greytrace v2.3 gets a real go or hold.
Same five questions, a completely different way the flip hides. This time the person with authority to decide simply wasn't in the room.

F · Alderyth Ruzicka, chief radiologist, who owned Greytrace's model review at Greymantle Health when it launched.
L · She stopped requiring the meeting to name a real decision before it started, once early Greytrace versions were informational only and nothing needed a true go or hold.
I · The delegation flip, a different shape from Perilune's. Old setting: the senior radiologist runs the meeting and holds ship authority in the same room, in the same breath. New setting: rotating residents run the meeting, present flagged wins, the room applauds, but nobody present can actually authorize a version rollout, so the real call silently waits for someone senior who isn't there. Nothing in between: either the person naming the decision has the authority to make it, or the meeting produces agreement with nowhere to go.
P · Delegating the meeting's facilitation to residents without keeping the naming of the decision with someone who had authority to make one, because early on nothing in that meeting ever needed a real yes or no.
S · Alderyth keeps decision-naming with herself even as residents keep running the logistics. The meeting opens on "does Greytrace v2.3 roll out hospital-wide," and the residents' agenda now has to include a trauma-complex case pulled on purpose, not just the week's cleanest catch.

What finally surfaced it Greytrace v2 missed a subtle rib fracture on a complex trauma case, the kind of messy, multi-injury film the review meeting had never once pulled in six months of praising clean catches. Treatment wasn't delayed for long, but it was close enough that the near miss reached Alderyth directly, not through the meeting that was supposed to be watching for exactly this.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: a review meeting that opens on a demo instead of a decision will always let a real ship risk sit untouched, no matter how good the demo is.
Cost: no budget to hunt a fresh hard case every week. Reuse last quarter's worst real misses as a rotating pool, it costs nothing extra, it's just deciding to keep them instead of letting them vanish once the meeting moves on.
The model got better, for real: say Greytrace's benchmark jumps on the next release. Doesn't matter, maybe it matters more. A model that's mostly excellent is exactly the one nobody thinks to keep checking on its worst ten percent.

Where people run it wrong.
They treat a good demo as proof nothing needs deciding this week, instead of asking what decision is actually due.
They let "the room agreed" stand in for "someone with authority said yes," especially once a meeting gets delegated to whoever's free that day.
They wait for a bad outcome to force the question, when the whole point of naming a decision up front is that you don't need one to show up first.

How to use it live. When an interviewer asks how you'd run a model review meeting, ask yourself one thing before answering: if someone pulled last month's meeting notes right now, could they find a written decision with a name next to it? If the honest answer is "probably just some nice examples," that's the whole question, answered.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
An over-trust flip. As Truce got better release over release, the room checked less and less, until it stopped forcing any real decision at all. Rare failures kept shipping unseen behind a run of good demos.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Perilune Strade, senior PM who has run Truce Review, Loomcastle's weekly model review meeting, for about fourteen months.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
She stopped deliberately hunting for a hard case to open the meeting with, once genuine wins started arriving on their own as Truce got better, month after month.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Old: the meeting opens on a hard case chosen on purpose to force a decision. New: it opens on whichever recent case looks best, and closes with applause and nothing written down. No version opens on a highlight reel and still ends in a decision.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Building Truce Review's whole shape, at launch, around keeping skeptical leadership sold on funding the model, opening on the best case with no named decision required. It made sense while buy-in was the real risk.
6 · THE NUMBER
Fill in the blank: Truce's overall accuracy sits at ___%, but on complex, multi-timezone conflicts it's only ___%.
Tap to flip
ANSWER
98% overall, 68% on complex cases. Complex cases are about 12% of weekly volume but 55% of the support tickets that reach a human.
7 · THE REPLAY
Same bad day, new design, what changes?
Tap to flip
ANSWER
Same version-four candidate, ready and waiting. The meeting opens on "does v4 ship to enterprise this cycle," pulls a hard case v4 gets right that v3 didn't, and closes: ship to 20% of enterprise accounts starting Monday, Perilune owns it, revisit in three weeks.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs FLIPS again on a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Greytrace, Greymantle Health's second-read X-ray model. The delegation flip: review got handed to rotating residents who could run the meeting but not authorize a version rollout, so decisions silently waited for someone senior who was no longer in the room.

Check yourself Score: 0 / 0

Multiple choice
1. What was the flip in Perilune's story, and what were its two settings?
  • A. The meeting stopped happening at all once Truce got reliable enough.
  • B. The meeting flipped from opening on a hard case chosen to force a decision, to opening on whichever recent win looked best, with nothing decided.
  • C. Truce itself switched from resolving conflicts automatically to only flagging them for a human.
  • D. Perilune started attending the meeting only every other week instead of weekly.
Show hint
Look at the I step in the framework recap.
Show answer
B. A flip is a two-setting switch in what the person does, not in whether the product or the meeting keeps running.
Fill in the blank
2. Truce's overall accuracy sits at ___%. On complex, multi-timezone conflicts, that number drops to ___%.
Show hint
Check the bar chart in "Let's learn."
Show answer
98%; 68%. The aggregate number kept climbing while a specific, riskier segment stayed stuck near one-in-three wrong.
True or false
3. True or false: if the review meeting had opened on a demo but still forced a vote at the end of every meeting, that alone would have fixed the real problem.
  • True
  • False
Show hint
Look at the rejected alternative in the story section.
Show answer
False. A forced vote on top of a highlight reel just rubber-stamps whatever demo got shown, since the evidence in the room was never chosen to test the actual decision.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Building Truce Review's shape around keeping skeptical leadership sold on funding the model, opening on the best case with no decision required. It made sense while leadership buy-in, not a real ship call, was the actual risk on the table.
Short answer, apply it yourself
5. Think of a recurring review or status meeting you run or sit in on, one that watches something get better over time. What decision has it been quietly avoiding?
Show hint
Look for a meeting that always feels productive but rarely produces something you could point to later.
Show answer
Model answer: A weekly "product health" sync that always shows the metrics trending up, but has never once decided whether to deprecate the one feature nobody uses. Naming that decision on the agenda, with the usage numbers pulled specifically to test it, would force the call the good metrics keep letting everyone avoid.
Multiple choice
6. Why couldn't Perilune have just "brought a hard case a bit more often" instead of redesigning what the meeting opens on?
  • A. Because hard cases were too rare to find at any frequency.
  • B. Because the flip has no middle setting: either the room's evidence is chosen to test a named decision, or it's chosen because it looks good, with nothing in between.
  • C. Because Truce's model architecture made hard cases impossible to detect.
  • D. Because Wisteria Duda required a specific meeting format by policy.
Show hint
This is the "flip versus dial" mistake the taxonomy warns about most often.
Show answer
B. "A bit more often" is a dial. A real flip means there's no version of the meeting that opens on a highlight reel and still reliably ends in a decision, because nothing on the agenda ever named one.
Before you close the answer
Why this works
Tests whether you'll design a review meeting around forcing evidence-based decisions, or keep running a ritual that feels productive while a live ship call quietly stalls. Most candidates describe better meeting hygiene; fewer notice that a room can look busy for months while the decision log stays empty.
Follow-up traps
"Isn't naming a decision every single week just busywork if most weeks nothing's actually ready to ship?" Response: a decision doesn't have to be "ship." Hold and kill are decisions too. The point is closing every meeting on one of the three, instead of on nothing.

"What if the hard case you pull is just a one-off, not a real pattern?" Response: that's exactly what the target rate and the eval set are for, checking one bad case against a labeled set of past complex cases before treating it as the whole story.
If pressed
The complex-case eval set gets rebuilt every quarter from a stratified sample of that quarter's actual complex-conflict tickets, not a fixed set frozen at launch, so it can't quietly go stale the same way the demo-only meeting did.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more