InterviewIntermediateModel Fluency & the AI PM Role / Working with ML engineers and researchers / #20

How would you build trust with an ML team that has been burned by a previous PM?

LEAD · why silence is the tell a burned team gives first, tested on Havistock's resume screener Riffle

Riffle is Havistock's screening tool. It reads a stack of resumes, matches them against a job's real requirements, and ranks who an ATS should move forward first. Ephrasie Culdrew took over its roadmap six weeks after the last PM shipped a change nobody on the ML team trusted anymore. Milovan Beckthorn leads that team. He flagged the problem three times and got overridden three times. Ephrasie's first three syncs with him went fine. That was the part that should have worried her.

The direct answer
Make one real, visible decision that follows your ML team's eval evidence, even when it costs you a deadline, and do it early. Then track how often the team starts raising pushback and bad news on its own. A rising number is proof trust is coming back. An open door and a kind word are not, no matter how much you mean them.
Do this, in order
  1. Make a real, visible decision based on the team's eval evidence, early, even if it costs a deadline.Why: this is the one thing a burned team actually checks. Words don't move it. A reversed decision does.
  2. Track how often the team raises pushback or bad news on its own, as your real trust metric.Why: it moves in weeks. A survey or a gut feeling moves in months, if it moves at all.
  3. Write down every time a decision goes against the team's eval-based pushback, and make the record visible above you.Why: a hidden pattern of overrides is exactly what let the last PM do it three times unnoticed.
  4. Never treat an open-door line or a one-time ask for candor as the fix by itself.Why: it changes no actual decision, and a team that's been burned once can feel the difference immediately.
  5. Watch the slow signals too, a trust survey, attrition, but don't steer by them.Why: by the time a quarterly number moves, you've either already fixed the problem or already lost the people.
  6. Leave alone any part of the team that the last decision never actually burned.Why: making them re-earn trust they never lost reads as suspicion, not leadership.

How to answer this, stage by stage

Nobody is grading whether you can be nice to a team that got hurt. They're grading whether you can name the actual thing trust moves on, instead of a list of gestures.

1
Ground it in one team, one product
Say it like this
"Let's ground this in one real team. Riffle is Havistock's resume screening tool. Ephrasie Culdrew just took over its roadmap. Milovan Beckthorn leads the ML team that got overridden three times by the last PM. I'll answer using them, not the idea of a burned team in general."
Why this works
Naming a real team and a real history stops "build trust" from turning into a vague pep talk.
2
Say your structure out loud
Say it like this
"I'll run this as LEAD. Link, what trust should actually be measured against. Early signal, the number that shows it's coming back weeks before anyone feels it. Abuse, how a trust-building move gets faked. Decision, what I'd actually do, and what I'd track after."
Why this works
Two seconds of structure tells the interviewer you have a method, not just a feeling about being a good manager.
3
Name what trust actually looks like, before you plan anything
Say it like this
"Here's the part most answers skip. Trust isn't a feeling you check for. It's whether the team tells you something's wrong before you ask. A team that only ever reports good news isn't calm, it's quiet on purpose. That's the real signal, and it's the one I'd watch."
Why this works
This reframes the question before any plan gets named, which is what separates a real answer from a list of gestures.
4
State the one decision you'd actually make
Say it like this
"So here's what I'd actually do. The first time the team's eval numbers say don't ship, I ship late instead. Not a speech. An actual launch date that moves, in public, because of their evidence. That's the whole test, and there's no substitute for passing it."
Why this works
This is the direct answer, said out loud, with no hedge in it.
5
Prove it with the real reversal, numbers first
Say it like this
"Here's what actually happened at Havistock. Nine days before a client launch, Milovan's team ran our 640-resume eval set against a new scoring model and found recall on flagged resumes had dropped from 88 to 65 percent, the same shape of gap the last PM had shipped through three times. I delayed the launch nine days, fixed the threshold, and told the client and the team the real reason. Twelve weeks later, the number of times the team raised something unprompted had gone from zero to seven a week."
Why this works
Two real numbers, weeks apart, beat any amount of talk about being a trustworthy manager.
6
Name the fake version of this move, then close
Say it like this
"In my first week, before I had a real decision to make, I tried the easy version, an open door, come tell me anything. Nobody did, and they were right not to, since I hadn't proven anything yet. So: trust comes back from one real decision made on their evidence, not from an invitation. Watch how often they push back after that. That number is the actual answer to whether it's working."
Why this works
Naming the version that doesn't work, and admitting you tried it, is what proves this is judgment and not a script.

Let's learn

Here is what happens when a team stops telling you the truth, and it looks exactly like everything going fine.

Riffle is the tool Havistock, an applicant tracking company, sells to recruiters. It reads a stack of resumes against a job's real requirements and ranks who should move forward first. Milovan Beckthorn's team builds the model underneath it.

For six weeks last spring, that team raised a problem three separate times. Their own eval set, 640 real resumes with a recruiter's own answer already known for each one, showed that a new scoring change was quietly punishing resumes with a real employment gap and a non-standard job title. Reservist. Caregiver. Career break. Recall on that group of resumes fell from 89 percent to 62 percent. The team said so in writing, three times. The PM at the time, Adelric Verlach, shipped it anyway, to hit a big client's renewal deadline.

Hand sketched timeline titled Three flags, one client audit, with four milestones: Flag one, gap recall drops, week 2. Flag two, same drop again, week 4. Flag three, overridden a third time, this milestone emphasized. Client audit, the complaint lands, month four.
Three warnings, in writing, with real numbers attached. Every one of them changed nothing.
Knowledge spark: what is recall, and what is an eval set? An eval set is a stack of real cases with a real answer already known for each one, so a model's guess can be checked against something true. Recall is how much of the real "yes" cases a model actually catches. High recall means it rarely misses one that should have gotten through.

Two months after launch, the client ran its own audit. Candidates with a real employment gap were advancing 40 percent less often than similar candidates without one. That became a formal complaint, then a remediation project, then an apology nobody at Havistock wanted to write. Adelric moved off the team a few weeks later.

Ephrasie Culdrew took the roadmap six weeks after that. Her first three syncs with Milovan's group went fine. Status only. Nothing flagged. Green across the board.

In her first week, before she understood what quiet actually meant here, she tried the easy fix. An open door. "Come tell me anything, anytime." Nobody did. They were being perfectly polite, and they told her nothing real, because nothing had happened yet to prove her door meant something different from the last one.

Hand sketched icon list titled Which one actually counts as trust rebuilt, with three rows: Says my door is always open, in a warning color. Asks for candor once, at an all-hands, in a warning color. Changes one real decision, in public, in a healthy green color.
Two of these cost nothing to say. Only one of them a burned team can actually check.

Here is the turn. Three clean syncs were not good news. They were the same shape as the six weeks before the audit, a team saying everything is fine right up until it very much was not. The extra silence wasn't the real problem either. The real problem was that Ephrasie had no way to tell a team that trusted her from a team that had simply learned not to bother.

Hand sketched comparison diagram titled What trust actually leaves behind. Left panel, a question mark icon labeled Only good news, caption status reports, nothing flagged, week after week. Right panel, a document icon labeled Pushback, out loud, caption a real number, raised before anyone asked for it.
Both looked calm on a status update. Only one of them was actually calm.
Three clean syncs were not the sound of trust. They were the sound of a team that had already learned what happens when it speaks up.

Week four, Milovan's team ran the same 640-resume eval set against a new feature, a title-match confidence score, nine days before it was due to ship to a new client. Recall on the exact same flagged group had fallen from 88 percent to 65. Milovan raised it the smallest way possible, one line in a chat message: "Might be noise, wanted you to see it before Thursday."

For about an hour, Ephrasie wanted to ship anyway. The deadline was real, the client was big, and 65 percent still sounded like a working model. That is the exact pull that got Adelric three times.

She pulled the eval breakdown herself instead. The drop was real, and it repeated on a second sample. Someone on the team floated a faster fix: route the flagged resumes to a person for manual review instead of retraining the score. She turned it down. A manual gate hides the same bias behind a curtain. It does nothing for every gap resume the gate happens to miss, and it leaves the ranking itself untouched for the next feature built on top of it.

So she delayed the launch nine days, fixed the threshold, and told the client the real reason instead of a vague one. Then she told Milovan, in the same channel where he'd raised it, that the delay was because of his numbers, not despite the deadline.

Unprompted pushback raised by Milovan's team, per week
8 4 0 0 week 4: the flag 7 Week 1 Week 4 Week 12
Unprompted pushback or bad news, per week
Zero for three weeks straight. The flag that changed the launch date sits right there in week four. By week twelve the team was raising something on its own seven times a week.

Something moved that a survey never would have caught in time.

"I can raise a concern without it costing me," quarterly survey score
100% 50% 0% 34% Last, under previous PM 71% Two quarters after the delay
BeforeAfter
The survey only runs once a quarter. It didn't confirm what the weekly pushback count had already been saying for two more quarters.
Hand sketched comparison diagram titled Which clock rings first. Left panel, a gauge icon labeled Weekly pushback count, caption moves inside four weeks. Right panel, a gauge icon labeled Quarterly trust survey, caption moves after two quarters, if ever.
One of these numbers is ready the same week a real decision happens. The other one waits for a survey to go out.

What it costs at its worst, if a PM can't tell the difference between a quiet team and a trusting one: the next bad decision looks exactly like the last one, right up until the same kind of audit finds it.

The choice I would take back Long before any of this, nobody at Havistock kept a shared record of the times a PM shipped against the ML team's own eval evidence. That made sense when it was rare enough that nobody thought to write it down. It stopped making sense the moment it happened three times to the same team, and nobody above Adelric could see the pattern, because each override lived inside its own private conversation.
Hand sketched labeled parts diagram titled What the override log actually holds. A central document icon labeled Override Log, with four labeled callouts: What the eval set showed, What the PM decided, Why, in the PM's own words, Who else can see it.
Four parts, every time a decision runs against the team's own evidence. Leave one out and it goes back to being three unconnected conversations.

What I would leave alone: Havistock's data engineering team, the group that keeps resumes flowing into Riffle's pipeline, never dealt with Adelric's overrides directly. Running them through the same "prove yourself" ritual would cost time, and it would read as suspicion aimed at people who never did anything to lose it.

The lesson: a team that stops arguing with you isn't necessarily convinced. Sometimes it's just tired of being right and ignored. The fastest way to tell the difference isn't a survey. It's whether they say something unprompted the next time it actually matters.

Now here is the same thing as a story

The short version above is what you'd actually say in an interview. Read this one for the six weeks it took before anyone could tell whether Milovan's silence meant calm or damage.

Milovan Beckthorn has led Havistock's ML team for four years. Ask anyone on it and they'll tell you the same thing: he doesn't raise a flag unless the eval set backs it up, and when he does, he's usually right.

Under Adelric, being right three times in six weeks changed nothing. Milovan's team built the 640-resume eval set themselves, over a weekend, specifically because they suspected the new title-consistency scoring change was quietly hurting resumes with an employment gap. Week two, recall on that group came back at 71 percent, down from a baseline of 89. They wrote it up and sent it to Adelric. He said the sample was probably small and asked them to keep an eye on it.

Week four, the same eval set, run fresh: 65 percent. Same warning, more detail this time. Adelric held an actual meeting about it. He said the right things. "I want us to be a team that surfaces this stuff." He asked good questions. Then he asked whether it could wait until after the client renewal.

Week six, the number held at 62 percent. Adelric shipped the feature to the client's renewal deadline that Friday, with a note in the changelog that called it "a minor scoring refinement."

Two months later, the client's own compliance team ran a routine audit. Candidates with a documented employment gap were advancing 40 percent less often than similar candidates without one. Their next email to Havistock had the word "discriminatory" in the subject line.

We did not lose a client for six weeks. We lost a team's belief that speaking up ever changed anything.

I want to say the problem was the scoring change itself. It wasn't, not really. The scoring change was a mistake anyone could make. The real damage was that Milovan's team told the truth three times, in writing, with real numbers, and watched it not matter three times. By the time Ephrasie showed up, they weren't hiding anything. They just weren't going to bother saying it out loud again until they had a reason to believe it would land differently.

Ephrasie's first three syncs with the team went fine. Status updates only. "On track." "Green." Nothing raised, nothing flagged, three weeks running.

Hand sketched timeline titled Ephrasie's first three syncs, with four milestones: Sync one, status only, all green. Sync two, status only, still green. Sync three, one hedge, walked back fast. She notices, this milestone emphasized, caption the quiet itself is the tell.
Nobody decided to hide anything on any single Tuesday. The team had simply already learned what speaking up used to cost.

For about a week, she read that as a good sign. Then she remembered what she'd been told in her own handoff conversation about why Adelric left, and the quiet started to look like something else.

Her first move was the easy one. She called a session, said her door was always open, meant every word of it. Milovan's team thanked her and told her nothing real. Not because they doubted her sincerity. Because sincerity had never been the thing that was missing.

The real test showed up in week four, and it wasn't dramatic. A new client wanted the title-match confidence score fast-tracked, nine days out. Milovan ran the same 640-resume eval set his team had built two years earlier, the one that had caught the last problem and been ignored. Recall on the same flagged group, gap resumes with non-standard titles: 65 percent, down from 88.

He almost didn't send it. He told her later he'd sat on the message for twenty minutes, half convinced it would go exactly the way it always had. Then he wrote one line: "Might be noise, wanted you to see it before Thursday." The hedge was doing real work. It was a man giving his new PM an easy way to ignore him, because that's what the last nine months had taught him to expect.

Ephrasie read it at her desk with the client contract open in the next tab. For about an hour, she thought about shipping anyway. Sixty-five percent still sounded like a working number, the deadline was real, and nobody would have blamed her for a decision made under someone else's timeline.

Then she pulled the breakdown herself, instead of taking Milovan's word or her own instinct. The drop was real. It held on a second sample. Someone suggested routing the flagged resumes to manual review instead, a faster fix that would still hit the deadline. She almost took it. It would have worked, technically. It also would have hidden the same bias behind a curtain and done nothing for the next feature built on the same score.

She delayed the launch nine days. She told the client why, plainly, no soft language. Then she wrote back to Milovan in the same channel: "Delaying nine days. Your numbers are why. Thank you for sending it even hedged."

Adelric had asked for the truth and filed it away. Ephrasie asked for the truth and let it move a real date.

The decision she'd take back sits further back than her own six weeks. It sits in whatever meeting, years earlier, decided that a PM's call to override the ML team's own evidence didn't need to be written down anywhere two people could read it later. That gap is what let Adelric do the same thing three times and have it look, from above, like three separate ordinary judgment calls instead of one pattern.

Ephrasie built the fix the following week. An override log. Any time a product decision runs against the ML team's own eval-based pushback, whichever way it goes, the evidence, the decision, and the actual reasoning get written down in one place her own boss can read. Not a private chat thread. A record.

Twelve weeks after the delay, Milovan's team was raising something unprompted about twice a week, then three times, climbing toward seven by week twelve. The quarterly survey question, "I can raise a concern about a product decision without it costing me," had sat at 34 percent the last time anyone measured it under Adelric. Two quarters after the delay, it read 71 percent. The number that would have told her the truth fastest moved in weeks. The number everyone actually tracks moved in quarters, and only after the faster one had already said so.

What I'd tell myself, back in that first week with the open door and the good intentions: an invitation to speak up is not the same thing as proof that speaking up works. The team already knew that. I was the only one who needed convincing.

LEAD, so a burned team gets proof and not a promise

Not a way to prove Adelric was a bad person. LEAD is what forces you to name the actual signal trust moves on, and to catch a fake version of rebuilding it before a team quietly checks out for good.

LLink. The business outcome that actually matters.
Not whether the team says nice things about you in a survey. What actually matters is whether Havistock ships fewer decisions like Adelric's three overrides, and the only way to know that is whether the ML team tells you when something's wrong before you go looking for it.
Riffle's real trust problem was never Ephrasie's likability. It was whether Milovan's team would ever again spend a weekend building an eval set nobody was going to read.
EEarly signal. The thing that moves weeks before the outcome does.
Here the early signal is the plainest number in this whole answer: how many times a week the team raises something nobody asked them to raise. A quarterly survey needs a quarter to move. A weekly pushback count moves the same week a real decision does.
It went from zero for three weeks straight to seven a week by week twelve. The survey question didn't confirm any of that for two more quarters.
AAbuse. How a trust-building move gets faked.
Two ways, both real at Havistock. First, Adelric held his own listening session after the first override and said all the right words, then overrode the same evidence two more times anyway. Second, Ephrasie's own open door in week one, sincere and completely empty, because sincerity with no decision behind it changes nothing a burned team can feel.
"My door is always open" and "I want us to be a team that surfaces this" sound identical to a real reversal, right up until the next deadline arrives and only one of them survives it.
DDecision. What you'd actually do differently, at each threshold.
Make the first real decision follow the team's eval evidence, in public, even when it's expensive. Build a record of every time a decision goes against that evidence, so a pattern can be seen from above instead of hiding inside separate private calls. Track pushback frequency and bad news, not a feeling, as the number that says whether trust is actually returning.
Nine days of delay and one written log turned three unconnected overrides into a pattern nobody could miss twice.

The recap, one line per letter: link trust to whether the team tells you something's wrong before you ask, not to whether they like you. The early signal is the weekly pushback count, because it moves the same week a real decision does and a quarterly survey does not. Name both abuses plainly, a listening session with no decision behind it, and a one-time ask for candor that gets forgotten the moment it's inconvenient. And the decision is what makes it real: one visible reversal, a written record above you, and a number you track instead of a feeling you hope for.

Two things worth saying outright, since the real judgment sits here. Ephrasie considered the faster fix first, manual review for the flagged resumes instead of a delay. She rejected it, because a manual gate hides the same bias behind a curtain and does nothing for the ranking itself, or for the next feature built on the same score. The AI-specific failure worth naming by name is proxy bias through job-title phrasing: "reservist," "caregiver," and "career break" aren't a protected category by name, but they correlate closely enough with one that a scoring model can learn the pattern without anyone teaching it to. The guardrail that catches it is a sliced eval set, checked by subgroup before every launch, never trusted from one aggregate accuracy number alone. And the trade-off was real and taken on purpose: nine days of delay against a paying client's deadline, accepted because a second bias incident inside a year would have cost Havistock far more than nine days, and because the delay was the only proof Milovan's team was ever going to believe.

And if you want to be sure it really works, try it somewhere else

Same four letters, a farm co-op instead of an ATS, and this time the override hid inside a camera instead of a job title.

Sedgemarsh runs Wrenmark for its member farms. A grower photographs a leaf, and Wrenmark flags disease before an agronomist ever drives out to look. Ikaika Solveth took over its roadmap after a decision that cost several smallholder members part of a season's crop.

Hand sketched flow diagram titled Sedgemarsh's own version of the same override, with five connected steps: Photo, old phone, Wrenmark scores it, Leaf rust missed, this step emphasized, No visit booked, Crop lost.
A different cause than Havistock's job-title scoring change. The same shape of silence let it reach a real field before anyone caught it.

The agronomy team's eval set had shown, months earlier, that Wrenmark's recall on leaf rust fell hard on photos taken with older phones, the kind most common among smallholder members with the smallest plots. The team flagged it before a government subsidy program's enrollment deadline. The feature shipped anyway. Real plots went unflagged. Real crop was lost before anyone caught it.

The decision Ikaika would take back Sedgemarsh never wrote down that the enrollment deadline had been chosen over the agronomy team's own evidence. From the co-op board's side, it looked like an ordinary launch, not a call anyone had actually weighed against a warning.

Ikaika inherited Reneke Nairnwood's team, and their first month of updates read exactly like Milovan's: on track, nothing flagged. The real test came when a drought-stress feature was set to roll out co-op wide. Reneke's team's eval set showed 91 percent recall on high-resolution photos and 58 percent on the older phones, the same camera gap as before. Ikaika delayed the rollout, retrained the threshold per camera band, and told the co-op board directly why, not just the engineers.

Mapped onto LEAD, the shape holds. The link is whether the agronomy team tells Ikaika about a gap before a season proves it the hard way. The early signal is the same one, how often they raise something unprompted, not a co-op-wide satisfaction survey that only runs once a year. The abuse Ikaika avoided was the board's own instinct to just apologize publicly and move on, a gesture with no decision behind it. And the decision matched Ephrasie's exactly: one real reversal, in public, tracked afterward by whether the team keeps talking.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: make one real decision follow the team's evidence, then track how often they speak up on their own.
Cost: no time this quarter to build a formal override log. Whoever owns the call writes the decision and the reason in the team's own channel, in public, the same day. A visible sentence beats an invisible spreadsheet nobody built.
The model got better, for real: say the scoring model's accuracy genuinely improves and the gap closes on its own. Track the pushback count anyway, because a team that got burned once needs more than one good number to believe the next one.

Where people run it wrong.
They treat an open door or a listening session as the fix itself, instead of the thing you do before you have a real decision to point to.
They measure trust with a survey that only runs once a quarter, and steer by a number that's already months behind the truth.
They make the reversal privately, to be kind, and lose the entire point, since a decision nobody sees can't prove anything to a team that's watching for proof.

How to use it live. When an interviewer asks how you'd build trust with a burned team, ask yourself one question before answering out loud: "what would they actually see me do differently, not hear me say." That's the whole answer in one line.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question that asks how you'd build trust through a real, checkable signal?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. Built for metric questions, it ties "trust" to something observable instead of a feeling, and names the number that moves before anyone can feel it.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Ephrasie Culdrew, who takes over Riffle's roadmap at Havistock, and Milovan Beckthorn, who leads the ML team burned three times by the previous PM, Adelric Verlach.
3 · THE LINK
What should "the team trusts me again" actually be measured against?
Tap to flip
ANSWER
Whether the team proactively raises pushback and bad news on its own, not whether they say kind things in a listening session.
4 · THE EARLY SIGNAL
What moved first here, weeks or quarters, and what took longer to catch up?
Tap to flip
ANSWER
The weekly pushback count moved first, climbing from zero to seven over twelve weeks. The quarterly trust survey score only confirmed it two quarters later.
5 · THE OLD DECISION
What decision would Havistock take back?
Tap to flip
ANSWER
Never keeping a shared, visible record of times a PM overrode the ML team's eval-based pushback, so three separate overrides looked like three unconnected one-offs to anyone above Adelric.
6 · THE NUMBER
Fill in the blank: recall on the flagged resumes fell from ___ percent to ___ percent under the new confidence-score model.
Tap to flip
ANSWER
88 to 65 percent. The same shape of drop the previous PM had overridden three times before.
7 · THE REPLAY
Same nine-day deadline pressure, but a real decision follows the evidence. What changes?
Tap to flip
ANSWER
The launch delays nine days, in public, logged and visible. Weekly pushback climbs from zero to seven over twelve weeks, and the quarterly trust score doubles from 34 to 71 percent two quarters later.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs LEAD again on a different product. Which one, and what's the parallel failure there?
Tap to flip
ANSWER
Wrenmark, a crop-disease model at the farm co-op Sedgemarsh, run by Ikaika Solveth. There the same override pattern hid false negatives on leaf rust for photos from older, low-resolution phones common among smallholder members.

Check yourself Score: 0 / 0

Short answer, name the reversal
1. What old decision would Havistock take back, and why did it make sense the first time nobody thought to write it down?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Never keeping a shared, visible record of times a PM overrode the ML team's own eval evidence. It made sense when overrides were rare enough that nobody thought to log them. It stopped making sense once the same team got overridden three times and nobody above the PM could see the pattern.
Multiple choice
2. Why did the weekly pushback count matter more here than the quarterly trust survey?
  • A. Because the survey was measured incorrectly.
  • B. Because the pushback count moved the same week a real decision happened, while the survey only ran once a quarter.
  • C. Because Milovan preferred chat messages over surveys.
  • D. Because Havistock didn't own a survey tool.
Show hint
Check the "E, early signal" step in the framework recap.
Show answer
B. The pushback count is the leading indicator. It's ready the same week a decision moves, while the survey lags by a full quarter or more.
True or false
3. True or false: holding an open "ask me anything" session in week one was what actually started rebuilding trust with Milovan's team.
  • True
  • False
Show hint
Look at the "A, abuse" step in the framework recap.
Show answer
False. The team stayed quiet and told her nothing real. Trust only started moving in week four, once a real launch date changed because of their evidence.
Fill in the blank
4. Recall on the flagged group fell from ___ percent to ___ percent under the new title-match confidence score.
Show hint
Look at stage 5 of the walkthrough, or the line chart's chart note.
Show answer
88 percent to 65 percent. The same shape of drop the previous PM had already overridden three times.
Short answer, apply it yourself
5. Think of a team you've worked with that stopped pushing back on something. What's one real decision you could make, based on their actual evidence, that would prove you're listening, not just say so?
Show hint
Think of something with a date, a dollar amount, or a scope that would visibly change, not a meeting you'd hold.
Show answer
Model answer: If a support team stopped flagging a recurring bug because it was never prioritized, actually moving that fix ahead of a planned feature, in public, and saying why, would be the visible proof. Asking them to "keep raising things" changes nothing they can check.
Short answer, work the number
6. If Ephrasie had waited for the quarterly trust survey instead of watching the weekly pushback count, about how many weeks would have passed before she had any real evidence that trust was returning?
Show hint
A quarter runs about thirteen weeks. Compare that to when the pushback count first started climbing.
Show answer
About thirteen weeks, roughly a full quarter. The pushback count was already climbing by week five. Waiting on the survey alone would have cost her two extra months of not knowing.
Before you close the answer
Why this works
Tests whether you know that trust isn't measured by how a team behaves toward you, but by whether they'll risk telling you something you don't want to hear. Most candidates answer with gestures, listening sessions, one-on-ones, an open door, and never name a number that would tell them if any of it worked.
Follow-up traps
"Isn't delaying a launch every time the team pushes back just letting them run the roadmap?" Response: no, it's following real evidence once, early, on purpose, to prove the mechanism works. Not every future pushback earns a delay, but the first real one has to, or the team never has a reason to believe the next one will land differently.

"What if the team never pushes back again, even after you reverse a decision for them?" Response: then the pushback count itself is the answer, staying flat at zero after a real reversal is real information, and it means the fix didn't land, not that the team ran out of problems to raise.
If pressed
The override log Ephrasie built doesn't just record that a decision went against the team's evidence, it records the eval numbers themselves at the time of the call, not a summary written after the fact, because a PM revising the numbers in hindsight to look more reasonable is its own quiet way of losing the same trust a second time.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more