How would you build trust with an ML team that has been burned by a previous PM?
Riffle is Havistock's screening tool. It reads a stack of resumes, matches them against a job's real requirements, and ranks who an ATS should move forward first. Ephrasie Culdrew took over its roadmap six weeks after the last PM shipped a change nobody on the ML team trusted anymore. Milovan Beckthorn leads that team. He flagged the problem three times and got overridden three times. Ephrasie's first three syncs with him went fine. That was the part that should have worried her.
- Make a real, visible decision based on the team's eval evidence, early, even if it costs a deadline.Why: this is the one thing a burned team actually checks. Words don't move it. A reversed decision does.
- Track how often the team raises pushback or bad news on its own, as your real trust metric.Why: it moves in weeks. A survey or a gut feeling moves in months, if it moves at all.
- Write down every time a decision goes against the team's eval-based pushback, and make the record visible above you.Why: a hidden pattern of overrides is exactly what let the last PM do it three times unnoticed.
- Never treat an open-door line or a one-time ask for candor as the fix by itself.Why: it changes no actual decision, and a team that's been burned once can feel the difference immediately.
- Watch the slow signals too, a trust survey, attrition, but don't steer by them.Why: by the time a quarterly number moves, you've either already fixed the problem or already lost the people.
- Leave alone any part of the team that the last decision never actually burned.Why: making them re-earn trust they never lost reads as suspicion, not leadership.
How to answer this, stage by stage
Nobody is grading whether you can be nice to a team that got hurt. They're grading whether you can name the actual thing trust moves on, instead of a list of gestures.
Let's learn
Here is what happens when a team stops telling you the truth, and it looks exactly like everything going fine.
Riffle is the tool Havistock, an applicant tracking company, sells to recruiters. It reads a stack of resumes against a job's real requirements and ranks who should move forward first. Milovan Beckthorn's team builds the model underneath it.
For six weeks last spring, that team raised a problem three separate times. Their own eval set, 640 real resumes with a recruiter's own answer already known for each one, showed that a new scoring change was quietly punishing resumes with a real employment gap and a non-standard job title. Reservist. Caregiver. Career break. Recall on that group of resumes fell from 89 percent to 62 percent. The team said so in writing, three times. The PM at the time, Adelric Verlach, shipped it anyway, to hit a big client's renewal deadline.
Two months after launch, the client ran its own audit. Candidates with a real employment gap were advancing 40 percent less often than similar candidates without one. That became a formal complaint, then a remediation project, then an apology nobody at Havistock wanted to write. Adelric moved off the team a few weeks later.
Ephrasie Culdrew took the roadmap six weeks after that. Her first three syncs with Milovan's group went fine. Status only. Nothing flagged. Green across the board.
In her first week, before she understood what quiet actually meant here, she tried the easy fix. An open door. "Come tell me anything, anytime." Nobody did. They were being perfectly polite, and they told her nothing real, because nothing had happened yet to prove her door meant something different from the last one.
Here is the turn. Three clean syncs were not good news. They were the same shape as the six weeks before the audit, a team saying everything is fine right up until it very much was not. The extra silence wasn't the real problem either. The real problem was that Ephrasie had no way to tell a team that trusted her from a team that had simply learned not to bother.
Week four, Milovan's team ran the same 640-resume eval set against a new feature, a title-match confidence score, nine days before it was due to ship to a new client. Recall on the exact same flagged group had fallen from 88 percent to 65. Milovan raised it the smallest way possible, one line in a chat message: "Might be noise, wanted you to see it before Thursday."
For about an hour, Ephrasie wanted to ship anyway. The deadline was real, the client was big, and 65 percent still sounded like a working model. That is the exact pull that got Adelric three times.
She pulled the eval breakdown herself instead. The drop was real, and it repeated on a second sample. Someone on the team floated a faster fix: route the flagged resumes to a person for manual review instead of retraining the score. She turned it down. A manual gate hides the same bias behind a curtain. It does nothing for every gap resume the gate happens to miss, and it leaves the ranking itself untouched for the next feature built on top of it.
So she delayed the launch nine days, fixed the threshold, and told the client the real reason instead of a vague one. Then she told Milovan, in the same channel where he'd raised it, that the delay was because of his numbers, not despite the deadline.
Something moved that a survey never would have caught in time.
What it costs at its worst, if a PM can't tell the difference between a quiet team and a trusting one: the next bad decision looks exactly like the last one, right up until the same kind of audit finds it.
What I would leave alone: Havistock's data engineering team, the group that keeps resumes flowing into Riffle's pipeline, never dealt with Adelric's overrides directly. Running them through the same "prove yourself" ritual would cost time, and it would read as suspicion aimed at people who never did anything to lose it.
The lesson: a team that stops arguing with you isn't necessarily convinced. Sometimes it's just tired of being right and ignored. The fastest way to tell the difference isn't a survey. It's whether they say something unprompted the next time it actually matters.
Now here is the same thing as a story
The short version above is what you'd actually say in an interview. Read this one for the six weeks it took before anyone could tell whether Milovan's silence meant calm or damage.
Milovan Beckthorn has led Havistock's ML team for four years. Ask anyone on it and they'll tell you the same thing: he doesn't raise a flag unless the eval set backs it up, and when he does, he's usually right.
Under Adelric, being right three times in six weeks changed nothing. Milovan's team built the 640-resume eval set themselves, over a weekend, specifically because they suspected the new title-consistency scoring change was quietly hurting resumes with an employment gap. Week two, recall on that group came back at 71 percent, down from a baseline of 89. They wrote it up and sent it to Adelric. He said the sample was probably small and asked them to keep an eye on it.
Week four, the same eval set, run fresh: 65 percent. Same warning, more detail this time. Adelric held an actual meeting about it. He said the right things. "I want us to be a team that surfaces this stuff." He asked good questions. Then he asked whether it could wait until after the client renewal.
Week six, the number held at 62 percent. Adelric shipped the feature to the client's renewal deadline that Friday, with a note in the changelog that called it "a minor scoring refinement."
Two months later, the client's own compliance team ran a routine audit. Candidates with a documented employment gap were advancing 40 percent less often than similar candidates without one. Their next email to Havistock had the word "discriminatory" in the subject line.
I want to say the problem was the scoring change itself. It wasn't, not really. The scoring change was a mistake anyone could make. The real damage was that Milovan's team told the truth three times, in writing, with real numbers, and watched it not matter three times. By the time Ephrasie showed up, they weren't hiding anything. They just weren't going to bother saying it out loud again until they had a reason to believe it would land differently.
Ephrasie's first three syncs with the team went fine. Status updates only. "On track." "Green." Nothing raised, nothing flagged, three weeks running.
For about a week, she read that as a good sign. Then she remembered what she'd been told in her own handoff conversation about why Adelric left, and the quiet started to look like something else.
Her first move was the easy one. She called a session, said her door was always open, meant every word of it. Milovan's team thanked her and told her nothing real. Not because they doubted her sincerity. Because sincerity had never been the thing that was missing.
The real test showed up in week four, and it wasn't dramatic. A new client wanted the title-match confidence score fast-tracked, nine days out. Milovan ran the same 640-resume eval set his team had built two years earlier, the one that had caught the last problem and been ignored. Recall on the same flagged group, gap resumes with non-standard titles: 65 percent, down from 88.
He almost didn't send it. He told her later he'd sat on the message for twenty minutes, half convinced it would go exactly the way it always had. Then he wrote one line: "Might be noise, wanted you to see it before Thursday." The hedge was doing real work. It was a man giving his new PM an easy way to ignore him, because that's what the last nine months had taught him to expect.
Ephrasie read it at her desk with the client contract open in the next tab. For about an hour, she thought about shipping anyway. Sixty-five percent still sounded like a working number, the deadline was real, and nobody would have blamed her for a decision made under someone else's timeline.
Then she pulled the breakdown herself, instead of taking Milovan's word or her own instinct. The drop was real. It held on a second sample. Someone suggested routing the flagged resumes to manual review instead, a faster fix that would still hit the deadline. She almost took it. It would have worked, technically. It also would have hidden the same bias behind a curtain and done nothing for the next feature built on the same score.
She delayed the launch nine days. She told the client why, plainly, no soft language. Then she wrote back to Milovan in the same channel: "Delaying nine days. Your numbers are why. Thank you for sending it even hedged."
Adelric had asked for the truth and filed it away. Ephrasie asked for the truth and let it move a real date.
The decision she'd take back sits further back than her own six weeks. It sits in whatever meeting, years earlier, decided that a PM's call to override the ML team's own evidence didn't need to be written down anywhere two people could read it later. That gap is what let Adelric do the same thing three times and have it look, from above, like three separate ordinary judgment calls instead of one pattern.
Ephrasie built the fix the following week. An override log. Any time a product decision runs against the ML team's own eval-based pushback, whichever way it goes, the evidence, the decision, and the actual reasoning get written down in one place her own boss can read. Not a private chat thread. A record.
Twelve weeks after the delay, Milovan's team was raising something unprompted about twice a week, then three times, climbing toward seven by week twelve. The quarterly survey question, "I can raise a concern about a product decision without it costing me," had sat at 34 percent the last time anyone measured it under Adelric. Two quarters after the delay, it read 71 percent. The number that would have told her the truth fastest moved in weeks. The number everyone actually tracks moved in quarters, and only after the faster one had already said so.
What I'd tell myself, back in that first week with the open door and the good intentions: an invitation to speak up is not the same thing as proof that speaking up works. The team already knew that. I was the only one who needed convincing.
LEAD, so a burned team gets proof and not a promise
Not a way to prove Adelric was a bad person. LEAD is what forces you to name the actual signal trust moves on, and to catch a fake version of rebuilding it before a team quietly checks out for good.
The recap, one line per letter: link trust to whether the team tells you something's wrong before you ask, not to whether they like you. The early signal is the weekly pushback count, because it moves the same week a real decision does and a quarterly survey does not. Name both abuses plainly, a listening session with no decision behind it, and a one-time ask for candor that gets forgotten the moment it's inconvenient. And the decision is what makes it real: one visible reversal, a written record above you, and a number you track instead of a feeling you hope for.
Two things worth saying outright, since the real judgment sits here. Ephrasie considered the faster fix first, manual review for the flagged resumes instead of a delay. She rejected it, because a manual gate hides the same bias behind a curtain and does nothing for the ranking itself, or for the next feature built on the same score. The AI-specific failure worth naming by name is proxy bias through job-title phrasing: "reservist," "caregiver," and "career break" aren't a protected category by name, but they correlate closely enough with one that a scoring model can learn the pattern without anyone teaching it to. The guardrail that catches it is a sliced eval set, checked by subgroup before every launch, never trusted from one aggregate accuracy number alone. And the trade-off was real and taken on purpose: nine days of delay against a paying client's deadline, accepted because a second bias incident inside a year would have cost Havistock far more than nine days, and because the delay was the only proof Milovan's team was ever going to believe.
And if you want to be sure it really works, try it somewhere else
Same four letters, a farm co-op instead of an ATS, and this time the override hid inside a camera instead of a job title.
Sedgemarsh runs Wrenmark for its member farms. A grower photographs a leaf, and Wrenmark flags disease before an agronomist ever drives out to look. Ikaika Solveth took over its roadmap after a decision that cost several smallholder members part of a season's crop.
The agronomy team's eval set had shown, months earlier, that Wrenmark's recall on leaf rust fell hard on photos taken with older phones, the kind most common among smallholder members with the smallest plots. The team flagged it before a government subsidy program's enrollment deadline. The feature shipped anyway. Real plots went unflagged. Real crop was lost before anyone caught it.
Ikaika inherited Reneke Nairnwood's team, and their first month of updates read exactly like Milovan's: on track, nothing flagged. The real test came when a drought-stress feature was set to roll out co-op wide. Reneke's team's eval set showed 91 percent recall on high-resolution photos and 58 percent on the older phones, the same camera gap as before. Ikaika delayed the rollout, retrained the threshold per camera band, and told the co-op board directly why, not just the engineers.
Mapped onto LEAD, the shape holds. The link is whether the agronomy team tells Ikaika about a gap before a season proves it the hard way. The early signal is the same one, how often they raise something unprompted, not a co-op-wide satisfaction survey that only runs once a year. The abuse Ikaika avoided was the board's own instinct to just apologize publicly and move on, a gesture with no decision behind it. And the decision matched Ephrasie's exactly: one real reversal, in public, tracked afterward by whether the team keeps talking.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: make one real decision follow the team's evidence, then track how often they speak up on their own.
Cost: no time this quarter to build a formal override log. Whoever owns the call writes the decision and the reason in the team's own channel, in public, the same day. A visible sentence beats an invisible spreadsheet nobody built.
The model got better, for real: say the scoring model's accuracy genuinely improves and the gap closes on its own. Track the pushback count anyway, because a team that got burned once needs more than one good number to believe the next one.
Where people run it wrong.
They treat an open door or a listening session as the fix itself, instead of the thing you do before you have a real decision to point to.
They measure trust with a survey that only runs once a quarter, and steer by a number that's already months behind the truth.
They make the reversal privately, to be kind, and lose the entire point, since a decision nobody sees can't prove anything to a team that's watching for proof.
How to use it live. When an interviewer asks how you'd build trust with a burned team, ask yourself one question before answering out loud: "what would they actually see me do differently, not hear me say." That's the whole answer in one line.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the team never pushes back again, even after you reverse a decision for them?" Response: then the pushback count itself is the answer, staying flat at zero after a real reversal is real information, and it means the fix didn't land, not that the team ran out of problems to raise.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Working with ML engineers and researchers
- #1 How do you write a requirement for a team whose output is a probability distribution?
- #2 An engineer says the model cannot do that. What questions do you ask before accepting it?
- #3 Describe how you would run a planning session when effort estimates are genuinely unknowable.
- #4 What does a healthy PM-to-research relationship look like when research timelines are open-ended?
- #5 How do you keep a research team connected to user problems without constraining their exploration?
- #6 Your ML team wants three months to improve accuracy by two points. How do you evaluate that ask?