How do you decide when a model is good enough to ship over the objections of the team that built it?
Ferrous Pay processes card payments for online merchants, and Driftcatch is the model that scores every transaction for fraud risk before it gets approved. Ursina Thorvaldsen is the product manager who owns the call on whether a new version of Driftcatch ships. Jokull Ochoa leads the team that builds it. Two days before a release that cleared every number they'd agreed on, Jokull asked her to hold it anyway, over a pattern his team could feel but not yet prove.
- Ship when the model clears the pre-agreed bar, judged against the eval set and the business case, not the team's mood in the room.Why: a felt confidence level isn't a metric, and it moves for reasons that have nothing to do with whether the model actually works.
- Treat every objection as a request for evidence: ask for the specific eval result or failure case it's tied to.Why: without that question, you can't tell a real gap from a bad feeling, and both feel identical from across the table.
- If the objection has no evidence yet, test it fast before the ship date, don't just wave it off or grant it on faith.Why: a few hours of a targeted backtest turns an unfalsifiable feeling into a number either side can act on.
- Never let an objection with no evidence block a release that clears the bar, once it's actually been tested and held up.Why: that's the same failure as ignoring a real objection, just running in the other direction, and it teaches everyone that feelings outrank numbers.
- Write down every override, what was raised, whether it was evidence based, and what happened once it got tested.Why: without a record, every release re-fights the same trust argument from zero, instead of learning from the last twenty.
- Carve out the cases that skip this whole test: a hard regulatory hold never needs to clear an eval bar first.Why: some costs aren't measured in dollars you can recover, so those get an automatic stop regardless of what any number says.
How to answer this, stage by stage
Nobody is grading whether you can define "false positive." They are grading whether you can point at one real ship call and say, out loud, what would actually change your mind.
Let's learn
Driftcatch is the model that looks at every card payment Ferrous Pay processes and scores it, in under a second, for how likely it is to be fraud.
Before Ferrous Pay had a real eval bar, ship calls got made in a room. Someone felt good, or someone felt uneasy, and whoever felt more sure carried the day. It worked, until a version shipped on full team confidence and missed a fraud ring running small test charges through newly opened merchant accounts. Nobody had checked for that pattern specifically, because nobody had checked for anything specifically. It cost three hundred eighty thousand dollars over six weeks before an outside chargeback alert caught it, not Driftcatch.
After that, the team agreed on a bar, in writing, before any release: catch at least ninety percent of confirmed fraud dollars on a frozen, labeled eval set, and keep the false decline rate, real customers wrongly blocked, under two point two percent. Same test, every release, decided before anyone saw a number.
Here's what that bar actually looked like the week Driftcatch v3 came up for release.
By every number the team had agreed to check, v3 was ready. It was better than v2 in both directions at once, more fraud caught, fewer good customers wrongly blocked. There was no tradeoff to argue about on paper.
Two days before the release, Jokull Ochoa's team, doing their normal weekly manual review, found something the eval set had never seen. Forty three confirmed cases over the last twelve days, card testing bursts, a dozen or so small charges fired in quick succession against freshly issued virtual cards, all tied to Candlewick Gifts, a digital gift card merchant that had only joined Ferrous Pay five weeks earlier. The frozen eval set had been locked eleven weeks before that, so not one of these forty three cases was in it.
Jokull couldn't point to a failed eval score, because there wasn't one to point to. He had forty three case files and a bad feeling that v3, trained mostly on the old mix of merchants, might not have generalized to a brand new card type from a brand new vertical. He asked Ursina to hold the launch until the team could be sure.
This is where a lot of ship decisions go wrong in one of two directions. Override Jokull outright, and if his team's read had been right, Ferrous Pay ships something that misses a real, growing pattern, on purpose, with a date as the only reason. Grant the hold on the feeling alone, and Ferrous Pay just recreates the exact failure the eval bar was built to stop, except this time the mistake points the other way: a good model sits on a shelf because nobody asked for a test.
Instead of picking a side on faith, Ursina asked for a third option: test the actual worry, fast, before the ship date. Jokull's team pulled the forty three known cases and re-ran them through v3, in isolation, scoring each one the way the model would have scored it live.
They ran it in batches as the cases came in, and the early numbers looked bad on purpose to nobody in particular, just because small samples wobble.
Thirty nine of the forty three known card testing cases scored above the block threshold. Four slipped through, totaling thirteen hundred sixty dollars attempted, each one under the five hundred dollar cap that Ferrous Pay's chargeback insurance already covered. The whole check took about three hours.
Ursina shipped v3 on schedule. Not because she overruled Jokull, and not because she took his word for it either. Because once the actual worry got tested, it turned out Driftcatch v3 handled it about as well as the bar already required, and the gap that was left was small enough to absorb.
What I would leave alone: anything compliance or legal flags for a hard regulatory reason, a sanctions list match, an anti-money-laundering flag, skips this whole bar. Those get an automatic hold no matter what the eval numbers say, because the cost of being wrong there isn't a chargeback you can insure against. It's regulatory exposure, and that's not the kind of thing a catch rate can weigh against a launch date.
The lesson: a pre-agreed number exists so a ship call doesn't get decided by whoever is loudest or most anxious in the room that day. But the number isn't a wall either. When someone brings a specific, testable worry, the right move is never to overrule them by decree or grant them on faith. It's to make the worry pay for itself, fast, in hours, and let whatever it turns up decide the call instead of how anyone happened to feel walking in.
Now here is the same thing as a story
The short version above is what you actually say out loud. Read this one for the three hours nobody wanted to spend, and the twenty two minutes it took to find out they were worth spending.
Jokull Ochoa has led the team that builds Driftcatch for two and a half years. Ask anyone at Ferrous Pay and they'll tell you the same thing: if Jokull's team ships something, it works. He reads a confusion matrix the way other people read a weather report, fast, and mostly right about what it means for tomorrow.
Ursina Thorvaldsen has run product on fraud risk for about as long. She built the ship criteria herself, in the weeks after the three hundred eighty thousand dollar miss that nobody at the company likes to bring up by name anymore. Five releases had run clean against that bar since. Nobody had needed to raise their voice in a launch meeting in almost a year, because there was finally a number to check instead of a feeling to defend.
Driftcatch v3 was supposed to be an easy one. Better on both numbers the team tracked, more fraud caught, fewer good customers wrongly blocked. Ursina had the release notes half written on Monday.
Then Wednesday, Jokull asked for fifteen minutes.
His team's weekly manual review had turned up something they couldn't stop looking at: forty three cases, all in the last twelve days, all tied to a new merchant, all shaped the same way. A dozen tiny charges, fired fast, against a virtual card that had been issued that same morning. It read like someone testing stolen card numbers to see which ones still worked, at a merchant Driftcatch had barely seen before. Jokull didn't have a failed test to hand her. He had case files, a gut that had been right more often than wrong for two and a half years, and a launch date in two days.
"I just don't want to find out the hard way that we trained this thing on the wrong customers," he said.
Here's the part worth being honest about: he wasn't being difficult, and he wasn't being cautious for its own sake either. Ferrous Pay's whole reason for having an eval bar in the first place was a room full of people who'd felt confident once and been wrong. Jokull's discomfort was doing exactly the job it was supposed to do. The only thing missing was a way to check it that didn't come down to whose gut people trusted more.
Ursina didn't say yes and she didn't say no. She asked Jokull's team to do one thing before Thursday's decision: pull the forty three cases, and run them back through v3, exactly as the model would have scored them live, and see what came back.
They started at eleven that morning. The first batch, eleven cases, came back rough. Eight caught, three missed. Seventy two percent. Someone in the room said, quietly, that maybe the hold was the right call after all.
Nobody made the decision on eleven cases. They kept going. Twenty two cases, eighteen caught, just under eighty two percent. Thirty three cases, twenty nine caught, almost eighty eight percent. By two in the afternoon, all forty three were in: thirty nine caught, ninety point seven percent, four missed for thirteen hundred sixty dollars combined, every one of them small enough that Ferrous Pay's own insurance already covered it without a claim being questioned.
Ninety point seven percent, against a bar of ninety. The worry Jokull couldn't put a number on turned out to have almost exactly the number the whole model already had to clear.
The alternative that actually got proposed in that Wednesday meeting, and it's worth saying because it sounded reasonable for about thirty seconds, was to give engineering a standing veto over any release they weren't comfortable with, no test required. It lost, because Ferrous Pay had already lived that version of the process, just aimed the other way, back when unchecked confidence shipped something broken. An unchecked veto doesn't fix a trust problem. It just moves who gets to skip the evidence.
Ursina shipped v3 Thursday morning, on schedule. Same launch date the release notes had said since Monday. The difference wasn't that Jokull got overruled, and it wasn't that his team got a blank check either. It was that a feeling nobody could quite defend on Wednesday became a number by two o'clock, and the number is what actually decided it.
What I'd tell myself, standing in that Wednesday meeting again: the three hours it took to test Jokull's worry weren't the cost of the delay. They were the thing that made the launch defensible either way it had gone.
PICK: the difference between a number and a feeling
PICK only earns its place here if you can name the exact question that separates a real objection from an understandable one that still isn't evidence.
The trade off worth saying out loud: the backtest cost about three hours and pushed the ship decision half a day later than planned. That's a real cost, and it's not one you pay on every minor tweak, only on a version change big enough, and a worry specific enough, that testing it beats guessing either way. Reserve the full live check for exactly those moments, or the discipline itself becomes the bottleneck it was built to prevent.
And if you want to be sure it really works, try it somewhere else
Same four letters, a pet clinic's X-ray reader instead of a payment processor, and this time the untested pattern is a breed, not a merchant.
Legwatch reads leg X-rays for a chain of pet clinics and flags likely fractures for a vet to confirm. Ronja Salazar runs product there. The ship bar: catch at least 92% of vet-confirmed fractures on the frozen labeled set, keep the false-flag rate, a normal bone marked "possible fracture," under 8%. The new version, v4, cleared it clean: 94% caught, 6% false-flagged.
Two days before launch, Toivo Kallstrom, who leads the radiology team, raised a hand. Sighthound breeds, greyhounds and their relatives, have naturally different bone density on an X-ray, and the frozen eval set, built mostly from mixed-breed clinic visits, had almost no sighthound images in it. He couldn't point to a failed score. He had a dozen borderline cases from memory and a feeling the model might read a normal sighthound leg as a fracture, or miss a real one, because it had barely seen the shape before.
Mapped onto PICK: the position holds its shape, ship on the bar and real business cost, not on how ready anyone feels. The impact splits the same way too, overriding Toivo risks a real miss on a breed the eval set never covered, and granting the hold on the feeling alone recreates the same failure the bar exists to prevent, just aimed at a different vertical. The cost asymmetry lands the same: a small known gap in a rare breed is something a vet can catch on manual review, cheap and reversible, while holding the launch for weeks with no test attached delays every other breed's better read too, and that delay doesn't come back either.
The kill test transfers without changing a word. Ronja asked for the dozen sighthound cases the clinics could pull from the last month, backtested v4 against exactly those, and got an answer within the afternoon: eleven of twelve real fractures caught, no extra false flags on the naturally denser normal bones. The worry was reasonable. It just wasn't evidence yet, and once it was, the number won.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Say the whole test in one line: does the objection point at a specific eval failure, or is it a feeling, and if it's a feeling, test it fast before you decide either way.
Cost: there's no time this sprint for a live backtest. Fine, but at minimum ask what the smallest possible test would be, and run that instead of skipping the question entirely.
The model got better, for real: say the underlying model's overall accuracy genuinely improved on a fresh benchmark. Doesn't change the argument. A better aggregate number still needs to be checked against whatever specific pattern someone is actually worried about, because a benchmark score is not a substitute for testing the one thing somebody can actually name.
Where people run it wrong.
They treat any objection from the team that built the model as automatically credible, because they built it, and skip asking for the evidence.
They treat any objection with no evidence yet as automatically dismissible, and miss the one time in ten it was actually right.
They let "we don't have time to test it" become a reason to skip the test instead of a reason to make the test smaller.
How to use it live. Before answering, picture the actual room, two days before a real launch, one person with a number and one person with a feeling. If you can say out loud what would make you change your mind on the spot, you've already answered the question.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the backtest had come back at 60%, not 90.7%?" Response: then the objection would have had real evidence behind it, and the launch holds until that specific gap is fixed, no different from any other failed eval result.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Working with ML engineers and researchers
- #1 How do you write a requirement for a team whose output is a probability distribution?
- #2 An engineer says the model cannot do that. What questions do you ask before accepting it?
- #3 Describe how you would run a planning session when effort estimates are genuinely unknowable.
- #4 What does a healthy PM-to-research relationship look like when research timelines are open-ended?
- #5 How do you keep a research team connected to user problems without constraining their exploration?
- #6 Your ML team wants three months to improve accuracy by two points. How do you evaluate that ask?