What is the right cadence for reviewing model quality with your engineering team?
Fenharrow Sports Media builds Ridgecut, a tool that watches full hockey game footage, scores every play from 0 to 100 for how reel-worthy it is, and auto-cuts the eight highest-scoring clips per player into a recruiting reel within a couple hours of the final horn. Five weeks after an update meant to catch more subtle skill plays, a Sallowick Hockey Academy player's series-winning overtime goal scored 41 and never made his reel. Wrenna Brennock, the product manager, has to decide how often Ridgecut's model quality gets checked against Mirelle Ossley's four-person engineering team, so this never happens quietly again.
- Build the always-on dashboard, checked within a day of every model or prompt push, against a golden set with real messy footage.Why: everything else here only works once this exists.
- Trigger a full engineering review off a real signal on that dashboard, never off a date on a calendar.Why: this is what actually catches a regression before a family or a recruiter sees it.
- Stop trusting the quarterly snapshot as the main safety net.Why: it only ever compared two flat photos three months apart, with nothing logged between them to show a number climbing.
- Reject a standing weekly all-hands review as the default.Why: four of the last six quarterly checks found nothing worth four engineers' time, and it's the one ritual nobody can quietly cancel once other teams plan around it.
- Keep a light twice-a-year retro for what a score can't catch.Why: a brand new sport or a new camera rig is a blind spot in the golden set itself, not a number moving.
- Rebuild the golden set with real, messy in-venue footage.Why: the dashboard is only as honest as what it gets checked against, and clean broadcast footage is exactly what let this drift hide.
How to answer this, stage by stage
Nobody is grading whether you can say "weekly" or "quarterly" with confidence. They're grading whether the number of weeks you pick actually tracks how fast the thing underneath it changes.
Let's learn
Before Ridgecut ever watched a single game, four assistant coaches at Sallowick Hockey Academy stayed after every game with a laptop, splicing phone footage together by hand so a recruiting reel could go out by Sunday night. It took most of a Saturday evening, for one team, one week's worth of games.
Ridgecut is Fenharrow Sports Media's tool for doing that automatically. It watches the full game feed, scores every single play from 0 to 100 for how reel-worthy it is, and stitches the eight highest-scoring clips per player into a reel that lands in an inbox within a couple hours of the final horn.
For a long stretch, that worked exactly as promised. A reel that used to cost a coach four hours on a Saturday night landed in a parent's inbox by nine that same night, picked clean: the goals, the saves, the one hit that made the whole bench jump up. Coaches stopped double-checking it. They'd glance at the thumbnail, see the score matched the final, and send it along.
Then, five weeks before a game that mattered, Mirelle's team shipped an update to widen what counted as reel-worthy, so scouting-minded coaches would get more than just goals: smart positioning, a good zone entry, the plays a college scout actually watches for. It passed every check before it shipped. Nobody watched it after.
Here's the part that matters. It isn't that the model got a little worse at picking goals. It's that two ordinary things started happening at once, and nobody was watching either of them. Routine zone clears and neutral-zone entries started scoring higher, because that's exactly what the update asked for. And a shaky, crowded, camera-jostled celebration clip, the kind every real overtime goal actually looks like from the stands, started scoring lower, because the same update leaned harder on a calm, steady frame as a sign of quality.
On a Tuesday in March, Sallowick's top defenseman scored the overtime goal that won his team's league semifinal, in front of two college scouts, on camera. Ridgecut's reel for him that night had eight clips: two routine zone clears, three neutral-zone entries, a shift change, and two early shots on goal. The goal itself scored 41, one spot below the cut. It never made the reel.
What I would leave alone: Ridgecut also generates a short outro card at the end of every reel, a fixed template with the team crest and the final score. Nobody has touched that template's logic in a year, and nothing about it depends on a model. A slow, fixed cadence is completely fine for a part of the product that never changes, so it stays off this whole plan.
The lesson: a review cadence is not a promise about how often you meet. It's a bet about how fast the thing you're checking actually changes. Get that bet wrong, and you either burn a team's week watching a needle that hasn't moved, or you find out about a real one from a parent instead of a dashboard.
Now here is the same thing as a story
The short version above is what you'd actually say out loud. Read this one for the five weeks nobody at Fenharrow knew a number was moving.
Emrys Kettleby can read a rink from the stands like most people read a room. Nine years running video for Sallowick Hockey Academy, and he can tell you, before a whistle even blows, which of his players is about to get skated around. He didn't fight Ridgecut when it arrived two years ago. He tested it hard for a month, checked its picks against his own, and once it kept agreeing with him, he let it take over the Sunday-night splicing he used to do himself.
For a long stretch, that was the best trade Sallowick ever made. A reel that used to cost Emrys four hours on a Saturday night landed in a parent's inbox by nine that same evening, picked clean. He stopped watching every reel before it went out. He'd glance at the thumbnail, check the score matched the final, and send it along.
Five weeks before the semifinal, nothing about that changed on his end. He had no reason to look closer. Nobody had told him anything was different, because inside Fenharrow, nothing looked different either. The model still passed its own checks. The quarterly review wasn't due for another seven weeks. Between one flat snapshot and the next, there was no place anyone was actually watching a number move.
Then came the Tuesday in March that decided Sallowick's whole season. Overtime, the league semifinal, two college scouts sitting behind the Sallowick bench with clipboards. Sallowick's top defenseman, a kid who'd been quietly great all year and loud about nothing, picked the puck off the boards, walked it the length of the ice, and buried it top shelf. The building came up out of its seats. Emrys, filming from the mezzanine, jumped with everyone else, and the camera jumped with him.
Ridgecut had the footage processed and the reel built before Emrys got home. He didn't check it that night. Why would he. He'd stopped checking.
The next morning, one of the scouts called the academy's office asking for a cleaner copy of the winning goal, for his file. Emrys pulled up the reel Ridgecut had already sent the family. Two routine zone clears. Three neutral-zone entries. A shift change. Two shots from the second period. No goal.
He watched it twice before he believed it. Then he wrote one line back to Fenharrow's support inbox: you cut the goal that got him looked at by two schools.
I want to say the problem is that the model got worse. It did, on that one clip. But that's not really what happened. Two ordinary things drifted in opposite directions at once, five weeks earlier, and there was no place either of them was visible until a parent's coach found out from a scout instead of from us.
So here's the decision I'd take back. Every quarter, Mirelle's team ran the live model against a golden set and compared the number to the one from three months before. Two clean snapshots, nothing logged between them. That made sense when Fenharrow was three people and nothing shipped between two checks. Nobody ever built anywhere to watch the number while it moved.
I'd put that history back. Not a meeting. A dashboard that logs the score, by category, every time a model or a prompt ships, checked against footage that's actually messy, actually shaky, actually what a real overtime goal looks like on a phone from the stands. Checked within a day of any push, by whoever's on call that week. And a full team review only when that dashboard actually moves, not because a date came around on someone's calendar.
That's the whole difference. One design waits for a photo every three months. The other one watches the film run.
And the part I'd want to tell my past self: we never once asked how often the model would change. We just picked a cadence and got used to it, the same way Emrys got used to not checking.
ORDER, for picking a cadence that tracks the model, not the calendar
PICK would fit if this were two options and one clean tradeoff. There are four real candidates here, and the job is ranking them by what's hardest to undo. That's ORDER's job.
Three things worth stating directly, since the real judgment sits here. The alternative worth naming and rejecting is the standing weekly all-hands review, which looks like the safest, most committed option on paper. It loses because the evidence check already answered the question: four of the last six quarters had nothing worth a room full of engineers, and the two quarters that did have something still weren't caught any faster by a fixed date, only by finally watching the number move. The AI-specific failure worth naming is silent scoring drift after a compound model change: two ordinary shifts, one making routine plays score higher, one making messy-but-real footage score lower, that cancel each other out inside a single blended average and only show up once you track categories separately. The guardrail is exactly that: a dashboard that tracks reel-worthiness by category, not one flattened number, checked against footage that's genuinely messy. And the trade-off is accepted on purpose: real, ongoing engineering time goes into maintaining that golden set and someone checking the dashboard within a day of every push, in exchange for catching the next drift in days instead of finding out from a scout.
And if you want to be sure it really works, try it somewhere else
Same five letters, a hospital corridor instead of a rink, and the sanitized summary wears a different coat: the annual case-mix review.
Grimscar Diagnostics runs FirstRead across a chain of small radiology clinics, a model that flags which chest X-rays in the day's queue need a radiologist's eyes first. Bramric Vantrick, product manager, has to decide how often FirstRead's flagging quality gets checked against the radiologists actually reading the films, after a routine model update meant to catch more subtle findings started drifting the same way Ridgecut did: an ordinary shadow started scoring as urgent more often, while a real urgent finding, one with a slightly unusual shape, started scoring lower, because the update leaned harder on textbook presentation as a sign of confidence.
Same steps, mapped onto Grimscar. Outcome: protect whether a genuinely urgent film gets flagged before a radiologist would have found it anyway, not whether the review looks thorough. Reversibility: a monthly automated flagging-rate check is trivial to pause for a week; folding a data scientist into every clinic's daily reading list is not, once clinics start scheduling around that person being there. Dependency: the flag only means something checked against a golden set of real, difficult scans, not the clean teaching set the original model was validated on. Evidence: Bramric's team pulled the last several formal reviews and found real drift twice, both already weeks old by the review date, the same shape Wrenna found. Rank: the dashboard-plus-trigger ships first, the standing weekly review stays off the table, and the old annual case-mix review survives only as a light look at scan types the golden set doesn't cover yet, not as the main safety net.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the rank: ship the dashboard first, it's the one thing every other option secretly depends on anyway.
Cost: there's no budget yet to build the full automated dashboard this quarter. Ship a manual version, one engineer spot-checks the same golden set by hand once a week, and say plainly that it's smaller, rather than pretending the old quarterly audit alone still covers it.
The model got better, for real: say the pre-launch eval score already looked great. Keep the same cadence anyway. A good score on a clean golden set was exactly what let this drift hide the first time.
Where people run it wrong.
They pick the heaviest ritual because it looks like the most commitment to quality, not because it fixes anything.
They treat "we have a quarterly review" as proof of a cadence, when the real question is whether anything meaningful happens between reviews.
They let "always-on monitoring" quietly mean a dashboard nobody is actually assigned to look at, which is the same as having none.
How to use it live. Before ranking anything, ask out loud: "does this cadence track how often the thing being reviewed actually changes, or does it track the calendar?" If the honest answer is the calendar, say so, then fix it.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the dashboard has the same kind of blind spot that caused this in the first place?" Response: that's exactly why the golden set gets rebuilt with real, messy footage, and why a light twice-a-year retro stays alive for whatever a category score still can't see, like an entirely new sport.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Working with ML engineers and researchers
- #1 How do you write a requirement for a team whose output is a probability distribution?
- #2 An engineer says the model cannot do that. What questions do you ask before accepting it?
- #3 Describe how you would run a planning session when effort estimates are genuinely unknowable.
- #4 What does a healthy PM-to-research relationship look like when research timelines are open-ended?
- #5 How do you keep a research team connected to user problems without constraining their exploration?
- #6 Your ML team wants three months to improve accuracy by two points. How do you evaluate that ask?