…
AI Product Case Questions

Models gave inconsistent predictions. How did you work with scientists?

A worked answer to a real AI PM interview question: the model gave inconsistent predictions. How did you work with the scientists?

Transcript

Read the full transcript (1,324 words)

[INTERVIEWER] Models gave inconsistent predictions. How did you work with scientists? "The models gave inconsistent predictions. How did you work with the scientists?" This question is really asking whether you can work with research, not over them. A strong answer shows you took a vague complaint about model inconsistency and turned it into a shared, measurable definition of success that the PM and the scientists both agreed on.

Instead of arguing across two different dashboards, you were suddenly solving the same problem on the same side. The core test here is whether a PM and a research team can align on a shared evaluation, because that metric is what makes the entire collaboration work. The classic failure is the standoff. The PM sees angry users, the scientists see a metric that looks completely fine, both of them are right, and nobody moves.

In the next few minutes I'll give you the scaffold, the collaboration moves that carry the answer, a full worked example, and the follow-ups. Pick a time the output of the model was inconsistent and you had to work with data scientists or researchers to sort it out. Start with the Situation, taking about forty-five seconds. Describe the inconsistency concretely.

You might say the same input gave different risk scores across runs, or predictions swung week to week with no product change while support received complaints. Concrete details beat saying the model was unreliable, because you can't build a metric around a vague feeling. Then the Task, taking thirty seconds. This is your job. The important framing here is that your role wasn't to fix the model yourself.

It was to align the team on what good means and drive towards it. Name the friction plainly. The PM sees user complaints while the science team sees a metric that looks healthy. Two truths, no shared language. Now the Action, taking three and a half minutes, covering three collaboration moves. Move one is to reproduce and locate the issue together.

How did you turn the complaint into something specific? Was it real inconsistency, non-determinism from sampling, data drift, or a threshold problem? Notice the word together. You don't diagnose it for the scientists and hand them a verdict. You bring the failing cases and work through them collaboratively. Move two is to co-define a shared metric and evaluation set. This is the core move of the whole answer.

You and the scientists agree on a golden set and a metric that captures what users actually experience. The ideas that the model is fine and users are unhappy stop being two separate truths and collapse into one number you both own. Measure the consistency explicitly, such as the variance across runs on a fixed set, not just average accuracy, because average accuracy is exactly what hid the problem.

Move three is to respect the boundary. You own the product bar and the user impact. They own the modelling. You bring the cases, the priorities, and the acceptance threshold. You don't walk in and tell them which loss function to use. Crossing that line is how you lose a research team for good. Then the Result, taking about a minute and a half.

Cover the aligned metric, the improvement, and a specific number. Show the team stopped talking past each other because they finally shared one definition of success. Here's the whole thing joined up. "Our lending product showed applicants a risk tier from a model the data science team owned. Support started getting complaints that the same applicant, re-checking a day later with no new actions on their account, would see a different tier.

The science team dashboard looked completely healthy and AUC was steady, so at first we were just talking past each other. I saw angry users, they saw a fine metric, and both of us were correct. My job wasn't to fix their model. It was to get us aligned on what good even meant. First, we reproduced the issue together.

I pulled two hundred real cases where the tier had flipped and sat with the lead scientist to find the causes. Some flips were genuine data updates, which is fair enough. But a lot of them were the model sitting right on a tier boundary, where the daily data refresh nudged the score just enough to cross the line and flip the label, even though the underlying risk had barely moved.

The dashboard hid this completely because average accuracy doesn't care about a handful of boundary cases. We then co-defined a new shared metric called tier stability. This was the percentage of unchanged applicants who kept the same tier across a week, measured on a fixed golden set of five hundred cases we built together. That number was the thing users actually felt, and now it was a number both teams owned.

The scientists then added hysteresis around the boundaries and recalibrated the thresholds, which was their call, not mine. I owned the acceptance bar, meaning stability had to clear ninety-five percent before we shipped. The result was that tier flips on unchanged applicants dropped from about twelve percent to under three percent. The complaints basically fell away. Honestly the bigger win was that stability went onto the shared dashboard permanently, so we never again had that standoff between the model being fine and users being unhappy.

We had one number and we were both looking at it." That is the complete answer, and the move that carries it is co-defining the metric because it ended the standoff. Here's what makes them lean in. First, you reproduced the issue with the scientists and brought failing cases rather than opinions. Two hundred real flipped cases beats users saying it is flaky every single time.

Second, you established a co-defined shared metric that captured the user pain the old dashboard was blind to. That's the PM adding something the research metric couldn't see. Third, you showed clear respect for the boundary. You owned the product bar and they owned the modelling. That's the difference between a PM research wants to work with and one they quietly route around.

Now for the ways people sink this. The first trap is telling the scientists how to fix the model or dismissing their metric as simply wrong. The moment you say their AUC is meaningless, you've made an enemy and learned nothing. Their metric wasn't wrong, it was just measuring a different thing than what users felt. The second trap is leaving the word inconsistent vague, meaning there's nothing to measure and no way to agree, which ensures the standoff never ends.

The third trap is having no shared evaluation at all, leaving the PM and research pointing at two different dashboards forever. Expect a follow-up question. They'll ask what you would do if the scientists disagreed that stability was the right metric. A good answer is that you don't impose it, but instead ground it in the failing cases you both looked at.

The metric has to explain the complaints or it's not the right one, making it a shared and checkable standard rather than a mere opinion. They might also ask how you decided ninety-five percent was the bar. You can say you set it against the complaint volume, specifically the threshold where flips stopped generating a meaningful support load. To bring the whole picture together, take the vague complaint and reproduce it with the scientists using real failing cases.

Co-define a shared metric on a shared golden set that captures what users actually feel. Respect the boundary by owning the product bar while they own the model. Then land the number and show the standoff ended. The one line to carry in is to turn the complaint about model inconsistency into a shared metric on a shared golden set.

Once you both own one number, you're on the same side of the problem instead of sitting across the table from each other.

Keep learning