How would you handle disagreement between two human reviewers?
PICK the product is Suresight, Redshale Mutual Insurance's fraud-flagging tool for auto claims
Redshale Mutual Insurance runs Suresight, a tool that flags auto claims worth a second look for possible fraud before an adjuster opens the file. Cassandra Voss has adjusted claims for eleven years. Nikolai Petrenko joined the claims floor eight months ago.
The direct answer
Route disagreement to a named third adjudicator who sees both reviewers' actual reasoning, not an average of their two scores and not an automatic win for whoever's more senior. Track the disagreement rate by claim type. Once it climbs past roughly one in three, stop treating it as two people failing to agree, since at that rate the policy itself is the thing that's actually unclear.
Do this, in order
Send disagreement to a real third adjudicator, never an average and never an automatic default to seniority.Why: averaging or deferring throws away the one signal a disagreement actually carries, that the case is genuinely hard.
Give that adjudicator both reviewers' full reasoning, not just their two final scores.Why: the number tells you they disagreed. The reasoning tells you why, and why is what actually resolves it.
Track disagreement rate per claim type, not just per case.Why: a rising rate on one claim type is a policy problem, not a one-off argument between two adjusters.
Treat a disagreement rate above roughly one in three as a signal to rewrite the guidance, not to retrain reviewers.Why: at that rate, the two people reading the same rule are reading it correctly, and it's the rule that's ambiguous.
Never let a newer reviewer's disagreement get silently overridden without a real look.Why: seniority is a decent prior, not proof, and a system that always sides with it stops learning from its newer reviewers entirely.
Leave the claims where both reviewers already agree exactly as they are.Why: adding a third reviewer there would slow the eleven cases in twelve that were never actually in question.
How to answer this, stage by stage
Nobody is grading whether you can name a tiebreak rule. They're grading whether you know disagreement is information, not noise to smooth over.
Stage 1
Scope it to one real tool
Say it like this
"I'll answer this for Suresight, Redshale Mutual's tool that flags auto claims for a possible-fraud second read before two adjusters ever look at one."
Why this works
Gives the interviewer a concrete case instead of "reviewers" in the abstract.
Stage 2
Say your structure out loud
Say it like this
"I'll use PICK. Position, my pick before any reasoning. Impact, who feels each kind of error. Cost asymmetry, which error is actually worse. Kill criteria, what evidence would change my mind."
Why this works
Signals a method for a question that otherwise invites a rambling policy discussion.
Stage 3
State the position first
Say it like this
"When two reviewers disagree, I send it to a named third adjudicator who sees both of their reasoning, not an average of their two scores and not an automatic win for whoever's more senior."
Why this works
This is the direct answer, committed to before any of the reasoning behind it, exactly what PICK's first letter is for.
Stage 4
Name who feels each error
Say it like this
"A wrongly delayed real claim costs one customer weeks of frustration. A real fraud that slips through because two reviewers split costs Redshale money spread thin across many other claims, and nobody ever traces it back to this one case."
Why this works
Names both sides in real units instead of leaving the tradeoff abstract.
Stage 5
Name the cost asymmetry
Say it like this
"Averaging two scores or letting seniority win looks efficient, but it throws away the one thing disagreement actually tells you, that this case is genuinely hard. I'd optimize against that hidden cost, not the loud, visible one."
Why this works
The hardest step in PICK, and the one that separates a real answer from a policy-sounding one.
Stage 6
Give the kill criteria
Say it like this
"If disagreement on a specific claim type climbs past around one in three, I stop treating it as an individual escalation problem. At that rate, the policy itself is the ambiguous thing, not the two people reading it."
Why this works
Shows the pick isn't stubborn, there's a real number that would flip it into a different kind of fix.
Stage 7
Prove it with the story, compressed
Say it like this
"Cassandra and Nikolai once split on a claim that turned out to be a staged-accident pattern. Averaging their two scores would have let it pay out at the exact midpoint. Only a real third look would have caught it."
Why this works
Grounds the abstract cost asymmetry in one real, compressed failure.
Stage 8
Close on the one line
Say it like this
"Disagreement is a signal, not noise. Averaging it away or handing it to seniority throws that signal out. Routing it to a real third look keeps it, and tracking it turns a recurring argument into an actual policy fix."
Why this works
Restates the direct answer in one breath, ready for whatever the interviewer pushes on next.
Let's learn
What happens when two people who are both good at their job look at the exact same claim and land in different places?
Suresight is Redshale Mutual's fraud-likelihood tool. It reads a filed auto claim and flags the ones worth a closer look before an adjuster opens the file.
Before Suresight, an adjuster reviewed every claim alone, on instinct and a shared training manual, with no second read unless a claim got kicked upstairs for its size.
Now Suresight flags roughly one claim in twelve as worth a second look, and on most of those, two adjusters land on the same call within minutes of each other.
Only one corner of this grid is where a disagreement actually needs a named third look.
Cost per error type, per 100 disputed claims
The missed-fraud bar is more than three times the size of the delayed-claim bar. That's the asymmetry PICK asks you to optimize against.
The disagreement itself was never the actual problem. The problem was what Redshale did the moment two honest, careful people looked at the same case and didn't agree.
At its worst: two adjusters split on a genuinely hard, ambiguous claim, the system quietly averages their two scores into one number in the middle, and a case that needed a real third look instead gets the exact answer nobody actually gave it.
The decision I would take back
We let disagreement resolve by simple average of the two adjusters' fraud scores, because it was the easiest rule to build and it matched how the two scores were already stored. That made sense when disagreements were rare and usually close calls anyway. It stopped making sense the day two scores landed far apart for a real reason.
What I would leave alone: the roughly eleven of twelve flagged claims where both adjusters already agree. Adding a third reviewer there would slow down cases that were never actually in question.
The lesson: disagreement between two careful reviewers is not two failures to average away. It's the one signal telling you a case is genuinely hard, and averaging it away throws out the exact information you most needed to keep.
Now here is the same thing as a story
The short version above is what you'd say defending this design to Redshale's claims committee. Read this one for how the averaging rule actually failed.
Cassandra Voss has adjusted auto claims at Redshale Mutual for eleven years. She can spot a staged-accident pattern from a repair estimate alone, long before Suresight ever flags one. Nikolai Petrenko joined the claims floor eight months ago, still building that same instinct, one file at a time.
For its first year, Suresight's flags mostly agreed with themselves: whichever adjuster looked at a flagged claim usually landed close to where a second adjuster would have.
Then a rear-end collision claim came in with a repair estimate that read almost too clean, and Suresight flagged it for a second read.
Four days from filed to paid. The split on day two never got a real second look before day four.
Cassandra read the claim as a likely staged pattern, the estimate matched three other claims from the same body shop over the past year. Nikolai read it as a legitimate accident, the customer's story checked out clean and their own claims history was spotless.
Knowledge spark: why would averaging two honest scores be worse than trusting either one alone?
Averaging isn't neutral, it throws away the fact that a real disagreement happened at all. A score exactly in the middle looks calm and confident. It isn't. It's two people who each had a real reason to land somewhere else entirely.
Suresight's system, as built, took the average of Cassandra's high fraud score and Nikolai's low one, landed in the "approve with monitoring" band, and the claim paid out.
Nobody at Redshale decided to approve that claim. A formula that had never met either adjuster did.
Three months later, the same body shop showed up on two more flagged claims, this time with a state investigator's letter attached. The pattern Cassandra had seen the first time was real, and the average had quietly erased it.
The old design skipped straight from the split to a formula. The new one adds one real box in between.
The adjudicator gets four things, not two numbers to split the difference on.
With a real adjudicator step, the same disagreement routes to a claims supervisor who reads both of their notes, not just their two numbers, before any payout goes out. Run the same rear-end claim forward: the supervisor sees Cassandra's body-shop pattern match sitting right next to Nikolai's clean customer history, asks the body shop for records on the other claims, and holds the payout until that comes back.
The old design asked two honest scores to cancel each other out. The new one asks a person to actually look at why they didn't.
I built the averaging rule because it felt fair, neither adjuster's judgment outweighed the other's. It took a body shop's second and third claim to see that fair-sounding math had thrown away the one piece of information that mattered most: that someone with eleven years of pattern-matching had seen something real.
PICK, in one screenNot a tiebreak rule. PICK is what tells you why averaging two honest scores is the wrong instinct.
P
Position. The pick, before the reasoning.
Route disagreement to a named third adjudicator with both reviewers' full reasoning, not an average and not automatic seniority.
Commits to an answer before the case for it, exactly what an interviewer is listening for.
I
Impact. Who feels each error, in real units.
A wrongly delayed real claim costs one customer weeks of frustration and an appeal. A real fraud that slips through costs Redshale money spread thin across many other claims, traced back to no single decision.
Names both sides in units a person actually feels, not abstract error rates.
C
Cost asymmetry. The one that's actually worse.
Averaging or defaulting to seniority throws away the one signal disagreement carries: that the case is genuinely hard. That hidden cost, a real pattern quietly erased into a calm-looking average, outweighs the visible cost of one slower claim.
The hardest step and the direct answer: this is what separates PICK from a coin flip.
K
Kill criteria. What would change the pick.
If disagreement on a specific claim type passes roughly one in three, it's no longer an individual escalation problem. It's the policy itself that needs rewriting.
Separates a confident answer from a stubborn one, since a real number could flip it.
One error is loud and lands on one customer. The other is quiet and lands everywhere at once.
One number decides whether this is still an adjudication problem or has become a policy problem.
The recap, one line per letter: position is a named adjudicator over an average, impact is one delayed customer against fraud spread thin, cost asymmetry is the hidden pattern an average erases, and kill criteria is the one-in-three line that turns individual escalation into a policy rewrite.
And if you want to be sure it really works, try it somewhere elseSame four letters, two radiologists instead of two adjusters. A completely different field, the same averaging trap.
Ferrymount Imaging Center uses an AI system to flag scans with a possible finding worth a second read. Dr. Helena Brandt has read scans for sixteen years. Dr. Amadou Traore joined the practice last year.
Mapped onto PICK: position is a named third radiologist for a real tiebreak read, never an average of two confidence scores and never automatic deference to seniority. Impact is a missed finding costing a patient a delayed diagnosis, against an over-called finding costing a patient an unnecessary biopsy and weeks of fear. Cost asymmetry is that averaging two radiologists' confidence on a genuinely ambiguous nodule erases the exact information that this particular case needed a specialist's eyes, not a formula. Kill criteria is watching disagreement rate by nodule type: if one specific category keeps splitting readers, the read protocol for that category needs rewriting, not another round of second opinions.
Reviewer disagreement rate by claim type, against the policy-rewrite line
Only the staged-accident category crosses the line. That's the one claim type whose written policy actually needs a rewrite.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "route disagreement to a real third look, never an average, and watch the rate by claim type," and stop.
Cost: there's no budget for a dedicated adjudicator role this quarter. Say so honestly, and rotate the adjudicator duty among senior adjusters for now, since the point is a real second look, not a new headcount line.
The model gets better, for real: if Suresight's fraud score becomes far more accurate, that's still not a reason to trust an average of two disagreeing humans more, the model getting better doesn't tell you which of the two people was right on this specific case.
Where people run it wrong.
They build a tiebreak rule that treats disagreement as noise to smooth over, instead of information to chase down.
They let seniority silently win by default, which quietly stops the system from ever learning anything from its newer reviewers.
They watch overall disagreement rate instead of disagreement rate by category, and miss the one category that's actually broken.
How to use it live. When someone asks how you'd handle two reviewers disagreeing, ask yourself one question first: does my fix use the disagreement, or does it make the disagreement disappear. If it makes it disappear, you've built an average with a fancier name.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a "how would you handle disagreement" tradeoff question?
Tap to flip
ANSWER
PICK: position, impact, cost asymmetry, kill criteria. The cost asymmetry step is what rules out averaging as an answer.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Cassandra Voss, an eleven-year adjuster at Redshale Mutual, and Nikolai Petrenko, who joined the claims floor eight months ago.
3 · THE OLD RULE
How did Suresight resolve a disagreement before the fix?
Tap to flip
ANSWER
It averaged the two adjusters' fraud scores into one number, which landed the claim in an "approve with monitoring" band that neither adjuster had actually chosen.
4 · THE ASYMMETRY
Which error is worse: a wrongly delayed claim, or a missed fraud?
Tap to flip
ANSWER
Missed fraud. It costs about three times as much per 100 disputed claims, and it's spread thin across other cases instead of landing visibly on one customer.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Letting disagreement resolve by simple average, because it was the easiest rule to build and matched how the two scores were already stored.
6 · THE NUMBER
Fill in the blank: once disagreement on a claim type passes roughly ___, treat it as a policy problem, not an adjudication problem.
Tap to flip
ANSWER
One in three. Below that, individual escalation works. Above it, the two people disagreeing are both reading an ambiguous rule correctly.
7 · THE REPLAY
Same rear-end claim, redesigned adjudication. What changes?
Tap to flip
ANSWER
A supervisor sees both adjusters' reasoning side by side, asks the body shop for its other claim records, and holds the payout until that comes back, instead of letting an average quietly approve it.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the parallel?
Tap to flip
ANSWER
Ferrymount Imaging Center's scan-flagging system. Same PICK shape: two radiologists' disagreement gets a named third read, not an averaged confidence score.
Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: Redshale's old rule resolved a disagreement by taking the ___ of the two adjusters' fraud scores.
Show hint
Look at "the decision I would take back."
Show answer
Average. Averaging felt fair, since neither adjuster's score outweighed the other's, but it erased the fact that a real disagreement had happened at all.
Multiple choice
2. Why is averaging two disagreeing reviewers' scores the wrong move, according to this answer?
A. Because averages are always mathematically incorrect.
B. Because Nikolai's score is always less trustworthy than Cassandra's.
C. Because it erases the one signal disagreement gives you, that the case is genuinely hard and needs a real look.
D. Because PICK requires every disagreement to escalate no matter what.
Show hint
Look at the cost asymmetry step.
Show answer
C. A calm-looking average number hides the fact that two people, each with a real reason, landed somewhere completely different.
True or false
3. True or false: once disagreement on a claim type passes the one-in-three line, the fix is to retrain the reviewers.
True
False
Show hint
Look at the kill criteria step.
Show answer
False. At that rate, the reviewers are reading the same ambiguous policy correctly. The fix is to rewrite the policy, not retrain the people reading it.
Short answer, where it wouldn't matter
4. Name a case where this exact adjudication step would not need to apply.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The roughly eleven of twelve flagged claims where both adjusters already agree. Nothing about those cases is actually in dispute.
Short answer, apply it yourself
5. Pick a product you use yourself. Name a habit it built in you that you'd stop doing if the product got a little worse.
Show hint
Think of an app you stopped double-checking because it was reliable for a long stretch.
Show answer
Model answer: Many people stop double-checking a GPS route, an auto-categorized expense, or a spell-checker's suggestion, once it's been right often enough in a row.
Short answer, name the reversal
6. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Resolving disagreement by averaging the two scores. It made sense while disagreements were rare and usually close calls, and stopped making sense the day two scores landed far apart for a real reason.
Before you close the answer
Why this works
Tests whether you treat disagreement as information worth chasing down, or noise worth smoothing over. Most candidates reach straight for a tiebreak rule.
Follow-up traps
"Isn't a named adjudicator just adding more process and cost?" Response: only on the small slice of claims that already disagree, roughly one in twelve times one in several. The eleven of twelve claims with no disagreement never touch this step at all.
"What if the third adjudicator also disagrees with both of them?" Response: that's exactly the signal that pushes a claim type toward the policy-rewrite side of the kill criteria, three-way disagreement is stronger evidence the guidance itself is ambiguous, not weaker.
If pressed
Redshale's real adjudication queue also logs which of the two original adjusters the adjudicator sided with, over time. That log is what eventually flags a reviewer who's drifting from the group consensus, not just a case that was hard.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.