CalculationAdvancedDesigning for Uncertainty & Trust / Feedback loops and data flywheels / #11

How would you sample feedback to avoid a vocal minority skewing your view?

BOUND the product is Kestrel Mutual's claims copilot, which drafts a plain-language summary of every insurance claim for the adjuster reviewing it

Kestrel Mutual's claims copilot reads a claim file and drafts a plain-language summary for the adjuster handling it, with a thumbs up or down under every draft. Petra Vansova is the senior adjuster who also owns the feedback dashboard leadership reads every month.

The direct answer
Don't just collect more feedback. Design who gets counted. Cap how many ratings any one adjuster can contribute to the official view, then actively solicit a fixed, random quota across claim type and tenure every week, instead of waiting for whoever feels strongly enough to click. A sample built this way looks like the real caseload. The one Kestrel had before only looked like its loudest adjuster.
Do this, in order
  1. Cap any one adjuster's contribution to the official sample, no matter how many they submit.Why: one adjuster alone supplied 28% of last week's ratings, all from one claim type he happens to specialize in.
  2. Stratify by claim type and adjuster tenure before sampling, not after.Why: a flat average can hide a claim type that never gets rated at all, buried inside a healthy-looking overall number.
  3. Actively solicit a fixed random quota each week, instead of waiting on organic clicks.Why: three out of four adjusters submit zero ratings most weeks, so waiting for them just recreates the same skew.
  4. Pick a margin of error the real weekly volume can actually support, and say so plainly.Why: a tight 5% margin needs 385 ratings a week; the honest volume points to a looser, still useful bar closer to 150.
  5. Re-check the sample's shape every week, not just once at rollout.Why: whoever dominated last week can dominate again the moment nobody's watching the cap.

How to answer this, stage by stage

Nobody's grading whether you know the phrase "sampling bias." They're grading whether you can show the actual arithmetic behind a trustworthy number.

Stage 1
Scope it to one real dashboard
Say it like this
"I'll answer this for Kestrel Mutual's claims copilot, where every AI-drafted summary gets a thumbs up or down from the adjuster reviewing it."
Why this works
Grounds an abstract sampling question in one real, numbers-having dashboard.
Stage 2
Say your structure out loud
Say it like this
"I'll use BOUND. Break it down, own the numbers, use a range, nail the sanity check, name the direction it swings."
Why this works
Signals a real estimation method instead of a hand-wavy "get more data" answer.
Stage 3
Break down the equation
Say it like this
"Sample size for a trustworthy proportion equals z-squared times p times one-minus-p, divided by margin of error squared. That's the equation deciding how many ratings I actually need."
Why this works
States the arithmetic out loud before touching a single number.
Stage 4
Own the assumptions with real numbers
Say it like this
"220 adjusters, about 2,400 claims a week. At 95% confidence and a 5% margin, that equation needs 385 ratings. We only get 340, and one person supplies 96 of them."
Why this works
BOUND's own numbers step: every figure stated, and where it came from.
Stage 5
Give a range, not a single number
Say it like this
"Depending on how tight a margin I demand, the real target sits somewhere between about 150 and 385 ratings a week, properly distributed, not raw."
Why this works
A single number here implies a confidence the data doesn't actually support.
Stage 6
Run the sanity check
Say it like this
"340 raw ratings a week sounds like plenty, until you notice one adjuster alone supplied 96 of them. Volume without design isn't the same thing as a sample."
Why this works
Compares the number to something everyone already understands is wrong.
Stage 7
Name what swings the estimate most
Say it like this
"The margin of error I demand swings the needed sample size by about nine times, from 43 to 385. That single choice matters more than anything else in this design."
Why this works
What a good estimator says, and a weaker one skips entirely.
Stage 8
Close on the one line
Say it like this
"Design who gets counted before you worry about how many. A capped, stratified sample of 150 beats a raw pile of 340 dominated by one person, every time."
Why this works
Restates the direct answer in one breath, ready for a follow-up push.

Let's learn

Picture ten adjusters, and only one of them ever clicks anything at all.

Kestrel Mutual's claims copilot reads a claim file, the incident report, photos, and prior notes, and drafts a plain-language summary for the adjuster to review before making a decision. A thumbs-up and thumbs-down icon sits under every draft.

Hand sketched flow diagram titled How the sample gets built. Five steps: pull all claims, stratify by type and tenure highlighted, sample within each stratum, weight back to true mix, report.
A trustworthy sample isn't collected. It's built, on purpose, in this order.

Before anyone looked closely, Kestrel's monthly dashboard reported one flat number: an average thumbs-up rate across every rating anyone ever submitted, no cap, no weighting, whoever clicked, counted.

Where last week's 340 ratings actually came from
340 170 0 1 rater: 96 Next 9: 124 Other 210: 120 340 total ratings, last week
Ten adjusters out of 220 supplied two thirds of last week's ratings. One of them alone supplied more than a quarter.

The turn: the size of that 340 was never the problem. A flat number that size can look completely healthy while still describing almost nobody's actual experience.

Three hundred and forty ratings sounds like a real sample. It's actually one adjuster's opinion, counted ninety six times, plus whatever was left over.
Knowledge spark: what's a stratified sample? Instead of pulling feedback from whoever shows up, you split the population into known groups first, claim type, tenure, region, and pull a fixed amount from each group on purpose. It's the difference between fishing wherever the fish happen to bite and checking every part of the pond.

At its worst: leadership reads a healthy 91% thumbs-up rate in a board deck, unaware it's built almost entirely on one adjuster's opinion about one claim type, while a different claim type Kestrel is actively expanding into has never been rated by anyone at all.

The decision I would take back We built the dashboard to average every rating ever submitted, with no cap on how many times one adjuster could rate and no weighting for how typical their caseload was. That made sense when the copilot only handled a small pilot group and everyone's clicks were roughly equal in number. It stopped making sense once the tool rolled out to all 220 adjusters and a handful of them started rating dozens of times a week while most rated never.

What I would leave alone: a single adjuster flagging one specific claim as flat-out wrong is still worth acting on immediately, on its own merits. The sampling design is about the aggregate quality number, not about ignoring an individual, specific, correct complaint.

The lesson: a bigger raw number of ratings isn't the same thing as a trustworthy one. The honest question was never "how do we get more feedback." It was "who exactly is allowed to count, and how much."

Now here is the same thing as a story

The short version above is what you'd say defending the redesigned dashboard to Kestrel's leadership. Read this one for how the gap actually got found.

Petra Vansova has adjusted auto and property claims at Kestrel Mutual for eleven years, and for the last two she's also owned the monthly quality dashboard for the claims copilot.

Hand sketched comparison diagram titled Loud versus quiet. Left panel a person icon labeled 1 adjuster, caption 40 ratings, all outliers. Right panel a person icon labeled 40 adjusters, caption 0 ratings, typical week.
Both of these describe a real week at Kestrel. Only one of them ever showed up on the dashboard.

A new data analyst, three weeks into the team, was building her first version of the weekly report when she asked Petra a plain question: "wait, are we just counting whoever complained loudest?"

Hand sketched quadrant titled Who actually shows up to rate. Axes how often they rate from never to constantly, and how typical their caseload from atypical to typical. Power rater sits constantly and atypical. Silent majority sits never and typical.
The adjusters actually rating things are, on average, the least typical part of the whole caseload.

Petra pulled the raw numbers to check. Two hundred twenty adjusters handle roughly 2,400 claims a week between them. Only 340 of those claims ever got a rating, and one adjuster, who specializes in a narrow slice of auto-body estimates, had submitted 96 of them himself.

Knowledge spark: what does a margin of error actually buy you? It's how far off your estimate could plausibly be from the true number. A 5% margin means the real value is very likely within 5 points of what you measured. A looser margin, say 15%, needs far fewer ratings, but the number it gives you can wander a lot further from the truth.

Using the standard formula for a trustworthy percentage, a 95% confidence level with a 5% margin of error needs 385 ratings. Kestrel had 340 total, and a huge share of those came from one person.

Ratings needed for a trustworthy estimate, by margin of error
5% margin 385 8% margin 150 15% margin 43
Choosing the margin of error is the single decision that swings the needed sample size the most, by close to nine times end to end.

Petra ran the sanity check out loud in the room: 340 sounds like a real number, until you notice a quarter of it came from one man who rates almost every claim he touches, on a claim type most adjusters barely see.

A sample isn't real just because the count is big. It's real when nobody in it can single-handedly move the number.

Petra's team redesigned the collection: cap any one adjuster's counted ratings at three per week, then actively prompt a random, fixed quota across five claim types and three tenure bands, targeting about 150 properly distributed ratings a week, an honest 8% margin instead of a fantasy 5% the current volume could never support.

Hand sketched labeled parts diagram titled What's inside a stratified sample. Center document icon labeled Sample plan, with four callouts: claim type, adjuster tenure, region, fixed size per group.
Four small design choices, and together they turn 340 unequal opinions into one number worth trusting.

Run the same new hire's question forward under the new design: the answer is simply no, because no single adjuster can supply more than three of the counted ratings in any given week, whatever the raw click count behind the scenes says.

Petra built the flat, uncapped dashboard because it felt neutral, just averaging whatever came in. It took one new analyst's plain question to see that "count everything equally" isn't neutral at all when some people show up forty times more often than others.

BOUND, the actual arithmeticNot a vibe check. BOUND is what makes a sampling design defensible with real numbers, not just good intentions.

B
Break it down. State the equation.
Sample size equals z-squared times p times one-minus-p, divided by margin of error squared.
Names the actual math before touching a single number.
O
Own the numbers.
220 adjusters, 2,400 claims a week, 340 raw ratings, 96 of them from one person.
Every figure stated plainly, with where it came from.
U
Use a range.
Roughly 150 to 385 ratings a week, properly distributed, depending on how tight a margin is demanded.
A single number here would imply confidence the data can't back up.
N
Nail the sanity check.
340 sounds like plenty, until a quarter of it turns out to be one person's opinion, repeated.
Compares the number to something everyone already knows is off.
D
Direction. What swings it most.
The margin of error you choose swings the needed sample from 43 to 385, roughly nine times over.
The hardest step, and what a good estimator says that a weaker one skips.

The recap, one line per letter: break it down is the standard sample-size equation, own the numbers is Kestrel's real 220 adjusters and 340 weekly ratings, use a range is the honest 150 to 385 span, nail the sanity check is one man supplying a quarter of the total, and direction is the margin-of-error choice that swings the estimate the most.

And if you want to be sure it really works, try it somewhere elseSame five letters, an HVAC dispatch floor instead of a claims desk. A different trade, the same arithmetic.

Ductwise suggests a likely cause to an HVAC technician mid-repair, based on a description of the symptoms. Reuben Okafor runs field service operations for the company using it, and a small handful of senior technicians rate almost every suggestion, while most technicians rate almost none.

Mapped onto BOUND: break it down is the same sample-size equation, applied to however many technicians and repair calls the company runs each week. Own the numbers might look like 85 technicians, 900 repair calls a week, and 60 ratings, 22 of them from one senior technician who always rates Ductwise's compressor-fault suggestions specifically. Use a range gives a target between roughly 40 and 95 properly distributed ratings a week, depending on the margin of error the team can live with. Nail the sanity check is noticing that one technician's specialty, compressor faults, is wildly overrepresented while duct-leak calls, a third of all repairs, have barely been rated at all. Direction is the same finding: the margin-of-error choice swings the needed count far more than anything else in the design.

Hand sketched icon list titled Three signs your view is skewed. Items: one adjuster rates dozens of times, one claim type never gets rated, ratings spike right after one bad week.
The same three warning signs show up whether the caseload is insurance claims or HVAC repair tickets.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "cap each rater's contribution, then actively sample a fixed quota by group, at a margin of error the real volume supports," and stop.
Cost: there's no engineering time to build active solicitation prompts this quarter. Say so honestly, and start by simply capping and reweighting the existing organic ratings first, since even that beats a flat, uncapped average.
The model gets better, for real: if the copilot's summaries get consistently more accurate, the sampling design still matters, since even a small remaining error rate needs an honest, unskewed view to catch it at all.

Where people run it wrong.
They treat a bigger raw number of ratings as automatically more trustworthy, without checking who supplied them.
They demand a tight margin of error without checking whether real weekly volume can actually support it.
They stratify the analysis after the fact instead of designing the sample around the strata from the start.

How to use it live. When someone asks how you'd sample feedback fairly, ask yourself one question first: could any single person, on their own, move this week's number. If the honest answer is yes, that's the fix, right there.

Hand sketched timeline titled Same sample, checked weekly. Three milestones: raw feedback swings hard, stratified sample holds steady highlighted, reported weekly trusted view.
The design only earns the word trusted once it's been checked the same way, week after week.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "how would you sample this feedback" question?
Tap to flip
ANSWER
BOUND: break it down, own the numbers, use a range, nail the sanity check, direction. Direction is the hardest, most-skipped step.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Petra Vansova, an eleven-year claims adjuster at Kestrel Mutual who also owns the claims copilot's monthly quality dashboard.
3 · THE GAP FOUND
What did checking the raw ratings reveal?
Tap to flip
ANSWER
One adjuster alone had submitted 96 of the week's 340 ratings, almost entirely on one narrow claim type he specializes in.
4 · THE EQUATION
What's the sample-size equation this answer uses?
Tap to flip
ANSWER
Sample size equals z-squared times p times one-minus-p, divided by the margin of error squared.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Averaging every rating ever submitted with no cap per adjuster and no weighting, a design that made sense only back when everyone's click counts were roughly equal.
6 · THE NUMBER
Fill in the blank: at 95% confidence and a 5% margin of error, this answer's equation needs ___ ratings.
Tap to flip
ANSWER
385. Kestrel only had 340 raw ratings that week, and a large share of those came from a single adjuster.
7 · THE REPLAY
Same new-hire question, redesigned sample. What changes?
Tap to flip
ANSWER
The answer becomes a flat no, since any one adjuster is capped at three counted ratings a week, no matter how many they actually submit.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the parallel?
Tap to flip
ANSWER
Ductwise's HVAC repair suggestions. Same BOUND shape: one senior technician dominates the ratings on his own specialty.

Check yourself Score: 0 / 0

Multiple choice
1. Why doesn't Kestrel's raw 340 weekly ratings already count as a trustworthy sample?
  • A. Because 340 is too small a number to ever be useful.
  • B. Because the count is dominated by a small number of frequent raters, so it doesn't reflect the real caseload.
  • C. Because the claims copilot's model needs retraining.
  • D. Because thumbs-up ratings are never useful for insurance products.
Show hint
Look at the stacked bar chart of who supplied last week's ratings.
Show answer
B. A large count can still badly misrepresent the population if a small number of people supply most of it.
Fill in the blank
2. Fill in the blank: at 95% confidence and a 5% margin of error, this answer's equation needs ___ ratings a week.
Show hint
Look at "own the numbers."
Show answer
385. Kestrel only collected 340 raw ratings that week, and a large share of those came from a single adjuster.
True or false
3. True or false: the fix in this answer is simply to collect a much larger number of raw ratings.
  • True
  • False
Show hint
Look at the direct answer's opening line.
Show answer
False. Just collecting more would likely let the same vocal minority supply even more of the total. The fix caps and stratifies who gets counted.
Short answer, where it wouldn't matter
4. Name a case in this answer where a single, uncapped rating is still worth acting on immediately.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A single adjuster flagging one specific claim as flat-out wrong is still worth acting on right away. The sampling design is about the aggregate quality number, not individual, correct complaints.
Short answer, apply it yourself
5. Think of a review or rating system you've used. Who do you think actually left most of the reviews, and were they typical of everyone who used the product?
Show hint
Think about who bothers to leave a review: usually people with an unusually strong opinion, good or bad.
Show answer
Model answer: Most public review systems share this exact skew: reviewers tend to be the most satisfied or most frustrated users, rarely the typical, quiet middle.
Multiple choice
6. If the required margin of error is loosened from 5% to 15%, what happens to the needed sample size?
  • A. It stays exactly the same, since margin of error doesn't affect sample size.
  • B. It drops sharply, from 385 down to about 43, since a looser margin needs far less data.
  • C. It increases, since a looser margin needs more evidence to be safe.
  • D. It can only be answered by collecting live data first.
Show hint
Look at "direction," the step naming what swings the estimate most.
Show answer
B. Margin of error swings the required sample size dramatically, roughly nine times over between the 5% and 15% cases here.
Before you close the answer
Why this works
Tests whether you can turn "avoid a vocal minority" into real arithmetic, a cap, a stratified quota, and an honest margin of error, instead of a vague call to "collect more data."
Follow-up traps
"Isn't a smaller, capped sample of 150 just throwing away good data?" Response: no data is thrown away, the cap only limits how much one rater counts toward the official view; every rating can still be reviewed individually for real, specific issues.

"What if the power rater's ratings are actually the most accurate ones?" Response: that's a separate, worthwhile question, but it doesn't change the sampling math; even a highly accurate rater shouldn't single-handedly define the whole population's experience.
If pressed
The real rollout used a rotating, randomized prompt schedule so an adjuster never knows in advance which claim will be the one Kestrel actively asks them to rate that week, which keeps the active solicitation itself from becoming gameable.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more