ConceptAdvancedDesigning for Uncertainty & Trust / Feedback loops and data flywheels / #9
Explain the risk of optimizing directly on user feedback.
GUARD the product is Solvemate, an AI math tutor for high schoolers, where every step-by-step explanation carries a thumbs up or down
Solvemate walks a high schooler through a math problem, step by step, and a thumbs-up or thumbs-down sits under every explanation. Ines Calderon leads product for Solvemate and owns the decision of what those ratings are allowed to do to the model.
The direct answer
Never let raw student ratings decide which explanation gets reinforced. A student who needs the tutor most is exactly the student who can't tell a confident wrong answer from a correct one, only whether it felt clear. Gate every explanation behind an independent correctness check first, and only let ratings choose among explanations already confirmed correct.
Do this, in order
Gate correctness first, before student ratings can influence anything.Why: a struggling student can judge clarity, not correctness, so a rating alone can't tell you which one it actually is.
Track error rate by rating tier, not just the overall average.Why: a real error rate can hide entirely inside your best-loved explanations if you only ever look at the aggregate.
Run a subject-matter human review specifically on your highest-rated explanations.Why: popularity is exactly where an error hides best, since nobody suspects a well-liked answer.
Watch the confident-but-wrong share compound across retraining rounds, not just once.Why: this risk builds slowly and rarely announces itself with one dramatic failure.
Report correctness rate and clarity rate side by side on every dashboard.Why: quoting one number without the other is exactly how the two got confused for each other in the first place.
How to answer this, stage by stage
Nobody's grading whether you know feedback can be gamed. They're grading whether you can name who actually gets hurt when it is.
Stage 1
Scope it to one real product
Say it like this
"I'll answer this for Solvemate, a math tutor where a thumbs up or down sits under every step-by-step explanation a student sees."
Why this works
Grounds an abstract risk question in one real feedback loop, not a general warning.
Stage 2
Say your structure out loud
Say it like this
"I'll use GUARD. Groups, unequal, ability to contest, reduce, detect. Ability to contest is the hardest step, and the one this question is really testing."
Why this works
Signals a method for a question that could otherwise turn into a vague warning about bias.
Stage 3
Name the two groups and the one lever
Say it like this
"Whoever clicks the rating holds the lever. Every future student served by whatever explanation won gets the outcome, with no lever of their own."
Why this works
GUARD's opening move: naming the operator and the subject on the same page.
Stage 4
Show where the harm lands unevenly
Say it like this
"A confident, wrong explanation reads as clear to a student who doesn't have the background to catch the error. That student rates it well, and a correct but plainer explanation loses."
Why this works
Names exactly which group absorbs the risk, not a vague population.
Stage 5
Ask who never gets to contest it
Say it like this
"A student weak in the topic can't flag 'this reasoning is actually wrong,' because that's exactly the judgment they came to the tutor for help with."
Why this works
GUARD's hardest step. This is the real risk, not just "the model might be biased."
Stage 6
Give the specific design change
Say it like this
"Gate every explanation behind an independent correctness check before ratings can promote it at all. Ratings only choose among explanations already confirmed correct."
Why this works
A real product decision, not a policy statement or a training session.
Stage 7
Say how you'd detect it in production
Say it like this
"Track error rate by rating tier, and watch the share of confident-but-wrong explanations round over round, so it surfaces before a parent or a news story finds it first."
Why this works
Shows you'd catch this internally, not rely on someone outside the company to flag it.
Stage 8
Close on the one line
Say it like this
"Optimizing directly on feedback optimizes for whoever can rate. When the people who need the help most can't judge correctness, you end up rewarding confidence, not truth."
Why this works
Restates the direct answer, ready for a follow-up push.
Let's learn
Here's what happens when you let a rating decide what's true.
Solvemate walks a high schooler through a math problem one step at a time. After each explanation, a thumbs-up or thumbs-down sits at the bottom of the screen. Early on, the team used those ratings the simplest way possible: whichever phrasing of an explanation got more thumbs-up became the version shown more often to the next student on that same problem.
There's a gap in this loop where correctness should sit, and right now nothing fills it.
For a long stretch, this looked like it was working. Ratings climbed steadily, quarter over quarter, and the team read that as students getting more out of every explanation.
The turn: the ratings climbing was never proof the math was getting more correct. It was proof the explanations were getting better at sounding clear, and those are not the same thing.
A rating tells you an explanation felt clear. It never once tells you whether it was true.
Knowledge spark: what's a proxy metric?
A stand-in for the thing you actually care about, used because it's easier to measure. A thumbs-up is a proxy for "this helped." The risk shows up the moment a system starts chasing the proxy instead of the real goal.
Error rate in explanations, by student rating tier
The best-loved explanations had over four times the error rate of the ones students liked less. Popularity and correctness were pulling in opposite directions.
At its worst: a wrong method for solving a quadratic gets reinforced across thousands of students because it reads as confident and friendly, while a correct but slightly more formal explanation quietly loses ground and stops being shown at all.
The decision I would take back
We set the default so that whichever explanation phrasing won more thumbs-up became the one shown more often, assuming a satisfied student was the same thing as a correct explanation. That made sense when the tutor only covered simple arithmetic, where almost any clear explanation was also a correct one. It stopped making sense once Solvemate covered algebra and calculus, where a wrong method can sound just as clear as a right one.
What I would leave alone: using student ratings to choose between two explanations that are already both mathematically correct, say, a shorter one versus a more detailed one, is exactly the right use of a rating. The risk is only in letting a rating decide correctness itself.
The lesson: a feedback signal is only safe to optimize on directly when the people giving it can actually judge the thing you're asking them to judge. The moment that stops being true, optimizing on it rewards confidence over correctness, quietly and for a long time before anyone notices.
Now here is the same thing as a story
The short version above is what you'd say defending this fix to Solvemate's leadership. Read this one for how the gap actually got found.
Solvemate's engineering floor is quietest around nine at night, when the week's retraining job kicks off, folding in a fresh batch of student ratings from the past few days.
There was no single bad Tuesday here. No customer complaint, no viral screenshot, no one incident anyone could point to. The risk built up the ordinary way: a little at a time, every week, for the better part of a year and a half.
One of these two people holds a lever. The other one just lives with whatever the lever produced.
Ines Calderon ran a quarterly correctness audit that had been on the calendar for two years, nothing special about this particular quarter. A rotating group of subject-matter math teachers reviewed a sample of two hundred of Solvemate's highest-rated explanations, the ones with a thumbs-up rate above ninety percent.
The dangerous quadrant isn't the one anyone was watching for. It's the one that looks fine on a ratings dashboard.
Thirty four of those two hundred, seventeen percent, contained an actual error in the math reasoning. Every one of them had been shown to students hundreds of times, each time reinforcing the same wrong method a little further.
Seventeen out of every hundred of the tutor's best-loved explanations were teaching the wrong method, confidently enough that no one had thought to check.
The teachers noticed something else, too: the errors clustered in a specific pattern. A confident, friendly tone, paired with a subtly wrong step buried in the middle, exactly the kind of mistake a student who's still learning the topic has no way to catch.
A rating is one number standing in for four different questions, and only one of them is the one that matters here.
Ines's team gated the retraining loop: an independent correctness check now runs on every explanation first, and only explanations that pass it are eligible for ratings to promote at all.
Share of served explanations that are confident but wrong, round over round
Left alone, the confident-but-wrong share more than doubled across three retraining rounds. Nothing about that climb ever showed up in the thumbs-up average.
Run the same eighteen months forward with the gate in place: the confident-but-wrong share never gets to climb past its round-one baseline, because nothing reaches a student until a correctness check, not a rating, has already cleared it.
Ines built the ratings-first design because it felt like listening to students directly, and that felt like the right instinct at the time. It took a routine audit, not a crisis, to see that listening to students about clarity and trusting them to judge correctness were two very different things wearing the same thumbs-up icon.
GUARD, the two people in the roomNot a fairness checklist. GUARD is what forces you to name who's rating and who's living with the result.
G
Groups. Who's affected.
Whoever clicks the rating, and every future student served by whatever explanation the ratings promoted.
Names the operator and the subject on the same page, not a vague "users."
U
Unequal. Where the harm lands.
A confident, wrong explanation reads as clear to a student who can't catch the error, and it wins over a correct but plainer one.
The harm lands hardest on exactly the students who need the tutor most.
A
Ability to contest. Who can't push back.
A student weak in the topic can't flag "this reasoning is wrong," since that judgment is exactly what they came to the tutor for.
The hardest step, and the real risk in the question, not just "the model might be biased."
R
Reduce. The design change.
Gate every explanation behind an independent correctness check before ratings can promote it at all.
A concrete product decision, not a policy statement.
D
Detect. How you'd know.
Track error rate by rating tier and the confident-but-wrong share round over round, so it surfaces before a parent or a headline does.
The routine quarterly audit is what actually caught this, not a single incident.
Nothing about this needs a single bad day. A small compounding drift, left unchecked, gets you to the same place.
The recap, one line per letter: groups is the rater and the future student who never rated anything, unequal is the confident-but-wrong explanations winning over correct plain ones, ability to contest is a struggling student who can't tell the difference, reduce is the correctness gate before any rating counts, and detect is watching error rate by tier instead of trusting the average.
And if you want to be sure it really works, try it somewhere elseSame five letters, a warehouse instead of a classroom. A different floor, the same lever.
Shiftline suggests weekly shift schedules to a warehouse's operations lead, and workers rate each suggested schedule with a thumbs up or down before it's finalized. Godwin Asare runs operations at a distribution site using it.
Mapped onto GUARD: groups are the workers who bother to rate a schedule, usually whoever got an easy shift that week, and the workers who end up on whatever pattern the ratings favor, including the ones who rarely rate anything at all. Unequal is a schedule that quietly loads more weekend and night shifts onto a quieter subgroup, since the vocal raters keep reinforcing whichever pattern favors them personally. Ability to contest is the night-shift crew, who rarely log into the rating system at all and have no lever over which pattern keeps winning. Reduce is the same design move as Solvemate's: gate any schedule change behind a fairness check on hour distribution across the whole team, not just the loudest raters, before ratings can influence which pattern repeats. Detect is watching the actual spread of night and weekend hours by worker, not just the overall thumbs-up rate on the schedule tool.
The same three questions catch this whether the output is a math step or a shift schedule.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "gate correctness before ratings can promote anything, since raters who need help most can't judge correctness," and stop.
Cost: there's no budget yet for a full independent correctness pipeline. Say so honestly, and start with a smaller, manual spot-check on the highest-rated explanations first, since a thin gate still beats none.
The model gets better, for real: if Solvemate's base model gets sharper and makes fewer errors overall, the gate still matters, since even a small remaining error rate reinforced at scale over years is still real damage.
Where people run it wrong.
They treat a high rating as proof of quality without ever separately checking correctness.
They wait for a single dramatic failure instead of running a routine, scheduled audit.
They fix the model's confidence without ever gating what ratings are allowed to reinforce in the first place.
How to use it live. When someone asks about optimizing on user feedback, ask yourself one question first: can the person giving this feedback actually judge the thing I'm about to let it decide. If the honest answer is no, that's the risk, right there.
Only one of these three branches is the good outcome, and a raw rating can't tell you which branch you're actually on.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits an "explain this risk" question about optimizing on feedback?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. Ability to contest is the hardest step.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Ines Calderon, product lead at Solvemate, an AI math tutor for high schoolers with a thumbs-up widget on every explanation.
3 · THE TWO GROUPS
Who holds the lever, and who lives with the result?
Tap to flip
ANSWER
Whoever clicks the rating holds the lever. Every future student served by whatever explanation won lives with the result, with no lever of their own.
4 · WHO CAN'T CONTEST
Why can't a struggling student push back on a wrong explanation?
Tap to flip
ANSWER
Because judging whether the math is correct is exactly the skill they came to the tutor to build. They can rate clarity, not correctness.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Letting whichever explanation got more thumbs-up become the one shown more often, assuming a satisfied student meant a correct explanation.
6 · THE NUMBER
Fill in the blank: explanations rated above 90% thumbs-up had a ___% math error rate.
Tap to flip
ANSWER
17%. Lower-rated explanations, under 70% thumbs-up, had only a 4% error rate, the opposite of what the ratings implied.
7 · THE REPLAY
Same eighteen months, with the correctness gate in place. What changes?
Tap to flip
ANSWER
The confident-but-wrong share never climbs past its round-one baseline, since nothing reaches a student until correctness clears it first.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the parallel?
Tap to flip
ANSWER
Shiftline's warehouse shift scheduler. Same GUARD shape: vocal raters favor a pattern that quietly disadvantages a quieter shift group.
Check yourself Score: 0 / 0
Fill in the blank
1. Fill in the blank: explanations rated above 90% thumbs-up carried a ___% math error rate, versus 4% for explanations rated under 70%.
Show hint
Look at the grouped bar chart of error rate by rating tier.
Show answer
17%. Popularity and correctness moved in opposite directions here, which is exactly why optimizing directly on the rating is risky.
Multiple choice
2. Why can't the risk here be fixed by just asking students to rate more carefully?
A. Because students don't care about getting the right answer.
B. Because judging correctness requires the exact skill the student is still learning, not more effort or attention.
C. Because the rating widget is technically broken.
D. Because Solvemate doesn't collect enough ratings.
Show hint
Look at "ability to contest," GUARD's hardest step.
Show answer
B. A student weak in the topic can't verify correctness no matter how carefully they rate, since that judgment is the exact thing they're still learning.
True or false
3. True or false: it's always risky to let students choose between two different explanations using a thumbs-up rating.
True
False
Show hint
Look at "what I would leave alone."
Show answer
False. It's safe to let ratings choose between two explanations that are already both confirmed correct. The risk is only in letting a rating decide correctness itself.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Letting the highest thumbs-up explanation win by default. It made sense when Solvemate only covered arithmetic, where a clear explanation was almost always a correct one too.
Short answer, apply it yourself
5. Think of a product you use that asks for a star rating or a thumbs-up. Is there a case where the rater genuinely couldn't judge the thing they're rating?
Show hint
Ask whether the rater has the specific expertise the rating is really about.
Show answer
Model answer: This shows up anywhere a rater judges "did this feel right" as a stand-in for "was this actually right," which is common in medical, legal, or technical AI tools.
Multiple choice
6. What would be the first sign, in production, that this risk is actually happening?
A. The overall thumbs-up average starts dropping sharply.
B. Error rate specifically among the highest-rated explanations climbs, even while the overall average looks healthy.
C. Students stop using the tutor entirely within a week.
D. The company receives a formal legal complaint.
Show hint
Look at "detect," GUARD's step for catching this in production.
Show answer
B. This risk hides inside the well-liked explanations specifically, so the aggregate rating can look fine right up until an audit by rating tier catches it.
Before you close the answer
Why this works
Tests whether you can name who actually gets hurt when a proxy metric is optimized directly, instead of giving a generic warning about "bias" or "gaming metrics."
Follow-up traps
"Isn't a correctness gate just adding more process and slowing things down?" Response: it adds one verification step before promotion, not before every single interaction, so the added cost is one delayed rollout cycle in exchange for not reinforcing wrong methods at scale.
"Couldn't you just show students the correct answer key next to the rating?" Response: no, most students are using the tutor precisely because they can't independently verify the answer key themselves, so more context doesn't fix a rater who can't judge correctness.
If pressed
The real correctness gate runs a second, independent model checking each derivation step against the answer key, separate from the explanation-writing model, so the same system isn't grading its own homework twice.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.