CaseAdvancedDesigning for Uncertainty & Trust / Feedback loops and data flywheels / #1

Design the feedback mechanism for an AI feature where users rarely click thumbs down.

The direct answer
Do not add a thumbs up or thumbs down button. Make the user tap one plain "send" action before anything goes live, and count every edit made before that tap as the feedback. The tap proves she looked, and the edits show you what was wrong, without ever asking her opinion.
Do this, in order
  1. Require one plain "send to CRM" tap before the draft goes live, and count every edit made before that tap as the feedback.Why: an edit only means something once you know she actually looked. The tap is what turns silence into a real signal instead of a guess.
  2. Track corrections field by field, not the summary as a whole.Why: a single wrong name inside an otherwise clean draft can hide for weeks behind one overall pass or fail count.
  3. Never show the edit count against her name to anyone above her, on day one.Why: the moment fixing a draft looks like being graded, she stops fixing it in the open and starts hiding the correction instead, the exact silence this whole design exists to end.
  4. Never ask her to explain why she changed something.Why: a "why did you edit this" prompt on every fix rebuilds the same friction as the thumbs down button, just with extra steps.
  5. Leave the raw call recording and transcript alone.Why: nobody reads the transcript directly. Only the drafted summary gets acted on, so a feedback layer on the transcript would catch nothing anyone would use.
  6. Watch the correction rate on each field over time, not the number of thumbs down clicks.Why: it is the number that would have caught a wrong field three weeks before anyone ran a manual check by hand.

How to answer this, stage by stage

Seven moves. Say each one out loud once before you try to use it for real.

1
Put one real person in the question before you say anything else
Say it like this
"Let me make this concrete. Say it's a tool that listens to a sales call and drafts the summary, the next steps, and the CRM updates afterward. The person using it is an account executive on five or six calls a day, with about twenty minutes total to deal with the paperwork from every one of them."
Why this works
An interviewer cannot grade "make it easier to give feedback." Naming the product and the person buys you room for a real decision instead of a general one.
2
Say your structure out loud before you use it
Say it like this
"I'll cover why a thumbs down button fails on its own, the one thing I'd build instead, what breaks the first time I'm wrong about it, and what I'm not going to touch on day one."
Why this works
Two seconds of structure stops you rambling, and it lets the interviewer steer you toward the part they actually want to hear.
3
Reframe what a quiet button actually means
Say it like this
"The instinct is to make the button bigger, or nag her for a rating, or give her points for using it. But a click was never free. It costs a decision and a few seconds she doesn't have between two calls. Silence isn't agreement. It's just the cheaper option."
Why this works
Most candidates answer with "make feedback easier to give." This line names the real cost instead of the surface problem, and it's what separates you from them.
4
Give one concrete decision, not a principle
Say it like this
"Here's the one thing I'd build. One plain button that sends the draft to the CRM, nothing fancier. Every edit she makes before she taps it gets logged automatically as feedback. She never rates anything. Her corrections are the rating."
Why this works
A specific decision can be argued with, which is exactly what the interviewer wants to do next. "Make feedback easier" can't be argued with, because it doesn't commit to anything.
5
Prove it with a failure, in four sentences
Say it like this
"Here's what happens without it. She quietly fixes a wrong 'next step' field for weeks and nobody ever sees it. Then sales ops spot checks twenty of her records against the call recordings and finds three with a wrong field. She stops trusting any of the auto filled fields and spends an evening re-checking two weeks of calls by hand."
Why this works
A concrete failure persuades where a principle doesn't, and it proves you've thought past the happy path, which is most of what a design question is testing.
6
Say what you would not build, on purpose
Say it like this
"I wouldn't ask her to explain each edit, and I wouldn't put the edit count on a leaderboard with her name on it. Both turn a normal fix into something she has to defend, and the moment fixing looks like grading, people stop doing it where you can see it."
Why this works
Naming what you'd leave out shows judgment. A candidate who wants every safeguard at once hasn't actually decided anything.
7
Close on the one line, so it's the last thing they hear
Say it like this
"So: log the edit, not the opinion. A thumbs down button asks her to stop and grade something. An edit log grades it for her, using work she's already doing on every single call."
Why this works
Interviewers remember your first and last lines most clearly. Don't let a strong answer trail off into a list of features.
If you remember one thing Stages 3 and 4 are the answer. Reframe why silence isn't agreement, then name the one plain thing you'd build. Everything else in this answer is proof.

Let's learn

What does it mean when almost nobody presses a button?

Picture a tool that listens in on a sales call. When the call ends, it writes three things: a summary of what got said, a list of next steps, and the fields a CRM needs filled in, like which stage the deal just moved to.

Before a tool like this existed, the person on the call wrote all of that herself, right after she hung up. Six calls a day, about twelve minutes of typing after each one. That's seventy two minutes a day spent turning a conversation into a record.

Now the draft shows up about ninety seconds after the call ends. In the early weeks she reads every line, fixes the two or three things that are wrong, and taps send. That whole pass takes about three minutes instead of twelve. Seventy two minutes a day drops to about eighteen.

That's not the interesting part either.

About one in ten drafts has something wrong in it somewhere: a number misheard, a next step assigned to the wrong person, a name spelled the way it sounds instead of the way it's written. That rate barely moves over time. The model isn't getting worse. What's quietly changing is how closely anyone is still looking for the mistake, and there is no button on that screen that costs less than just fixing it and moving on.

Two panels comparing real mistakes in about one in ten drafts against reported mistakes, almost none ever
The gap the whole design has to close
Next step owner field, corrected out loud before send
13%
Week 2
2%
Week 11
Her own catch and fix rate fell hard. Whether the real mistakes fell with it is a different question, and nothing on this chart can answer it. That's exactly the gap that let three wrong records sit undetected until someone finally checked by hand.

Here is where it goes wrong at its worst. A next step field says "customer to send the signed order" when really the customer is waiting on a quote from her. Nobody flagged it, because there was never a cheap way to flag anything. Weeks later, a regional forecast built on those fields shows a deal as healthy that has actually gone quiet.

The mistakes were never the real problem. The problem is that nothing anywhere records when one happens.

The old system of handwritten notes was messy, but everyone knew to double check it. This one looks clean and confident and wrong, and that is worse than the messy version ever was.

The choice I would take back. When this feature was being built, someone added a thumbs up and a thumbs down icon next to every draft, because that's what almost every AI feature ships with. It looked like the safe, obvious choice, and nobody in the room argued with it. I would take it out, put one plain send button in its place, and count every edit made before that tap as the real feedback.

What I would leave alone. The raw recording of the call, and the rough transcript underneath the summary. Nobody opens either one to check it. The summary is the only part anyone reads and acts on, so that's the only part that needs a way to say something was wrong.

Knowledge spark: what a click actually costs A feedback button isn't free just because it's small. On a busy day it's competing with the next call on her list, not with silence, and the next call almost always wins.

The lesson. If a feature only gets graded when someone is angry enough to click something, you aren't measuring whether it works. You're measuring whether someone had a bad enough day to say so. Most people never have that day. They just quietly fix it and move on, call after call, and you never see any of it happen.

Now here is the same thing as a story

You don't need this part to answer the question. Read it when you want to feel why the short version is true.

Every weekday, right after her last call, Rina Okafor used to sit with a legal pad and turn what she remembered into six sets of CRM notes. She sells ovens, walk in coolers, and prep line equipment to restaurant chains, and she is good at the part of the job nobody can teach: she can hear from a facilities director's silence, on the phone, whether he is actually deciding or just being polite.

Then the tool arrived. For the first six weeks, it gave her back the best part of that ritual. A draft would land while she was still writing her own follow up email. She read every line of it anyway, at first out of habit, then because she liked catching the small things: a supplier's name typed wrong, a delivery date one week off. She fixed what needed fixing and sent it on.

By week seven she was reading the top and the middle and trusting the rest. By week ten she was reading the first line, mostly to check it had the right client name, and letting the rest go. By week eleven, most days, she never opened it at all. It synced itself. She was already dialing the next call before the draft had even finished loading.

A timeline in three beats: reads every line in weeks one to six, skims it in weeks seven to ten, never opens it from week eleven on
The same habit, three ways of doing it

Nothing dramatic caused any of that. Nobody told her to stop checking. She just had six calls and twenty minutes, and reading a finished draft stopped feeling like the twenty minutes' best use.

Then, on a Tuesday in her twelfth week, the sales ops lead, a man named Colin Ferraro, ran the audit he runs every month: he pulls twenty records at random and checks them against the actual call recordings. It's routine. He does it whether anything is wrong or not.

Three of Rina's twenty records had the next step field wrong. Not badly wrong. In each one, the field said the customer owed her something, when really she owed the customer a revised quote. Small, easy to miss, exactly the kind of thing a busy person glances past.

Colin didn't accuse her of anything, and he was right not to. She had done nothing careless. She had just stopped being able to afford the twenty minutes, the same as anyone would. But once he showed her the three records, something in her flipped, all at once, for every record, not just those three. She stopped trusting any of the auto filled fields, going back weeks. That evening, instead of going home, she pulled up recordings from the last two weeks and checked every single one by hand, the way she used to. It took her past nine o'clock.

The three wrong fields cost Colin twenty minutes to find. Getting her trust back cost her an entire evening she didn't have, and it didn't even fix the actual problem, because nobody could tell her how far back the mistakes really went.

She never had a number in her head for any of this. She had a feeling, and it only had two settings: I can glance at this, or I have to read all of it myself. Three records in a spot check flipped the whole switch at once, for records that were never even checked.

I keep thinking about the meeting where the thumbs up and thumbs down icons went in. Somebody on the team said, almost as an aside, "every AI feature has these, right," and three other people nodded, and it went into the spec that afternoon. Nobody in that room was wrong to think it was harmless. It just quietly promised something it could never deliver: that if something went wrong, someone would tell us.

So here is what I would build instead, back when this was still a screen in a design file.

Ask for a rating, where she has to stop and grade it, compared with logging the edit, where the fix is the signal
Stop asking. Start counting the edit.

Take the two icons off. In their place, one plain button: send to CRM. Every edit Rina makes to a field before she taps it gets logged, quietly, against that call. She never rates anything. Her corrections are the rating.

Now walk the same Tuesday forward with that version instead. By week two, the correction log already shows the next step owner field getting fixed on thirteen out of every hundred calls, a real number, moving, three weeks before Colin ever runs his audit. He doesn't need the audit to catch it anymore. He sees it on a dashboard on a Thursday morning, opens two of the flagged calls to listen, and asks someone to check the wording the model uses for that field. It gets fixed by Friday. Nobody loses an evening.

We didn't need her opinion. We needed her hands, and her hands were telling us everything, every single day. We just weren't writing any of it down.

A thumbs down button asks her to stop and grade something. An edit log grades it for her, using work she was already doing on every call, whether anyone was watching or not.

SPARK, step by step, on this exact draft

This is a design question, so the framework is SPARK. A question that opens with "what if the error rate doubled" would reach for FLIPS instead. Different question shapes need different tools.

SPARK laid out as five rows: situation, payoff, anchor, risk, and keep out
SPARK, applied to Rina's draft
S, situation. Rina Okafor, account executive, sells kitchen equipment to restaurant chains. Six calls a day. Wrote her own call notes and CRM updates by hand, about twelve minutes a call.
P, payoff. Not "get more feedback." The habit worth building is that she keeps fixing the two or three things that are actually wrong, in place, the moment she sees them, without ever having to decide if something is wrong enough to report.
A, anchor. One plain send button, and every edit made before that tap logged as feedback. No rating, ever asked for.
R, risk. The first day she stops reading the draft at all, zero edits looks exactly like a perfect draft. The system can't tell the difference between agreement and neglect.
K, keep out. Don't ask her why she changed something, and don't put the edit count anywhere near her name on day one. Both turn a normal fix into something she has to defend.
Why the anchor and the risk have to match Check them against each other: does one plain send tap actually defend against the day she stops reading? Only if the tap itself proves she looked, separate from any edit she makes. That's why the anchor isn't just "log the edits." It's "log the edits, and require the tap." Without the tap, a rushed day and a clean day produce the exact same silence.

And if you want to be sure it really works, try it somewhere else

A city building department drafts inspection reports with a similar tool. Same framework. A very different reason nobody flags a bad one.

An inspector flagging a report himself versus a report that gets kicked back later, the real signal already there
Same framework, a different reason to stay quiet

S. Dale Kestenbaum, one of about forty inspectors for a mid size city. Eight to ten site visits a day, a written report due on each one. He used to write them by hand at his desk, about twenty five minutes a report.
P. The habit worth building: he keeps lightly editing each drafted report before he files it, instead of either rubber stamping the whole thing or rewriting it from nothing.
A. A different anchor here, because flagging a bad draft isn't free the way it is for Rina. For Dale, marking one wrong opens his entire report history to a supervisor review. So instead of asking him to flag anything, the design watches which of his filed reports get kicked back for revision later, weeks after the fact, and treats that kickback as the real signal.
R. If kickbacks take three or four weeks to surface, and the drift started six weeks ago, the signal always lags the real problem by about a month. A design built for Rina's fast loop doesn't survive at Dale's slower one without a change.
K. Don't build a formal "why was this report inaccurate" review flow on day one. For someone whose job security runs through these reports, that flow guarantees he'll never touch it, which is exactly the silence this whole design exists to end.

Swap the trigger and it still runs

  • Speed: the draft takes eight seconds instead of ninety. Nothing about the fix changes, because a tap still costs less than typing the whole thing herself.
  • Cost: the company starts charging per summary generated. She starts skipping it for her smallest accounts and typing those by hand, so the edit log slowly stops covering her whole pipeline, not just her busiest week.
  • The model gets better: the draft gets so reliable that she stops reading even the fields she used to always check first. The one time a big account's contact title is subtly wrong, it ships straight into a forecast with nobody the wiser.

Where people run it wrong

  • Making the send button itself feel like a rating, with a star count or a color that changes based on how much she edited. That turns the one safe, plain action back into something to avoid.
  • Putting the correction log in a report somebody has to remember to open. If it isn't next to the draft, on the same screen, it won't get checked either.
  • Writing the flagged pattern back to her in the tool's language instead of hers. "Field mapping mismatch on next_step_owner" tells an account executive nothing. "This field gets fixed a lot: next step owner" does.

How to use it live

If you're asked this cold, buy yourself ten seconds by naming the person first. "Let me put someone real in this. Say it's an account executive on six calls a day with twenty minutes for all the paperwork." That's not stalling. It's stage one of the answer, and it buys you the time to find the actual design decision instead of reaching for the first feedback button you can think of.

Flashcards (click a card to flip it)

1 · THE SITUATION
Who is this answer about, and what does her day look like without the tool?
Tap to flip
ANSWER
Rina Okafor, an account executive who sells kitchen equipment to restaurant chains. Six calls a day. Before this tool, she wrote her own call notes and CRM updates by hand, about twelve minutes a call.
2 · THE REFRAME
Why doesn't a quiet thumbs down button mean the drafts are good?
Tap to flip
ANSWER
Clicking it costs a decision and a few seconds she doesn't have between calls. Silence is the cheaper option, not proof of agreement.
3 · THE ANCHOR
What's the one design decision in this answer?
Tap to flip
ANSWER
One plain send to CRM button, with every edit made before that tap logged as the feedback. She's never asked to rate anything.
4 · THE RISK
What breaks the first time this design is wrong?
Tap to flip
ANSWER
The day she stops reading the draft at all. Zero edits then looks exactly like a perfect draft, and the system can't tell agreement from neglect.
5 · THE PROOF
What actually happened to Rina's trust, in four sentences?
Tap to flip
ANSWER
Sales ops spot checked twenty of her records and found three with a wrong field. Nothing dramatic, just enough to flip her. She stopped trusting every auto filled field, not just those three, and spent an evening re-checking two weeks of calls by hand.
6 · THE NUMBER
Sales ops found ___ of Rina's 20 spot checked records had a wrong field.
Tap to flip
ANSWER
3 of 20. Small enough that nobody panicked. Big enough that she stopped trusting every field the tool filled in, not just those three.
7 · THE REPLAY
Same Tuesday, new design. What changes?
Tap to flip
ANSWER
By week two, the correction log already shows the next step owner field getting fixed on 13 of every 100 calls, three weeks before the manual audit would have caught it. Sales ops sees the dashboard, checks two calls, and fixes the wording by Friday.
8 · CROSS-PRODUCT
Section 4 runs SPARK again on a different product. Which one, and what changes about the anchor?
Tap to flip
ANSWER
A city building inspector's report drafts. The anchor isn't edit tracking. It's watching which filed reports get kicked back for revision later, because flagging a bad draft himself puts his whole report history under review.

Check yourself Score: 0 / 0

True or false
1. True or false: wrong data started reaching the CRM because the AI got worse at writing summaries.
  • True
  • False
Show hint
Look at what actually changed: the model, or how closely anyone was checking it.
Show answer
False. The real mistake rate barely moved, about one in ten drafts the whole way through. What changed was how closely Rina was still checking for it.
Multiple choice
2. Which of these is the real design decision in this answer, and which is a principle wearing an answer's clothes?
  • A. "I would make it easier for her to give feedback."
  • B. "One plain send to CRM button, and every edit made before that tap logged as the feedback."
  • C. "I would focus on building more trust with the user."
  • D. "I would run a survey after each call."
Show hint
Which one could an interviewer actually push back on?
Show answer
B. A and C sound reasonable but commit to nothing specific. D adds a new ask instead of removing one. Only B names something concrete enough to argue with.
Short answer
3. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Think about what almost every AI feature ships with by default.
Show answer
Model answer: The thumbs up and thumbs down icons next to the draft. They made sense at the time because that's the default almost every AI feature ships with, and nobody in the room had a reason yet to think they would sit there unused.
Fill in the blank
4. Zero edits on a draft only means something once you know she actually ______.
Show hint
It's not about what she changed. It's about whether she was even looking.
Show answer
Looked. Opened and read the draft before sending it. Without that, zero edits could mean a perfect draft, or it could mean nobody checked at all.
Multiple choice
5. Which of these is a place in this same product where losing the ability to flag mistakes would NOT matter?
  • A. The raw call recording and transcript, which nobody reads directly.
  • B. The next step owner field.
  • C. The CRM deal stage field.
  • D. The customer contact's name and title.
Show hint
Ask who actually reads this part, and what they do with it.
Show answer
A. Nobody acts on the raw transcript itself, only on the summary built from it, so a feedback layer there would catch problems nobody would ever see.
Short answer, apply it yourself
6. Pick something else that drafts text for you, like an email reply suggestion or an auto generated caption. What do you fix almost every time without ever telling anyone? What would that fix be worth as feedback if someone logged it?
Show hint
Think of the one small thing you always change out of habit, not because anyone asked you to.
Show answer
Model answer: "Take a phone's suggested reply to a text. I always delete the exclamation mark it adds to anything even slightly serious. I've never once tapped 'not helpful.' If someone logged that one edit across a few million messages, they'd learn the tone is wrong for serious conversations, a fact a thumbs down count would never surface, because nobody bothers to tap it over one exclamation mark."
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more