Design the feedback mechanism for an AI feature where users rarely click thumbs down.
- Require one plain "send to CRM" tap before the draft goes live, and count every edit made before that tap as the feedback.Why: an edit only means something once you know she actually looked. The tap is what turns silence into a real signal instead of a guess.
- Track corrections field by field, not the summary as a whole.Why: a single wrong name inside an otherwise clean draft can hide for weeks behind one overall pass or fail count.
- Never show the edit count against her name to anyone above her, on day one.Why: the moment fixing a draft looks like being graded, she stops fixing it in the open and starts hiding the correction instead, the exact silence this whole design exists to end.
- Never ask her to explain why she changed something.Why: a "why did you edit this" prompt on every fix rebuilds the same friction as the thumbs down button, just with extra steps.
- Leave the raw call recording and transcript alone.Why: nobody reads the transcript directly. Only the drafted summary gets acted on, so a feedback layer on the transcript would catch nothing anyone would use.
- Watch the correction rate on each field over time, not the number of thumbs down clicks.Why: it is the number that would have caught a wrong field three weeks before anyone ran a manual check by hand.
How to answer this, stage by stage
Seven moves. Say each one out loud once before you try to use it for real.
Let's learn
What does it mean when almost nobody presses a button?
Picture a tool that listens in on a sales call. When the call ends, it writes three things: a summary of what got said, a list of next steps, and the fields a CRM needs filled in, like which stage the deal just moved to.
Before a tool like this existed, the person on the call wrote all of that herself, right after she hung up. Six calls a day, about twelve minutes of typing after each one. That's seventy two minutes a day spent turning a conversation into a record.
Now the draft shows up about ninety seconds after the call ends. In the early weeks she reads every line, fixes the two or three things that are wrong, and taps send. That whole pass takes about three minutes instead of twelve. Seventy two minutes a day drops to about eighteen.
That's not the interesting part either.
About one in ten drafts has something wrong in it somewhere: a number misheard, a next step assigned to the wrong person, a name spelled the way it sounds instead of the way it's written. That rate barely moves over time. The model isn't getting worse. What's quietly changing is how closely anyone is still looking for the mistake, and there is no button on that screen that costs less than just fixing it and moving on.
Here is where it goes wrong at its worst. A next step field says "customer to send the signed order" when really the customer is waiting on a quote from her. Nobody flagged it, because there was never a cheap way to flag anything. Weeks later, a regional forecast built on those fields shows a deal as healthy that has actually gone quiet.
The old system of handwritten notes was messy, but everyone knew to double check it. This one looks clean and confident and wrong, and that is worse than the messy version ever was.
The choice I would take back. When this feature was being built, someone added a thumbs up and a thumbs down icon next to every draft, because that's what almost every AI feature ships with. It looked like the safe, obvious choice, and nobody in the room argued with it. I would take it out, put one plain send button in its place, and count every edit made before that tap as the real feedback.
What I would leave alone. The raw recording of the call, and the rough transcript underneath the summary. Nobody opens either one to check it. The summary is the only part anyone reads and acts on, so that's the only part that needs a way to say something was wrong.
The lesson. If a feature only gets graded when someone is angry enough to click something, you aren't measuring whether it works. You're measuring whether someone had a bad enough day to say so. Most people never have that day. They just quietly fix it and move on, call after call, and you never see any of it happen.
Now here is the same thing as a story
You don't need this part to answer the question. Read it when you want to feel why the short version is true.
Every weekday, right after her last call, Rina Okafor used to sit with a legal pad and turn what she remembered into six sets of CRM notes. She sells ovens, walk in coolers, and prep line equipment to restaurant chains, and she is good at the part of the job nobody can teach: she can hear from a facilities director's silence, on the phone, whether he is actually deciding or just being polite.
Then the tool arrived. For the first six weeks, it gave her back the best part of that ritual. A draft would land while she was still writing her own follow up email. She read every line of it anyway, at first out of habit, then because she liked catching the small things: a supplier's name typed wrong, a delivery date one week off. She fixed what needed fixing and sent it on.
By week seven she was reading the top and the middle and trusting the rest. By week ten she was reading the first line, mostly to check it had the right client name, and letting the rest go. By week eleven, most days, she never opened it at all. It synced itself. She was already dialing the next call before the draft had even finished loading.
Nothing dramatic caused any of that. Nobody told her to stop checking. She just had six calls and twenty minutes, and reading a finished draft stopped feeling like the twenty minutes' best use.
Then, on a Tuesday in her twelfth week, the sales ops lead, a man named Colin Ferraro, ran the audit he runs every month: he pulls twenty records at random and checks them against the actual call recordings. It's routine. He does it whether anything is wrong or not.
Three of Rina's twenty records had the next step field wrong. Not badly wrong. In each one, the field said the customer owed her something, when really she owed the customer a revised quote. Small, easy to miss, exactly the kind of thing a busy person glances past.
Colin didn't accuse her of anything, and he was right not to. She had done nothing careless. She had just stopped being able to afford the twenty minutes, the same as anyone would. But once he showed her the three records, something in her flipped, all at once, for every record, not just those three. She stopped trusting any of the auto filled fields, going back weeks. That evening, instead of going home, she pulled up recordings from the last two weeks and checked every single one by hand, the way she used to. It took her past nine o'clock.
The three wrong fields cost Colin twenty minutes to find. Getting her trust back cost her an entire evening she didn't have, and it didn't even fix the actual problem, because nobody could tell her how far back the mistakes really went.
She never had a number in her head for any of this. She had a feeling, and it only had two settings: I can glance at this, or I have to read all of it myself. Three records in a spot check flipped the whole switch at once, for records that were never even checked.
I keep thinking about the meeting where the thumbs up and thumbs down icons went in. Somebody on the team said, almost as an aside, "every AI feature has these, right," and three other people nodded, and it went into the spec that afternoon. Nobody in that room was wrong to think it was harmless. It just quietly promised something it could never deliver: that if something went wrong, someone would tell us.
So here is what I would build instead, back when this was still a screen in a design file.
Take the two icons off. In their place, one plain button: send to CRM. Every edit Rina makes to a field before she taps it gets logged, quietly, against that call. She never rates anything. Her corrections are the rating.
Now walk the same Tuesday forward with that version instead. By week two, the correction log already shows the next step owner field getting fixed on thirteen out of every hundred calls, a real number, moving, three weeks before Colin ever runs his audit. He doesn't need the audit to catch it anymore. He sees it on a dashboard on a Thursday morning, opens two of the flagged calls to listen, and asks someone to check the wording the model uses for that field. It gets fixed by Friday. Nobody loses an evening.
A thumbs down button asks her to stop and grade something. An edit log grades it for her, using work she was already doing on every call, whether anyone was watching or not.
SPARK, step by step, on this exact draft
This is a design question, so the framework is SPARK. A question that opens with "what if the error rate doubled" would reach for FLIPS instead. Different question shapes need different tools.
And if you want to be sure it really works, try it somewhere else
A city building department drafts inspection reports with a similar tool. Same framework. A very different reason nobody flags a bad one.
S. Dale Kestenbaum, one of about forty inspectors for a mid size city. Eight to ten site visits a day, a written report due on each one. He used to write them by hand at his desk, about twenty five minutes a report.
P. The habit worth building: he keeps lightly editing each drafted report before he files it, instead of either rubber stamping the whole thing or rewriting it from nothing.
A. A different anchor here, because flagging a bad draft isn't free the way it is for Rina. For Dale, marking one wrong opens his entire report history to a supervisor review. So instead of asking him to flag anything, the design watches which of his filed reports get kicked back for revision later, weeks after the fact, and treats that kickback as the real signal.
R. If kickbacks take three or four weeks to surface, and the drift started six weeks ago, the signal always lags the real problem by about a month. A design built for Rina's fast loop doesn't survive at Dale's slower one without a change.
K. Don't build a formal "why was this report inaccurate" review flow on day one. For someone whose job security runs through these reports, that flow guarantees he'll never touch it, which is exactly the silence this whole design exists to end.
Swap the trigger and it still runs
- Speed: the draft takes eight seconds instead of ninety. Nothing about the fix changes, because a tap still costs less than typing the whole thing herself.
- Cost: the company starts charging per summary generated. She starts skipping it for her smallest accounts and typing those by hand, so the edit log slowly stops covering her whole pipeline, not just her busiest week.
- The model gets better: the draft gets so reliable that she stops reading even the fields she used to always check first. The one time a big account's contact title is subtly wrong, it ships straight into a forecast with nobody the wiser.
Where people run it wrong
- Making the send button itself feel like a rating, with a star count or a color that changes based on how much she edited. That turns the one safe, plain action back into something to avoid.
- Putting the correction log in a report somebody has to remember to open. If it isn't next to the draft, on the same screen, it won't get checked either.
- Writing the flagged pattern back to her in the tool's language instead of hers. "Field mapping mismatch on next_step_owner" tells an account executive nothing. "This field gets fixed a lot: next step owner" does.
How to use it live
If you're asked this cold, buy yourself ten seconds by naming the person first. "Let me put someone real in this. Say it's an account executive on six calls a day with twenty minutes for all the paperwork." That's not stalling. It's stage one of the answer, and it buys you the time to find the actual design decision instead of reaching for the first feedback button you can think of.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Feedback loops and data flywheels
- #2 Explain the difference between explicit and implicit feedback signals.
- #3 What implicit signals tell you an output was bad?
- #4 How do you avoid a feedback loop that only captures complaints?
- #5 Describe how you would turn user edits into a quality signal.
- #6 What is the latency between collecting feedback and improving the product, and how do you shorten it?
- #7 Critique a thumbs up and down widget as a feedback mechanism.