InterviewAdvancedDesigning for Uncertainty & Trust / Feedback loops and data flywheels / #22

Design a feedback system for a product I name.

SPARK the named product is VetLoop, an ambient AI scribe that listens during a vet exam and drafts the clinical note for the vet to review at Kestrel Hollow Veterinary Clinic

The interviewer names the product: VetLoop, an ambient AI scribe. It listens during an exam and drafts a clinical note, findings, assessment, plan, for the vet to review, edit, and sign. Dr. Marisol Feng has practiced small-animal medicine for six years at Kestrel Hollow, sharing an exam-room tablet with two colleagues.

The direct answer
Design the feedback so it doesn't depend on anyone choosing to speak up. Every edit a vet makes to a VetLoop draft gets diffed and categorized automatically, at zero extra effort, whether it's a wording fix or a deleted finding that was never there. Add one optional, single tap a vet can use to mark an edit as clinically important, with no written explanation required and no name attached in any shared view. The automatic diff is the floor. The tap is the bonus. Neither one requires a vet to admit anything to anyone.
Do this, in order
  1. Capture every edit automatically, without requiring the vet to report anything.Why: a feedback design that depends on someone volunteering bad news will quietly stop working the moment the news feels embarrassing.
  2. Add one optional, single-tap flag for "this mattered," never a written form.Why: a vet mid-appointment will tap once; they won't stop to explain themselves.
  3. Keep the shared aggregate view unattributed to any individual vet by name.Why: the moment a correction count sits next to someone's name, reporting starts looking like an admission instead of a favor to the product.
  4. Keep a private, engineer-only link back to who reported what, for follow-up on real patterns.Why: full anonymity loses the ability to ask a reporting vet for case detail when something needs investigating.
  5. Hold off on full audio or video review, mandatory reports, or any named leaderboard for day one.Why: all three either cost too much trust or too much time, for a signal the automatic diff already gives you for free.

How to answer this, stage by stage

Nobody is grading whether you can name a feedback button. They're grading whether your design still works on the day someone would rather stay quiet.

Stage 1
Take the named product and scope it to one user
Say it like this
"You said VetLoop, so I'll design this for Dr. Marisol Feng, a vet at Kestrel Hollow who reviews and edits a VetLoop draft after every exam."
Why this works
Even with the product handed to you, the answer still needs one real person using it, not an abstract "clinicians."
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, how it's handled today. Payoff, the habit I want. Anchor, the one decision. Risk, what breaks it. Keep out, what waits."
Why this works
Signals a real method, so a feedback-system design doesn't turn into a list of buttons.
Stage 3
Show today, with no feedback design at all
Say it like this
"Right now, if VetLoop drafts something wrong, Dr. Feng just fixes it herself before signing off. Nothing about that correction goes anywhere else."
Why this works
The situation step. You can't design a fix for a gap you haven't shown existing.
Stage 4
Name the habit you actually want
Say it like this
"I want every real correction to leave a trace, without ever requiring a vet to feel like they're confessing something. The habit I want isn't 'report more.' It's 'never have to decide whether to.'"
Why this works
This is the payoff step, and it's the whole insight this question is testing for.
Stage 5
Give the one design decision
Say it like this
"The anchor: automatic diffing of every edit, zero effort, plus one optional single tap for 'this mattered,' with no name attached anywhere a peer or manager can see it."
Why this works
Matches the direct answer, stated as a concrete, arguable decision.
Stage 6
Prove the anchor survives the day someone stays quiet
Say it like this
"Say a vet is embarrassed by a recurring error and never taps the flag, ever. The diff still logs automatically. The pattern still shows up in aggregate, even from a vet who never says a word about it."
Why this works
The risk step. A feedback design that only works when people feel safe enough to speak isn't a design, it's a hope.
Stage 7
Say what you're not building yet
Say it like this
"Day one doesn't review full exam audio, doesn't require a written report, and doesn't put anyone's correction count on a shared leaderboard. All three cost more trust or more time than they're worth yet."
Why this works
The keep-out step. Naming real, deliberately excluded ideas shows judgment about scope.
Stage 8
Close on the one line
Say it like this
"So: capture every edit automatically, add one optional tap with no name attached, and never build a system that only works on the days people feel brave enough to use it."
Why this works
Restates the direct answer, ready to survive a follow-up push.

Let's learn

The tablet in Kestrel Hollow's second exam room gets passed between three vets a day, still warm from whoever used it last.

Before VetLoop, writing up a full exam note took Dr. Feng about twelve minutes per patient, typed between appointments. VetLoop cuts that to under two: it listens during the exam and drafts the note, findings, assessment, plan, for her to check and sign.

Hand sketched flow diagram titled Today, without a feedback design. Four steps: VetLoop drafts the note, Dr Feng spots an error, fixes it silently highlighted, nothing else happens.
Four steps, and the third one is where every bit of useful information currently disappears.

Right now, with no feedback design built at all, a wrong draft just gets fixed silently, the same way you'd fix a typo, and the correction goes nowhere else.

How VetLoop corrections were actually handled, before any redesign
100% 50% 0% 4% Reported 96% Fixed silently
Ninety-six percent of what actually went wrong left no trace anywhere the product team could ever see.

The turn: the silence was never about vets not noticing errors. Dr. Feng noticed every one of them; she's good at her job. The silence was about what reporting one would have said about her.

Hand sketched comparison diagram titled The day it's wrong. Left panel, a person icon labeled Vet says nothing, caption diff still logs automatically. Right panel, a gauge icon labeled Vet taps the flag, caption adds the why in three seconds.
Only one of these two paths depends on anyone feeling brave. The other one happens either way.
The decision I would take back We put every flagged VetLoop error into a shared quality log, reviewed monthly, sorted by which vet reported it. That felt like the transparent, accountable choice at launch. It stopped being reasonable the moment "flagged twelve AI errors this month" next to someone's name started reading as "relies on the tool too much," instead of what it actually was: a vet doing the product team a favor.

At its worst: the same rare-condition note error, VetLoop confidently writing an incorrect assessment for an uncommon presentation, gets independently caught and quietly fixed by several vets across the network, none of them telling anyone, for months, until it finally reaches a patient record uncorrected somewhere it shouldn't have.

What I would leave alone: a genuinely one-off typo or phrasing quirk doesn't need any of this. The automatic diff will show it as noise, and it should stay noise; the design exists for the patterns, not for every single keystroke.

The lesson: a feedback system that only works when people feel safe enough to use it isn't really a feedback system. It's a hope that nobody will ever have a reason to stay quiet.

Now here is the same thing as a story

The short version above is what you'd pitch to VetLoop's product team. Read this one for how the gap actually got found.

Dr. Marisol Feng has practiced small-animal medicine for six years, the kind of vet who can catch a subtle limp before an owner even mentions it.

For VetLoop's first several months, she used to mention it casually, in the break room, when a draft came out strange: "VetLoop wrote something weird about the Hendricks cat again." It was a joke, mostly, and it doubled as an accidental feedback channel nobody had designed on purpose.

Knowledge spark: why does a rare-condition blind spot count as an AI-specific failure, not bad luck? A model trained mostly on common cases will sound just as confident on a rare one it barely understands. The wrongness isn't random, it's the model's own long tail, the cases it's seen least, showing up as quiet, repeatable mistakes on exactly those cases.

VetLoop kept confidently misjudging the same rare, uncommon-in-general-practice condition, and as the joke wore thin, the "AI errors" log rolled out clinic-wide, reviewed monthly, sorted by name. Reporting an error stopped feeling like a joke and started feeling like an admission that she needed the tool's help too much.

Hand sketched timeline titled How the blind spot was found. Four milestones: errors start month 1, silent fixes spread months 2 to 5 highlighted, conference story month 6, feedback redesigned month 7.
Five months of quiet fixing, and the story from a colleague is the only reason anyone connected the pattern.
She didn't stop noticing the errors. She stopped telling anyone, the same month "flagged errors this month" started sitting next to her name where a colleague could see it.

At a regional veterinary conference, Dr. Feng ran into a colleague from a partner clinic in the same network, who mentioned, almost in passing, that a VetLoop-drafted note on a similar rare case had made it into a permanent record uncorrected, and had come up during an insurance audit.

Dr. Feng realized, standing there, that she'd caught and silently fixed the exact same kind of error at least five times over the past few months. If she'd said something the first time, an engineer might have caught the pattern before it ever reached her colleague's patient record at all.

The same rare-condition error, network-wide: silently fixed versus formally reported
41 20 0 Silently fixed: 41 Formally reported: 2 Month 1 Month 6
Forty-one quiet fixes across the network, and only two of them ever reached anyone who could actually do something about the pattern.

Run the same six months forward under the redesign: every one of those forty-one edits gets diffed and logged automatically, no vet ever needing to decide whether reporting makes them look bad. The pattern crosses a detection threshold by the ninth occurrence, network-wide, weeks before Dr. Feng's colleague's case ever reaches an audit.

I put the quality log out in the open, sorted by name, because transparency sounded like the responsible choice at launch. It took one conference conversation to see that a design depending on people feeling safe enough to admit a mistake isn't transparent. It's just quiet in a different, more dangerous way.

SPARK, in one screenNot a features list. SPARK is what forces a feedback design to survive the day someone would rather say nothing at all.

S
Situation. How it's handled today.
A wrong VetLoop draft gets fixed silently before signing, and the correction goes nowhere else at all.
Grounds the design in a real, current gap instead of a hypothetical one.
P
Payoff. The habit the design should build.
Every real correction leaves a trace, without any vet ever having to decide whether admitting it makes them look bad.
Not "report more," but never having to weigh that decision at all.
A
Anchor. The one design decision.
Automatic, zero-effort diffing of every edit, plus one optional, unattributed single tap for "this mattered."
The hardest step, and the direct answer: signal that doesn't depend on anyone choosing to speak.
R
Risk. What breaks it.
A vet embarrassed by a recurring error and staying silent forever. The diff still logs automatically either way, with or without her saying a word.
Proves the anchor survives the exact failure this question is really testing for.
K
Keep out. What waits for later.
Full audio or video review, mandatory written reports, a named leaderboard. Real ideas, all of them costing more trust or time than they're worth yet.
Shows judgment about scope, not a longer wish list.
Hand sketched labeled parts diagram titled The anchor close up, an edit event. Center document icon labeled Edit event, with four callouts around it: what changed, edit type, optional mattered, no name in shared view.
Four fields on every edit, and the fourth one is what makes the whole design survive silence.

The recap, one line per letter: situation is silent fixing with no trace, payoff is never having to decide whether to admit a mistake, anchor is automatic diffing plus an unattributed optional tap, risk is the design surviving total silence from any one vet, and keep out is the heavier tools deliberately left for later.

And if you want to be sure it really works, try it somewhere elseSame five letters, an HVAC dispatch tool instead of a vet clinic. Nothing else about the two jobs is alike.

Briarfield HVAC Services uses an AI tool that drafts a diagnosis and repair estimate for a technician to review before a customer sees it. Marguerite Bissette has dispatched technicians there for five years.

Mapped onto SPARK: situation is the same shape, a technician quietly correcting a wrong diagnosis before sending it to a customer, with the correction going nowhere else. Payoff: never needing to decide whether flagging a wrong estimate makes them look like they can't do the job without the tool. Anchor: the same two-part design, an automatic diff of every changed line item, plus one optional tap for "this mattered," unattributed in any shared view. Risk: a technician embarrassed by a recurring miss on a specific unit type stays silent forever, and the diff still surfaces the pattern regardless. Keep out: no mandatory written report, no named leaderboard, at least not yet.

Hand sketched quadrant titled Sorting feedback design options. Axes, effort required from the vet from none to heavy, and signal richness from thin to rich. Automatic diff sits low effort moderate to high richness. Optional one tap sits low effort high richness. Mandatory report sits heavy effort high richness. Do nothing sits low effort and low richness.
The two that made the cut both sit in the same corner: almost no effort, and still genuinely useful.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "capture every edit automatically, add one unattributed tap, never make anyone decide whether admitting a mistake is safe," and stop.
Cost: engineering says automatic diffing across every note costs a real build sprint. Say so, and start with the optional tap alone for v1, since even a partial signal beats none, while the automatic diff gets built next.
The model gets better, for real: if VetLoop's overall accuracy improves, this design still matters, because a rarer miss is even more likely to feel like an individual vet's fault rather than the model's, which makes silent concealment more likely, not less.

Hand sketched icon list titled What we kept out of day one. Three items: a box icon labeled full audio and video review, a document icon labeled mandatory written reports, a person icon labeled a public named error leaderboard.
Three real ideas, and every one of them would have cost more trust than it earned on day one.

Where people run it wrong.
They build a feedback channel that only works if someone volunteers embarrassing information about their own reliance on the tool.
They make the "responsible," transparent choice of full attribution, without noticing what that attribution actually costs socially.
They wait for a support ticket or a formal complaint, which a quietly concealed error, by definition, never generates.

How to use it live. When someone asks you to design a feedback system, ask yourself first: does this still work on the day the person using it would rather nobody knew? If the answer is no, the design isn't finished yet.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Concealment flip. Dr. Feng used to mention VetLoop's errors casually, then stopped mentioning them at all once a named quality log made reporting feel like an admission.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Dr. Marisol Feng, six years practicing small-animal medicine at Kestrel Hollow Veterinary Clinic.
3 · THE HABIT
What did Dr. Feng stop doing as the named quality log rolled out?
Tap to flip
ANSWER
She stopped casually mentioning VetLoop's odd drafts in the break room, and started fixing every error in total silence instead.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Mentioning a VetLoop error out loud to colleagues, versus silently fixing it and never telling anyone once reporting felt like an admission.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Putting flagged errors in a shared quality log sorted by vet name, meant to be transparent, which instead turned honest reporting into something to avoid.
6 · THE NUMBER
Fill in the blank: only ___% of real VetLoop corrections were ever formally reported to engineering.
Tap to flip
ANSWER
4%. Ninety-six percent were fixed silently, leaving no trace anywhere the product team could see the pattern.
7 · THE REPLAY
Same six months, redesigned feedback system. What changes?
Tap to flip
ANSWER
All forty-one edits get diffed and logged automatically, and the pattern crosses a detection threshold by the ninth occurrence, weeks before it ever reaches an audit.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what changed?
Tap to flip
ANSWER
Briarfield HVAC Services' diagnosis-drafting tool. Same automatic-diff-plus-optional-tap anchor, applied to repair estimates instead of clinical notes.

Check yourself Score: 0 / 0

Short answer, apply it yourself
1. Think of a mistake you've quietly fixed yourself rather than reported, in any tool you use for work. What would have had to be true for you to report it instead?
Show hint
Think about whether reporting it would have felt neutral, or would have said something about you.
Show answer
Model answer: Most people report more freely when doing so carries no social cost, no name attached, no implication that relying on the tool reflects poorly on them.
Multiple choice
2. Why does the automatic diff matter more than the optional one-tap flag, if you could only ship one?
  • A. Because the tap is harder to build than the automatic diff.
  • B. Because the automatic diff still produces a signal even from a vet who never chooses to say anything, while the tap depends on that choice.
  • C. Because vets never make mistakes worth flagging with a tap.
  • D. Because the tap requires a written explanation every time.
Show hint
Look at the risk step, R.
Show answer
B. The whole point of the anchor is a signal that doesn't depend on anyone feeling safe enough to speak up. The automatic diff is what survives total silence.
True or false
3. True or false: a public leaderboard showing which vet caught the most VetLoop errors would encourage more honest reporting.
  • True
  • False
Show hint
Look at "the decision I would take back."
Show answer
False. The named, visible quality log is exactly what turned honest reporting into something vets avoided. Attribution is the risk, not the fix.
Fill in the blank
4. Fill in the blank: across the network, the same rare-condition error was silently fixed ___ times by month 6, but formally reported only twice.
Show hint
Look at the line chart in the story.
Show answer
41. Forty-one quiet fixes across the network, and only two of them ever reached anyone who could act on the pattern.
Short answer, name the reversal
5. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Making the quality log shared and sorted by vet name, meant to be transparent and accountable. It made sense as a launch-time value; it broke once attribution turned honest reporting into a visible admission.
Short answer, where it wouldn't matter
6. Name a kind of VetLoop correction in this story that genuinely doesn't need any of this feedback design applied to it.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A one-off typo or phrasing quirk. The automatic diff shows it as noise, and it should stay noise; the design exists to catch patterns, not every keystroke.
Before you close the answer
Why this works
Tests whether your feedback design still works on the day someone has a real reason to stay quiet, not just on the easy days when everyone's honest for free.
Follow-up traps
"Doesn't full anonymity make it impossible to follow up on a real pattern?" Response: that's why the design keeps a private, engineer-only link back to the reporting vet, unattributed in every shared or peer-visible view, but still traceable for a genuine investigation.

"Won't vets just ignore the optional tap entirely?" Response: that's fine; the automatic diff is the floor the whole design depends on, and the tap is a bonus signal, never the thing everything else is built on.
If pressed
The actual pattern-detection threshold at Kestrel Hollow's network isn't a fixed count; it compares a given note pattern's edit rate against that same case type's historical baseline, so a genuinely rare condition doesn't need dozens of occurrences before it gets noticed.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more