Artifact critiqueIntermediateDesigning for Uncertainty & Trust / Feedback loops and data flywheels / #7

Critique a thumbs up and down widget as a feedback mechanism.

AUDIT the product is Corrigan Helpdesk, where DraftLine writes a reply an agent can send with one click, and a thumbs icon sits under every draft

Corrigan Helpdesk is a support-ticket platform. DraftLine, its AI assistant, writes a suggested reply under every incoming ticket. A thumbs-up and thumbs-down icon sits beneath each draft. Priya Kadam, senior support operations manager, owns the dashboard that turns those clicks into a number leadership quotes.

The direct answer
A thumbs up or down count is only worth anything next to its own response rate. Before quoting a percentage like "92% positive," force a fixed, representative sample of drafts to get rated no matter what, tag every rating to the exact model version that produced it, and always publish the score and the response rate side by side. A score with no denominator isn't a number. It's a slogan.
Do this, in order
  1. Force a representative sample of drafts to get rated, tagged to a model version, and always show the response rate next to the score.Why: a percentage with no denominator can hide almost anything, including the opposite of what it claims.
  2. Ask who is grading their own homework before quoting the number anywhere external.Why: the team that ships the feature and the team that reports its quality were the same six people.
  3. Split the score by ticket type before trusting a single average.Why: a 92% average was hiding a 58% on the ticket type that actually mattered most to renewal.
  4. Log every unrated draft as unknown, never as a silent positive.Why: folding silence into the denominator is how a self-selected number gets to look like a population number.
  5. Re-test the number on a fresh slice before it goes in a renewal deck, not after a customer asks.Why: the cheapest time to find a bad number is before someone else finds it for you.
  6. Watch the response rate itself over time, not just the score.Why: a shrinking response rate is a warning sign on its own, whatever the score still says.

How to answer this, stage by stage

Nobody is grading whether you can spot that thumbs widgets are flawed. Everyone knows that. They're grading whether you can say exactly what's missing and how you'd go check.

Stage 1
Scope it to one real dashboard
Say it like this
"I'll answer this for Corrigan Helpdesk's DraftLine dashboard, where a thumbs up or down sits under every AI-drafted reply an agent sees."
Why this works
Grounds a generic widget critique in one real report someone actually has to defend.
Stage 2
Say your structure out loud
Say it like this
"I'll use AUDIT. Ask who ran it, uncover the eval set behind it, demand the version pin, isolate what's missing, then test it myself."
Why this works
Signals a method instead of a loose list of complaints about the widget.
Stage 3
Ask who's grading their own homework
Say it like this
"The 92% lives in Corrigan's own renewal deck, built by the same team that ships DraftLine. Nobody outside that team checked it before it reached a customer."
Why this works
This is AUDIT's opening move, and the one most candidates skip entirely.
Stage 4
Uncover what "rated" actually means
Say it like this
"Only 14% of drafts ever get a click at all. The 92% is computed over that self-selected 14%, not over everything DraftLine ever sent an agent."
Why this works
A score means nothing until you can name the population it was scored over.
Stage 5
Demand the version and the date
Say it like this
"That number blends six months and three model versions, one of them retired back in March. A number with no version pin can't be checked again."
Why this works
Models get updated quietly. A stale version baked into a live number is a common, avoidable mistake.
Stage 6
Name what the report leaves out
Say it like this
"No split by ticket type, no range around the number, and every draft nobody rated just disappears instead of counting as unknown."
Why this works
What a report omits is usually more telling than what it shows.
Stage 7
Test it on a forced sample
Say it like this
"I'd pull one week where every single draft, no exceptions, ties to a rating. Run that and you'll get the real number, not the convenient one."
Why this works
AUDIT's strongest move: replicate the claim on your own slice before it reaches a signature.
Stage 8
Close on the one line
Say it like this
"A thumbs-up rate with no response rate next to it isn't a quality number. It's a highlight reel of whoever felt like clicking."
Why this works
Restates the direct answer, ready for a follow-up push.

Let's learn

What does a thumbs up actually tell you?

Corrigan Helpdesk sells a support-ticket platform to other companies. DraftLine, its AI assistant, reads an incoming ticket and writes a suggested reply an agent can send in one click or edit first. A thumbs-up and thumbs-down icon sits right under every draft.

Before anyone looked closely, Corrigan's leadership quoted "DraftLine drafts get a thumbs-up 92% of the time" in the sales deck and in every quarterly renewal conversation. It read like proof the AI was good.

The number in the deck, next to the number behind it
100% 50% 0% 92% Reported positive rate 14% Drafts that got any click
The 92% was never computed over every draft. It was computed over the one in seven drafts an agent bothered to click at all.

The turn: the drafts DraftLine gets wrong are not the real problem here. A wrong draft gets caught and rewritten in seconds. The real problem is a report built on a number nobody had checked the population behind.

The problem was never that DraftLine drafts are bad. A ninety-two with no denominator isn't a fact. It's a slogan wearing a percent sign.
Knowledge spark: what's self-selection bias? It's what happens when the people who bother to respond aren't a fair stand-in for everyone. If only agents with a strong opinion click, the number reflects strong opinions, not typical drafts.

At its worst: a customer's own vendor-risk team asks Corrigan to show the raw data behind the 92%, finds a self-selected 14% response rate and a 58% score on billing-dispute tickets specifically, and walks a three-year renewal to a competitor over a number Corrigan never actually checked.

The decision I would take back We built the dashboard to log a click when one happened, and quietly dropped every draft nobody rated instead of tracking the response rate itself. That made sense when DraftLine was new and almost every agent clicked something out of curiosity. It stopped making sense once clicking became optional enough that only the strongly-opinionated bothered.

What I would leave alone: the raw thumbs widget itself is still a fine, cheap pulse check for an agent's own team retro, where nobody's citing it externally and everyone in the room already knows its limits. The problem is only using it as an external proof point with no sampling design behind it.

The lesson: a feedback number isn't trustworthy because it's high. It's trustworthy because you can say exactly who it was measured over, on what date, on which version, and what happens when you go check it yourself.

Now here is the same thing as a story

The short version above is what you'd say defending the redesign to Corrigan's leadership. Read this one for how the gap actually got found.

Priya Kadam could smell a bad support metric from across the room. Seven years running Corrigan's support operations had taught her that any number that looks too clean usually has an asterisk nobody wrote down.

For the first year after DraftLine shipped, the 92% felt earned. Agents were curious about the new assistant, clicking thumbs up or down on nearly every draft just to see what would happen, and the score held steady in the low 90s.

Hand sketched flow diagram titled Where a click actually happens. Five steps: AI drafts reply, agent sends it, customer reads email, maybe clicks a thumb highlighted, counted as 92 percent.
Somewhere between "agent sends it" and "counted as 92%," most drafts quietly drop out of the picture.

Over the second year, clicking thinned out in three beats. First, agents stopped rating drafts they agreed with, since sending felt like approval enough. Then they stopped rating minor edits, since a small fix didn't feel worth a click. By the third beat, only a draft that felt clearly great or clearly bad got a thumb at all.

The trigger was small: a prospective enterprise customer's vendor-risk analyst, doing renewal due diligence, wrote back one line: "Can you show us the raw counts behind the 92%, not just the percentage?"

Knowledge spark: what's a version pin? The exact model version, tied to an exact date, that produced a given result. Models get quietly updated. Without a version pin, nobody can go back and check a number against what actually made it.

Priya pulled the raw counts herself. Fourteen out of every hundred drafts had ever gotten a click, across three different DraftLine versions, one of them retired back in March.

Hand sketched decision tree titled Reading the same 92 percent three ways. Root Thumbs up rate is high, branching to genuinely satisfied leads to good sign, clicked to dismiss fast leads to false positive, only happy ones bother leads to survivorship bias.
The same 92% supports three different stories. Only a forced sample tells you which one is true.
Fourteen percent of drafts were never a sample of anything. They were whoever happened to feel strongly enough to click.

Priya built a forced weekly sample: every draft tied to a randomly chosen ticket ID, no exceptions, got a rating from the agent handling it, tagged to the exact DraftLine version live that week.

Weekly response rate on DraftLine drafts, since launch
40% 20% 0% Week 1: 35% Week 8: 26% Week 16: 19% Week 24: 14%
The response rate had been quietly falling for six months before anyone thought to check it. The score never moved. The population behind it did.

The forced sample's real number came in at 71%, and it split hard by ticket type: 96% on simple order-status replies, 58% on billing disputes, the exact ticket type the vendor-risk analyst's own company sent Corrigan the most.

Run the same renewal conversation forward under the old dashboard: Corrigan quotes 92%, the analyst asks for raw counts, and the real 58% surfaces during due diligence instead of during a design review six months earlier.

I built that dashboard to be simple, one click, one number, because a clean metric felt like the whole point. It took one analyst's plain question to see that a number nobody can explain isn't clean. It's just unexamined.

AUDIT, the actual readNot a gut check. AUDIT is what tells you whether a report's headline number is evidence or decoration.

A
Ask who paid for it.
The 92% lives in Corrigan's own sales deck, built by the same team shipping DraftLine. Nobody outside that team checked it first.
The hardest step: most people never ask whose interest a good number serves.
U
Uncover the eval set.
Only 14% of drafts ever got a click. The 92% was computed over that self-selected slice, not over every draft DraftLine wrote.
A score means nothing until you can name the population behind it.
D
Demand the version pin.
Six months, three model versions, one retired in March, all blended into one number.
A number with no version and no date can't be checked again by anyone.
I
Isolate what's missing.
No split by ticket type, no confidence range, and every unrated draft vanished instead of counting as unknown.
What a report leaves out is usually the most informative part of it.
T
Test it yourself.
A forced weekly sample put the real number at 71%, and 58% specifically on billing disputes.
Replicate the claim on your own slice before it reaches a signature, not after.
Hand sketched labeled parts diagram titled What a 92 percent positive rate leaves out. Center document icon labeled 92 percent positive, with four callouts: why it was clicked, repeat clicks one person, no response counted as neutral, which model version.
Four gaps, and any one of them alone is enough to make the headline number meaningless.

The recap, one line per letter: ask is whose report this really is, uncover is the self-selected 14% behind the 92%, demand is the missing version pin across three model releases, isolate is the missing split and missing range, and test is the forced sample that found the real 71%.

And if you want to be sure it really works, try it somewhere elseSame five letters, a city permits office instead of a support desk. A different building, the same audit.

Thornecrest's permits office runs an AI chatbot that answers zoning questions for residents applying for renovation permits. A thumbs-up widget sits at the end of every chat, and the office's annual public report cites "89% found this helpful."

Mapped onto AUDIT: ask is whether the office's own communications team, who wants a good headline for the annual report, is the same team that ran the analysis. Uncover is finding that only 9% of residents who chatted ever rated the conversation, almost all of them people who got a fast, simple answer. Demand is checking which chatbot model version produced the number, since the office switched providers mid-year without updating the report. Isolate is noticing there's no split between "got an answer" and "got redirected to call the office," which residents rate very differently. Test is pulling a forced sample of one month's chats, rating every single one regardless of whether the resident bothered, and finding the real helpful rate closer to 61%, mostly dragged down by complex renovation cases the chatbot couldn't actually resolve.

Hand sketched icon list titled Five questions before you trust this widget. Items: who benefits if this number looks good, what counts as a rated reply, which model version made this number, what got left out of the report, would it replicate on a fresh sample.
The same five questions work on a support dashboard or a city hall report. Only the building changes.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "ask for the response rate next to the score, and who ran the analysis," and stop.
Cost: there's no time to build a full forced-sample pipeline before the next board meeting. Say so honestly, and hand-audit one week's raw data first, since a small honest check beats a big unchecked slogan.
The model gets better, for real: if DraftLine's draft quality genuinely improves, that's still not a reason to skip the audit. A better model just makes the missing response rate a smaller problem, not a solved one.

Where people run it wrong.
They quote a percentage from an internal dashboard externally without ever asking who built the dashboard.
They treat "no click" as if it were a quiet yes, instead of tracking it as unknown.
They average across model versions and time periods that shouldn't be blended into one headline number.

How to use it live. When someone hands you a report with a clean percentage in it, ask yourself one question first: could I name the exact population, version, and date behind this number, right now. If you can't, that's the audit, and that's your answer.

Hand sketched comparison diagram titled Who actually left feedback. Left panel a person icon labeled Rated it, caption small vocal group. Right panel a box icon labeled Never clicked, caption most customers silent.
A widget that only hears from a small, opinionated slice isn't measuring the room. It's measuring who felt like talking.
Hand sketched quadrant titled Sorting feedback signals by what they prove. Axes how easy to game from hard to fake to easy to fake, and how much detail it gives from almost none to a lot. Thumbs up down sits easy to fake and almost no detail. Written complaint and escalated to human sit hard to fake and a lot of detail.
A thumb is the fastest signal to collect and the easiest one to game. That trade is fine, as long as you remember you made it.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "critique this feedback report" question?
Tap to flip
ANSWER
AUDIT: ask who ran it, uncover the eval set, demand the version pin, isolate what's missing, test it yourself. The ask step is the hardest one.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Priya Kadam, senior support operations manager at Corrigan Helpdesk, seven years into running support and owner of the DraftLine dashboard.
3 · THE GAP FOUND
What did the forced sample reveal that the reported 92% hid?
Tap to flip
ANSWER
Only 14% of drafts had ever been rated, and the real forced-sample score was 71% overall, dropping to 58% on billing-dispute tickets specifically.
4 · THE HARDEST STEP
What's the hardest step in AUDIT, and why?
Tap to flip
ANSWER
Asking who ran the analysis. It's the step most people skip, since it means questioning your own team's number, not just a stranger's.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Logging a click only when one happened, and never tracking the response rate itself, so a shrinking, self-selected sample looked identical to a healthy one.
6 · THE NUMBER
Fill in the blank: the response rate on DraftLine drafts fell from 35% in week 1 to ___% by week 24.
Tap to flip
ANSWER
14%. The score stayed near 92% the whole time. Only the size and shape of the population behind it changed.
7 · THE REPLAY
Same renewal conversation, redesigned dashboard. What changes?
Tap to flip
ANSWER
The 58% on billing disputes surfaces in an internal design review months earlier, instead of surfacing during a customer's due diligence right before a signature.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the parallel?
Tap to flip
ANSWER
Thornecrest's permits-office chatbot, citing "89% helpful." Same AUDIT shape: a 9% response rate hiding under the headline, found the same way.

Check yourself Score: 0 / 0

True or false
1. True or false: the 92% thumbs-up rate was calculated over every draft DraftLine ever produced.
  • True
  • False
Show hint
Look at "uncover the eval set."
Show answer
False. It was calculated only over the 14% of drafts that ever got any click at all, a self-selected slice, not the full population.
Multiple choice
2. Why doesn't collecting even more of the same kind of clicks fix this number?
  • A. Because thumbs widgets are always technically broken.
  • B. Because more clicks from the same self-selected 14% still isn't a fair stand-in for the other 86% who never click.
  • C. Because DraftLine's model needs to be retrained first.
  • D. Because agents are not allowed to see the dashboard.
Show hint
Think about what a bigger pile of the same biased sample actually proves.
Show answer
B. Volume doesn't fix a biased sample. Only a forced, representative sample changes what the number actually measures.
Fill in the blank
3. Fill in the blank: the decision Priya would take back is that the dashboard never tracked the ___ itself.
Show hint
Look at "the decision I would take back."
Show answer
Response rate. Without it, a shrinking, self-selected sample looked exactly like a healthy, stable one on the dashboard.
Short answer, where it wouldn't matter
4. Name one place a raw, unaudited thumbs count is still fine to use.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: An internal team retro, where nobody's citing the number externally and the room already knows its limits. The risk is only in quoting it as external proof.
Short answer, apply it yourself
5. Think of a rating or review number you've trusted somewhere (an app store score, a restaurant rating). What question from AUDIT would you ask about it now?
Show hint
Try "uncover the eval set": who actually left that rating, and who didn't bother?
Show answer
Model answer: Most public ratings share this exact flaw: they only count people who felt strongly enough to respond, which is rarely the full, typical population.
Multiple choice
6. Given a 92% score computed over just a 14% response rate, what's the most honest way to describe that number?
  • A. A reliable population-level satisfaction score.
  • B. A score about the opinions of the 14% who clicked, not a measurement of every draft sent.
  • C. Proof the AI model itself has gotten worse.
  • D. A number that should be rounded up before reporting it.
Show hint
Go back to what "uncover the eval set" actually establishes.
Show answer
B. The 92% is real, it just describes a narrow, self-selected slice, not the full population the deck implied it covered.
Before you close the answer
Why this works
Tests whether you interrogate a report's sampling before trusting its headline score, instead of accepting a clean-looking percentage at face value.
Follow-up traps
"Isn't 71% worse than 92%, so did DraftLine actually get worse?" Response: no, the model didn't change; only the sample changed from self-selected to forced, so the real number was always around 71%, we just hadn't measured it honestly before.

"Won't forcing a rating on every draft slow agents down?" Response: it only applies inside one bounded weekly random sample, not every draft in production, so the added friction is small and temporary, not a permanent tax on the whole team.
If pressed
The real forced sample rotates automatically, chosen by a hashed ticket ID so agents can't predict in advance which tickets will land in that week's sample, which is what keeps the number honest instead of gameable.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more